A working agent, not a source dump
A genuine tool-calling loop, a streaming REPL, session history, and multi-turn execution. Ported from the real Claude Code TypeScript architecture and shipped as a CLI you actually run.
ClawCodex is a token-efficient Python rebuild of Claude Code — 230K lines of pure Python. A genuine tool-calling loop, a streaming REPL, skills, and one runtime in front of Anthropic, OpenAI, Z.ai GLM, MiniMax, OpenRouter, and DeepSeek, where prefix-cache reuse makes long DeepSeek sessions over 200× cheaper to run.
ClawCodex puts a real runtime around the model: a working loop, every provider, and code you can read and extend.
A genuine tool-calling loop, a streaming REPL, session history, and multi-turn execution. Ported from the real Claude Code TypeScript architecture and shipped as a CLI you actually run.
Claude Code targets Claude models only. ClawCodex puts Anthropic, OpenAI, Z.ai GLM, MiniMax, OpenRouter, and DeepSeek behind the same loop — so you swap vendor, region, and price tier without giving up tools or skills.
Idiomatic Python with full type hints, real test suites, and markdown-driven SKILL.md extensibility. Fork it, add a tool or a skill, and make it yours. MIT licensed.
Streaming, skills, permissions, sessions, MCP, and a scriptable headless mode — the same loop whichever model you point it at.
clawcodex web serves the browser client off the same clawcodex serve gateway the desktop app uses — one in-process agent, one config, one session store. A session started in the terminal resumes in a tab, and its Trajectory view accounts for every request: per-step tokens, TTFT, throughput, cache-hit rate, and model time separated from tool time.
clawcodex --nano trades the full tool surface for six tools and a ~2,000-token fixed payload instead of ~17,000. On terminal-bench 2.1 it solved the same tasks as default mode at 4.4× lower cost, and faster. Default-off and process-global: without the flag, nothing about your setup changes.
ClawCodex keeps the request prefix byte-stable so DeepSeek's prompt cache covers your whole system + tools + history span. Cache-hit input bills at about $0.0435 per 1M tokens, over 200× cheaper than Claude Fable 5, so long agentic sessions cost pennies. The longer you code, the more you save.
True API streaming for direct replies plus richer streaming during tool-driven loops. Toggle live output with /stream and re-render clean Markdown with /render-last.
Markdown-based SKILL.md slash commands with named arguments and per-skill tool limits. Project skills and user skills, loaded from .clawcodex/skills.
An inline prompt_toolkit + Rich REPL with history, tab completion, and multiline input. Launch the Textual TUI any time with clawcodex --tui or /tui.
Plan (read-only), acceptEdits, dontAsk, and an explicit bypass for sandboxes. The REPL, TUI, and headless -p mode all honor the same gates.
Save and reload conversations locally. auto_save writes each session; max_history caps retained turns. Everything stays on your machine.
Model Context Protocol tools and resources, plus pre/post tool-use lifecycle hooks for shell, prompt, agent, and HTTP automation.
Run -p for one-shot prompts, with --output-format json or stream-json for pipes, CI, and agent-to-agent workflows.
A TS-parity Read pipeline sniffs magic bytes and resizes to API limits. @-mention an image to inline it, with Anthropic-to-OpenAI block translation for vision-capable backends.
Swap vendor, region, and price tier without giving up tools or skills. All 30 providers ship today.
Seven native integrations (DeepSeek, Anthropic, OpenAI, Gemini, Z.ai GLM, MiniMax, OpenRouter) plus an OpenAI-compatible gateway registry and local servers that need no key — 30 in all. Add any other OpenAI-compatible endpoint with a custom base_url; keys resolve from config or the vendor's standard env var (e.g. TOGETHER_API_KEY).
# point a single run at a different backend
clawcodex --provider deepseek --model deepseek-v4-pro
clawcodex --provider zai --model glm-5.2 -p "refactor utils.py"
Read code, call tools behind permission gates, leave real evidence, and resume across sessions.
You type in the REPL or pipe a prompt with -p. Plan mode is read-only by default.
The selected provider streams the reply; tool calls overlap with the stream.
Reads, edits, bash, and web run behind permission gates, leaving real evidence.
Results feed the next turn. Sessions save locally so long work survives.
The heartbeat that orchestrates model calls and tool execution.
30+ tools: read files, run bash, search the web, drive sub-agents and MCP.
Background workflows and agent orchestration, journaled and resumable.
A bootstrap layer and an app layer, kept independent by design.
Relevance scanning over project, user, and team CLAUDE.md files.
Lifecycle events with shell, prompt, agent, and HTTP executors.
Full SWE-bench Verified split (499 instances), both agents driven by Gemini 2.5 Pro under one standardized harness.
Same CLI you just installed, same agent loop, same tools. No hand-edits.
A mini CRM with contacts, deals, a dashboard, and a full test suite.
demos/crm-app
Profile, network, jobs, and messaging in a familiar feed layout.
demos/linkedin-app
A browser voxel sandbox with terrain, mining, a HUD, and player controls.
demos/minecraft-app
Animated hero, live countdown, host nations, and 16 stadiums. Built with GLM-5.2.
demos/wc26-intro
v1.7.0 rebuilds the multi-agent layer. One session-scoped supervisor now admits, tracks and interrupts every subagent — foreground, background, team and workflow — and can pause new spawns; and agent teams actually run: TeamCreate makes your session the leader, and named teammates stay alive between assignments, message each other, and share a dependency-aware task board. Background workers genuinely resume with their history.
TeamCreate makes your session the leader; each named Agent call adds a teammate that stays alive between assignments with its context intact. SendMessage routes findings to a named peer or to the leader, a shared and locked task board hands out work in dependency order, plan approvals and shutdowns follow a matched request/response protocol, and interrupting a worker withdraws its pending permission prompt. Teams run in-process, one per workspace.
Foreground delegations used to register nowhere, so nothing could list or stop them, and the TUI's agents overlay called three RPCs that had no backend. A session-scoped supervisor now admits every worker behind two configurable backstops (32 concurrent, depth 3), with live status, per-agent interrupt and a session-wide pause on new spawns in the TUI overlay and in the web client's subagent list. A refused spawn comes back as a tool error the model can act on.
A follow-up to a finished background worker used to flip it back to running without ever starting a model loop. Resume now reloads the worker's history under the same ID, a correction accepted while it was finishing is no longer left unread, notifications reach the session and parent that own them, worktree isolation runs in a real Git checkout or fails before the model runs, and ending a session interrupts every worker it owns and waits a bounded time for them to stop. Verified end to end with a scripted provider driving real query loops and WebSocket connections, plus one live DeepSeek smoke run.
DeepSeek-V4.1-Flash becomes the DeepSeek default, with /cost following its peak/off-peak card; a model or effort pick is saved as your default on every interface; opt-in cost-aware auto-compaction; ChatGPT-subscription model discovery; and in the web client, subagents in the header with a child view per run, attachments of any file type, and saved sessions that open in milliseconds instead of ~45 s.
No CLA. No sponsor lockouts. Issues triaged in the open, releases cut from main, and the maintainer reads everything. Bring a real test and prose that tells the reviewer what you were thinking.