ISSUE 2026-09-23 · WEDNESDAY, SEPTEMBER 23 clawcodex · v1.7.0
80.9% on Terminal-Bench 2.1 · open source · multi-provider

A token-efficient coding agent for your terminal, on any model.

ClawCodex is a token-efficient Python rebuild of Claude Code — 230K lines of pure Python. A genuine tool-calling loop, a streaming REPL, skills, and one runtime in front of Anthropic, OpenAI, Z.ai GLM, MiniMax, OpenRouter, and DeepSeek, where prefix-cache reuse makes long DeepSeek sessions over 200× cheaper to run.

Mission for this build Keep the real Claude Code architecture, but make it Python-native and open to every provider. Stream replies, call tools, persist sessions, and let you choose the most flexible, cost-effective model stack for agentic coding. Now 1.0; shipping weekly.

Maintained by ClawCodex Team · v1.7.0 · 30 providers

30+ tools files, bash, web, agents, MCP
30 providers one runtime, swap any time
80.9% Terminal-Bench 2.1 on Opus 5 · k=1, ≈3rd
200× cost saving token-efficient on DeepSeek
ClawCodex running in a terminal: a streaming REPL with visible tool activity and a status row.
The actual inline REPL: streaming replies, visible tool calls, and a tool-aware status row.
Section 01 · why

A model writes text. An agent leaves consequences.

ClawCodex puts a real runtime around the model: a working loop, every provider, and code you can read and extend.

real runtime

A working agent, not a source dump

A genuine tool-calling loop, a streaming REPL, session history, and multi-turn execution. Ported from the real Claude Code TypeScript architecture and shipped as a CLI you actually run.

any provider

One runtime in front of every model

Claude Code targets Claude models only. ClawCodex puts Anthropic, OpenAI, Z.ai GLM, MiniMax, OpenRouter, and DeepSeek behind the same loop — so you swap vendor, region, and price tier without giving up tools or skills.

built to hack on

Readable Python you can extend

Idiomatic Python with full type hints, real test suites, and markdown-driven SKILL.md extensibility. Fork it, add a tool or a skill, and make it yours. MIT licensed.

Section 02 · what's inside

The full surface, in your terminal.

Streaming, skills, permissions, sessions, MCP, and a scriptable headless mode — the same loop whichever model you point it at.

Terminal, desktop, or browser tab

clawcodex web serves the browser client off the same clawcodex serve gateway the desktop app uses — one in-process agent, one config, one session store. A session started in the terminal resumes in a tab, and its Trajectory view accounts for every request: per-step tokens, TTFT, throughput, cache-hit rate, and model time separated from tool time.

Nano mode — a deliberately smaller harness

clawcodex --nano trades the full tool surface for six tools and a ~2,000-token fixed payload instead of ~17,000. On terminal-bench 2.1 it solved the same tasks as default mode at 4.4× lower cost, and faster. Default-off and process-global: without the flag, nothing about your setup changes.

Token-efficient on DeepSeek

ClawCodex keeps the request prefix byte-stable so DeepSeek's prompt cache covers your whole system + tools + history span. Cache-hit input bills at about $0.0435 per 1M tokens, over 200× cheaper than Claude Fable 5, so long agentic sessions cost pennies. The longer you code, the more you save.

Streaming agent experience

True API streaming for direct replies plus richer streaming during tool-driven loops. Toggle live output with /stream and re-render clean Markdown with /render-last.

Programmable skill runtime

Markdown-based SKILL.md slash commands with named arguments and per-skill tool limits. Project skills and user skills, loaded from .clawcodex/skills.

REPL by default, TUI on demand

An inline prompt_toolkit + Rich REPL with history, tab completion, and multiline input. Launch the Textual TUI any time with clawcodex --tui or /tui.

Permission modes

Plan (read-only), acceptEdits, dontAsk, and an explicit bypass for sandboxes. The REPL, TUI, and headless -p mode all honor the same gates.

Sessions you can resume

Save and reload conversations locally. auto_save writes each session; max_history caps retained turns. Everything stays on your machine.

MCP + hooks

Model Context Protocol tools and resources, plus pre/post tool-use lifecycle hooks for shell, prompt, agent, and HTTP automation.

Scriptable headless mode

Run -p for one-shot prompts, with --output-format json or stream-json for pipes, CI, and agent-to-agent workflows.

Image handling

A TS-parity Read pipeline sniffs magic bytes and resizes to API limits. @-mention an image to inline it, with Anthropic-to-OpenAI block translation for vision-capable backends.

Section 03 · any provider

One runtime in front of every model.

Swap vendor, region, and price tier without giving up tools or skills. All 30 providers ship today.

DeepSeek default · prefix cache
Anthropic Claude
OpenAI GPT-5.4
Gemini Google · 2.5 Pro
Z.ai GLM GLM-5.1 / 5.2
MiniMax MiniMax-M2.7
OpenRouter multi-vendor proxy
NVIDIA NIM
Together AI
Fireworks AI
Novita AI
SiliconFlow
SiliconFlow (China)
Moonshot / Kimi
DeepInfra
Hugging Face
Volcengine Ark
StepFun
Arcee AI
AtlasCloud
Xiaomi MiMo
Wanjie Ark
Meta muse-spark-1.1
Groq
Cerebras
Baseten
xAI (Grok) grok-4.5
Ollama local · no key
vLLM local · no key
SGLang local · no key

Seven native integrations (DeepSeek, Anthropic, OpenAI, Gemini, Z.ai GLM, MiniMax, OpenRouter) plus an OpenAI-compatible gateway registry and local servers that need no key — 30 in all. Add any other OpenAI-compatible endpoint with a custom base_url; keys resolve from config or the vendor's standard env var (e.g. TOGETHER_API_KEY).

swap models
# point a single run at a different backend
 clawcodex --provider deepseek --model deepseek-v4-pro
 clawcodex --provider zai --model glm-5.2 -p "refactor utils.py"
Section 04 · how it works

A reviewable line from prompt to patch.

Read code, call tools behind permission gates, leave real evidence, and resume across sessions.

Prompt

You type in the REPL or pipe a prompt with -p. Plan mode is read-only by default.

Stream

The selected provider streams the reply; tool calls overlap with the stream.

Tools

Reads, edits, bash, and web run behind permission gates, leaving real evidence.

Resume

Results feed the next turn. Sessions save locally so long work survives.

under the hood · six abstractions

Query loop

The heartbeat that orchestrates model calls and tool execution.

Tool system

30+ tools: read files, run bash, search the web, drive sub-agents and MCP.

Tasks

Background workflows and agent orchestration, journaled and resumable.

Two-tier state

A bootstrap layer and an app layer, kept independent by design.

Memory

Relevance scanning over project, user, and team CLAUDE.md files.

Hooks

Lifecycle events with shell, prompt, agent, and HTTP executors.

Read ARCHITECTURE.md →

Section 05 · swe-bench verified

clawcodex beats openclaude on the same model.

Full SWE-bench Verified split (499 instances), both agents driven by Gemini 2.5 Pro under one standardized harness.

clawcodex
58.2% 291/499
openclaude
53.0% 265/499
241 Both solved 50 Only clawcodex 24 Only openclaude 184 Neither
Reproduce locally — see eval/README.md for the full workflow.
Section 06 · proof

Every demo was built by ClawCodex itself.

Same CLI you just installed, same agent loop, same tools. No hand-edits.

React 18 + Vite + Vitest

CRM app

A mini CRM with contacts, deals, a dashboard, and a full test suite.

demos/crm-app
React 18 + Vite + Router

LinkedIn-style feed

Profile, network, jobs, and messaging in a familiar feed layout.

demos/linkedin-app
React + three.js

Minecraft sandbox

A browser voxel sandbox with terrain, mining, a HUD, and player controls.

demos/minecraft-app
Static HTML/CSS/JS

World Cup 2026 intro

Animated hero, live countdown, host nations, and 16 stadiums. Built with GLM-5.2.

demos/wc26-intro
Section 07 · new in v1.7.0

Shipping weekly.

v1.7.0 rebuilds the multi-agent layer. One session-scoped supervisor now admits, tracks and interrupts every subagent — foreground, background, team and workflow — and can pause new spawns; and agent teams actually run: TeamCreate makes your session the leader, and named teammates stay alive between assignments, message each other, and share a dependency-aware task board. Background workers genuinely resume with their history.

agents#950

Persistent teams — teammates that stay, talk, and share a board

TeamCreate makes your session the leader; each named Agent call adds a teammate that stays alive between assignments with its context intact. SendMessage routes findings to a named peer or to the leader, a shared and locked task board hands out work in dependency order, plan approvals and shutdowns follow a matched request/response protocol, and interrupting a worker withdraws its pending permission prompt. Teams run in-process, one per workspace.

agents#915, #933, #950

One supervisor for every agent

Foreground delegations used to register nowhere, so nothing could list or stop them, and the TUI's agents overlay called three RPCs that had no backend. A session-scoped supervisor now admits every worker behind two configurable backstops (32 concurrent, depth 3), with live status, per-agent interrupt and a session-wide pause on new spawns in the TUI overlay and in the web client's subagent list. A refused spawn comes back as a tool error the model can act on.

agents#950

Worker lifecycles that hold

A follow-up to a finished background worker used to flip it back to running without ever starting a model loop. Resume now reloads the worker's history under the same ID, a correction accepted while it was finishing is no longer left unread, notifications reach the session and parent that own them, worktree isolation runs in a real Git checkout or fails before the model runs, and ending a session interrupts every worker it owns and waits a bounded time for them to stop. Verified end to end with a scripted provider driving real query loops and WebSocket connections, plus one live DeepSeek smoke run.

also#903, #905, #913, #917, #922, #925, #930, #947, #949

Also in 1.7.0

DeepSeek-V4.1-Flash becomes the DeepSeek default, with /cost following its peak/off-peak card; a model or effort pick is saved as your default on every interface; opt-in cost-aware auto-compaction; ChatGPT-subscription model discovery; and in the web client, subagents in the header with a child view per run, attachments of any file type, and saved sessions that open in milliseconds instead of ~45 s.

Full changelog →   All activity →

join in

A small project. Your patch matters.

No CLA. No sponsor lockouts. Issues triaged in the open, releases cut from main, and the maintainer reads everything. Bring a real test and prose that tells the reviewer what you were thinking.