← Open Source
Enderfga

claw-orchestrator

Run Claude Code, Codex, Antigravity, Grok Build and OpenCode behind one runtime, and a run only counts as done when checks it didn't write pass. Cross-engine fan-out, councils, session handoff, spend caps, MCP server, ACP agent.

AI EngineeringGive agent toolsBuild agent workflow (SDK)TypeScript
Open on GitHub
Momentum
+-1stars in 24 hours-0.2%
586
Stars
102
Forks
+1
This week
13
Contributors
Created 2026-01-30 · Updated 2026-10-07 · #18561 today
Top developers
README

Claw Orchestrator

Claw Orchestrator

Run Claude Code, Codex, Antigravity, Grok Build and OpenCode behind one runtime, and a run only counts as done when checks it didn't write pass.

npm version npm downloads CI License: MIT

An agent reports that all tests pass; the runtime's own check refutes it. Five engines, three runs each, tallied by verdict.

One bug, one acceptance contract, five engines, three runs each, first attempt only. The runtime ran the checks itself: 9 verified, 4 refuted, 2 errored (Grok Build hit its free-tier usage limit). One task, not a benchmark. Real runs, recorded locally — how the film was made, and how to re-run it.

▶ The 40-second film, with sound

https://github.com/user-attachments/assets/5bf9bea8-e5ac-4632-a292-a80a45831dff

Coding agents grade their own work: "All tests pass ✓". Claw Orchestrator runs them as persistent sessions behind one API and keeps the verdict out of their hands. You declare the checks, the runtime executes them and reads the exit codes, and the evidence stays on disk.

npm install -g @enderfga/claw-orchestrator
clawo serve &    # local runtime + dashboard on 127.0.0.1:18796

clawo solve "The test fails. Fix price.js." --engine codex \
  --check "npm test" --check "node ../holdout/price.spec.mjs" --wait
clawo verify     # each check, why it failed, the files that changed
clawo runs --since 1h   # every turn on every engine: model, tokens, cost, verdict

What you get

You want to… Claw Orchestrator gives you
Stop trusting "All tests pass ✓" Acceptance contracts — commands, HTTP probes, screenshots, diff policy, file assertions — that the runtime runs itself. A run with a contract ends verified, refuted or unverified, never "the agent said so". Editing the tests a contract runs refutes the run. → verification
Use more than one vendor Persistent sessions for Claude Code, Codex, Antigravity, Grok Build, OpenCode and custom CLIs behind one interface. clawo fanout asks several of them the same question at once. → multi-engine
Switch engines mid-task session_handoff replays a live conversation onto another engine in the same workspace. → sessions
Know what it cost A durable per-turn ledger across engines (clawo runs) and a maxBudgetUsd cap the runtime enforces on every engine. → observability
Runs that survive a crash Durable workflows: every transition is checkpointed; resume, retry, timeout, cancel and steer come from the kernel. → workflow
Call it from your own tools An MCP server (clawo-mcp), an ACP agent (clawo acp) for Zed, JetBrains and other ACP editors, an OpenAI-compatible endpoint, a TypeScript SDK, and the CLI. → integrations

How it relates to other tools. claude-squad, Conductor and Vibe Kanban are interfaces for running agents side by side in worktrees; Claude Code subagents and Codex's parallel attempts stay within one vendor. Claw Orchestrator is a headless runtime: several vendors behind one API, durable workflows, and a verification step the agents cannot edit. They can be used together.


Quick Start

Requires Node 22+ and at least one engine CLI installed and logged in (claude, codex, agy, grok or opencode).

npm install -g @enderfga/claw-orchestrator
clawo serve   # dashboard at http://127.0.0.1:18796/dash

The server writes an access token to ~/.openclaw/server-token; the CLI reads it from there. Open http://127.0.0.1:18796/login?token=&redirect=/dash once to sign in the browser.

From a terminal:

clawo solve "Fix the failing tests" --engine claude --check "npm test" --wait
clawo fanout "How would you speed up this test suite?" --engines claude,codex,grok --synthesize --wait
clawo workflow list

From Claude Code, or any MCP host:

claude mcp add clawo -- clawo-mcp

Then ask for workflow_start with template solve, a task, and a contract such as {"checks":[{"type":"command","cmd":"npm","args":["test"]}]}.

From TypeScript:

import { SessionManager } from '@enderfga/claw-orchestrator';

const manager = new SessionManager();
await manager.startSession({ name: 'fix-tests', engine: 'claude', cwd: '/project' });
const result = await manager.sendMessage('fix-tests', 'Fix the failing tests');

All features

Capability What it does Reference
Persistent Sessions Long-lived coding agents kept alive across requests, with full context, tool, model, and worktree control. sessions.md
Multi-Engine Runtime One interface over Claude Code, Codex, Antigravity (agy), Grok Build, OpenCode, and arbitrary custom CLIs. multi-engine.md
Session Handoff Move a live conversation to another engine or model — a stuck Claude session into Codex, an expensive model into a cheaper one. The new session picks up where the old one stopped, in the same workspace; the old one keeps running. sessions.md
Multi-Agent Council Parallel agents in isolated git worktrees, voting on consensus until they agree. council.md
Fan-out Run one task across N engine/model agents in parallel and collect their answers, with an optional synthesis pass — the cross-engine best-of-N / diverse-perspective primitive (no rounds or worktrees). tools.md
ultracode session_start({ ultracode: true }) lets Claude orchestrate a dynamic JS workflow and fan out to subagents per task (Claude engine). tools.md
Autoloop Three-agent autonomous workspace iteration with independent engine/model selection for Planner, Coder, and Reviewer. Chat with the Planner; it spawns Coder + Reviewer into a self-iterating subloop and pushes you on regression, target-hit, or decision points. autoloop.md
Ultraapp A three-agent Opus council turns a short structured interview into a deployed web app — Tailwind UI, BYOK, file-queue runtime, smoke test, all live at localhost:19000/forge//. ultraapp.md
Embedded Dashboard Three-tab UI for Autoloop, Council, and Forge with sidebar lifecycle controls, per-run live event streaming, and cookie-based auth via a /login redirect. dashboard.md
OpenAI-Compatible Proxy POST /v1/chat/completions accepts OpenAI-format requests and routes them to persistent sessions on the engine the model name selects, streaming replies back in OpenAI shape. Point any OpenAI-SDK client or webchat at the orchestrator without changing call sites. openai-compat.md
Durable Run Kernel Declarative workflows over agent / fanout / council / verifier / human_gate / router / subflow / autoloop / ultraapp_* nodes. Every state transition is checkpointed, so a run survives a process restart and resumes at the node boundary. Retry, per-node timeout, cancel, steer, and bounded loops come from the kernel. workflow.md
Verification Plane Acceptance contracts the runtime executes itself — commands, HTTP probes, screenshots, diff policy, file assertions — producing an evidence bundle on disk. A run carrying a contract cannot reach completed unless it passes, and one without a contract completes as unverified rather than claiming success. The tests a contract runs are held to what the run started with, so a run cannot pass by editing them. verification.md
Run Ledger & Spend Caps Every turn on every engine is appended to a durable JSONL ledger — engine, model, tokens, cost, duration, and the council/fanout/autoloop it belonged to — queryable with clawo runs after a restart. Rows carry both the engine's self-report (ok) and the runtime's own measurement (verified), kept apart. maxBudgetUsd is enforced by the runtime, so a cap holds on Codex, Grok, agy and OpenCode too, not just Claude Code. observability.md

The full 78-tool surface is enumerated in tools.md.


Integrations

Standalone CLI

clawo serve                                            # dashboard + HTTP server on :18796
clawo session-start fix-tests --engine claude --cwd .  # start a session
clawo session-send fix-tests "Fix the failing tests"   # send into it

Every command is documented in cli.md.

OpenClaw Plugin

curl -fsSL https://raw.githubusercontent.com/Enderfga/claw-orchestrator/main/install.sh | bash

Installs via npm, registers the plugin in ~/.openclaw/openclaw.json, restarts the gateway. All 78 tools become available to every OpenClaw agent.

Model Context Protocol Server

npm install -g @enderfga/claw-orchestrator   # clawo-mcp is now on PATH

Register clawo-mcp with any MCP-compatible host: Hermes Agent, Claude Desktop, Cursor, Cline, Continue, Zed, Windsurf, Goose, and others. Per-host stdio-config snippets and the CLAWO_MCP_TOOLS allowlist for tight tool budgets are in mcp.md.

Agent Client Protocol Agent

clawo acp        # or the dedicated binary: clawo-acp

MCP gives tools to an agent; ACP makes you be the agent. clawo acp speaks Agent Client Protocol over stdio, so Zed, JetBrains, Neovim, Emacs, the VS Code ACP extension — or dsh via its subagent-acp provider — can drive Claw Orchestrator as their coding agent.

The model selector is grouped by engine, so one dropdown holds Claude, Codex and Grok models at once and switching it switches engine mid-session; /council, /ultraplan and /ultrareview run multi-agent orchestrations from the chat box. Setup, the dsh YAML block, and the cancellation and permission limitations are in acp.md.


Engine Compatibility

Engine CLI Tested Version
Claude Code claude 2.1.292
Codex codex 0.160.1
Antigravity agy 1.3.1
Grok Build grok 1.0.46
OpenCode opencode 1.18.35
Custom CLI any —

Any coding CLI that runs as a subprocess can be wired up as a custom engine — see multi-engine.md.


How the engine table stays current

The versions above are not typed in — they are what the weekly sweep last ran. scripts/sweep.ts measures each core engine's installed, pinned and upstream version, diffs the flags the wrapper passes against the binary's --help, runs one live turn through the real wrapper class, smokes the ACP and MCP entry points, and checks src/models.ts against both vendors' published price tables. It contains no model call, so the check that reports a wrapper as broken does not share the wrapper's failure modes.

scripts/sweep-workflow.json wraps the same script as a durable run on this project's own kernel: verifier → router → an agent that drafts the alignment on a sweep/ branch → a human gate. A person reviews and merges; the orchestrator does not edit its own code unattended.

Contributing

See CONTRIBUTING.md. Run npm run build && npm run lint && npm run format:check && npm run test before submitting.

License

MIT — see LICENSE.