← Open Source
Dicklesworthstone

pi_agent_rust

High-performance AI coding agent CLI written in Rust with zero unsafe code

ApplicationsCodingRust
Open on GitHub
Momentum
+6stars in 24 hours+0.3%
1.84k
Stars
218
Forks
+32
This week
2
Contributors
Created 2026-02-02 · Updated 2026-10-05 · #1365 today
Top developers
README

Pi Agent Rust

pi_agent_rust

pi_agent_rust - Native AI coding agent CLI written in Rust

Why Should You Care? • TL;DR • Methodology • Quick Start • Features • Installation • Commands • Configuration

Rust 2024 License: MIT + Rider No Unsafe Code

# Install latest release
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/pi_agent_rust/main/install.sh?$(date +%s)" | bash

The Problem

You want an AI coding assistant in your terminal, but existing tools are:

  • Slow to start: Managed runtimes can add noticeable startup overhead
  • Resource intensive: Electron apps or heavy runtimes can add substantial overhead
  • Unreliable: Streaming breaks, sessions corrupt, tools fail silently
  • Hard to extend: Closed ecosystems or complex plugin systems

The Solution

pi_agent_rust is a from-scratch Rust port of Pi Agent by Mario Zechner (made with his blessing!). Official release archives install the single end-user binary pi, with streaming responses and 36 built-in tools (19 in the default --tools list; 14 always in the model's schema, the rest reachable through the xdev dispatcher or enabled in settings).

Current product direction

This project is no longer trying to be a strict drop-in replacement for the legacy TypeScript Pi. That target became both impractical and undesirable as the legacy implementation evolved. Legacy Pi remains useful historical context, but it is not our compatibility authority or definition of completeness.

OMP is the closer reference for where the product is going: feature surface, agent workflows, look and feel, and overall UI/UX. Pi Rust still chooses Rust-native architecture and may intentionally differ from both projects when that produces a simpler, safer, or better coding agent. Historical drop-in and parity artifacts remain in the repository as records; they do not gate product work, releases, or user-facing claims.

All repository quality checks, builds, and releases run through Doodlestein Self-Releaser (DSR). Contributors must not invoke Cargo or RCH directly, and GitHub Actions is never an execution or evidence authority for this project. Use dsr quality --tool pi_agent_rust for the registered quality recipe and the canonical DSR build/release commands documented in docs/releasing.md.

For signed published payload verification and offline CLI smoke checks, see Published release consumer checks.

Rather than a direct line-by-line translation, this port builds on two purpose-built Rust libraries:

  • asupersync: A structured concurrency async runtime with built-in HTTP, TLS, and SQLite
  • rich_rust: A Rust port of Rich by Will McGugan, providing beautiful terminal output with markup syntax
# Start a session
pi "Help me refactor this function to use async/await"

# Continue a previous session
pi --continue

# Single-shot mode (no session)
pi -p "What does this error mean?" < error.log

Why Should You Care?

If you already use Pi Agent, especially through OpenClaw, this project keeps the core workflow while upgrading the engine under the hood:

  • A native single-binary design intended to minimize startup and runtime overhead
  • A bounded-resource architecture for long-running sessions
  • A capability-gated security model for extension/tool execution, including command-level blocking of dangerous extension shell patterns

Security is a first-class design goal here, not a bolt-on:

  • Capability-gated hostcalls (tool/exec/http/session/ui/events)
  • Two-stage extension exec enforcement: capability gate first, then command mediation that blocks critical shell classes by default (for example recursive delete, disk/device writes, reverse shell) and can tighten to block high-tier classes in strict/safe policy
  • Policy + runtime risk + quota enforcement on the execution path
  • Per-extension trust lifecycle (pending -> acknowledged -> trusted -> killed) with kill-switch audit logs and explicit operator provenance
  • Hostcall-lane emergency controls that can force compatibility-lane execution globally or for one extension when fast-lane behavior needs immediate containment
  • Structured concurrency via asupersync for more predictable cancellation/lifecycle behavior
  • Auditable runtime signals/ledgers and redacted security alerts for extension behavior

TL;DR (Pi/OpenClaw Users)

The Rust port is designed around large-session, multi-agent, and extension-heavy workloads. Release-facing performance numbers are published only when the checked-in evidence artifacts are current, have matching run provenance, and show data and a passing result for every declared budget, with no data-contract failures. Historical benchmark snapshots are retained in planning/evidence artifacts, but they are not treated as current README claims until the performance evidence gate is regenerated cleanly.

Extension runtime guarantees are also concrete:

Extension assurance signal Why you should care
Two-stage exec guard (exec capability policy + command-level mediation + DCG/heredoc AST signals) Dangerous shell intent is caught before spawn, including destructive payloads hidden in multiline wrappers
Trust lifecycle + kill switch (pending/acknowledged/trusted/killed) You can quarantine an extension instantly, log who pulled the switch and why, and require explicit re-acknowledgement before restoring access
Hostcall lane kill-switch controls (forced_compat_global_kill_switch, forced_compat_extension_kill_switch) Fast-path regressions can be contained immediately by forcing compatibility-lane execution without disabling the extension system
Deterministic hostcall reactor mesh (shard affinity, bounded SPSC lanes, backpressure telemetry, optional NUMA slab tracking) Runtime behavior stays predictable under contention; queue pressure and routing decisions are observable instead of opaque
Cold owner-isolated JS realms + persistent transpile cache Every reload gets a fresh realm while versioned disk-cached transpilation avoids treating mutable JavaScript state as safely reusable
Tamper-evident runtime risk ledger (verify / replay / calibrate) Security decisions are hash-linked and can be replayed or threshold-tuned from real runtime traces

Bottom line: Pi's architecture targets lower latency, lower memory use, and stronger extension runtime safety under real workload pressure; current numeric claims must come from fresh, provenance-matched evidence artifacts.

Data source: docs/planning/BENCHMARK_COMPARISON_BETWEEN_RUST_VERSION_AND_ORIGINAL__GPT.md (latest secure-path + full orchestrator checkpoints, 2026-04-23).

README Citation Convention

Release-facing numeric performance claims in this README include inline citations with format: *(from [artifact-path], run [correlation-id])*

Example: *(from [artifact-path], run [correlation-id])*

Two additional machine-recognized citation forms exist:

  • *(from [artifact-path])* — path-only citation. The cited artifact must exist and parse; freshness is enforced by the same 14-day rule as release-facing claims.
  • *(from [artifact-path]; historical snapshot)* — explicit historical contract. The citation itself declares the obligation a retained snapshot: existence and validity are still checked, but staleness is not enforced because such claims never satisfy current release-facing requirements.

scripts/check_readme_evidence_freshness.py (a pre-release check listed in docs/releasing.md, not part of the code quality recipe) checks file freshness and artifact content so stale, no-data, or correlation-mismatched evidence cannot back user-facing performance claims; for release-facing citations of budget_summary.json it validates the full pi.perf.budget_summary.v2 contract, including that the header counts equal the per-budget rows. It reports line-numbered proof obligations for cited claims and extracts claim-gated performance phrases for reviewer audit. Historical snapshot citations are mapped separately and do not satisfy current release-facing claims.

Performance-Oriented Architecture

In this README, we means the project owner and collaborating coding agents.
The design concentrates performance work in several runtime layers rather than assuming one optimization proves an end-to-end result.

Technique What we do Runtime intent
Cold-start minimization Single native binary, no Node/Bun runtime bootstrap, no JIT warmup, startup prewarm for extension runtime paths Reduce time-to-first-interaction
Less copying on hot paths Arc/Cow message flow, zero-copy hostcall/tool payload handling, reduced clone-heavy provider/session paths Reduce CPU and allocation pressure
Deterministic dispatch core Typed hostcall opcodes, fast-lane/compat-lane routing, bounded shard queues with reactor-mesh telemetry Reduce tail latency under concurrent extension load
Efficient long-session storage SQLite session index + v2 sidecar (segmented log + offset index) with O(index+tail) reopen path Avoid full-history work on eligible resumes
Streaming parser tuned for real networks SSE parser tracks scanned bytes, handles UTF-8 tails, normalizes chunk boundaries, interns event-type strings Reduce repeated scanning and parser stalls
Safe fast-path controls Shadow dual execution sampling, automatic backoff on divergence/overhead, compatibility-lane kill switches for containment Bound optimization risk and preserve fallback behavior
DSR performance governance Scenario matrices, strict artifact contracts, fail-closed perf gates Detect regressions before release

If you want the full implementation inventory, see Performance Engineering.

Benchmark Methodology and Claim Integrity

The benchmark evidence policy is designed to keep results realistic, reproducible, and hard to game.

What we measured:

  • Matched-state workloads: resume a large session and append the same 10 messages.
  • Realistic E2E workloads: resume + append + extension activity + slash-style state changes + forks + exports + compactions.
  • Scale levels: from 100k up to 5M token-class session states.
  • Startup/readiness: command-level readiness (--help, --version) separately from long-session workflows.

How we kept comparisons fair:

  • Two scopes in the benchmark report:
    • apples-to-apples (pi_agent_rust vs legacy coding-agent)
    • apples-to-oranges (legacy stack components included where legacy behavior is outsourced)
  • Release-mode binaries and repeated runs per matrix cell.
  • No paid-provider noise in core latency/footprint tables (provider-call costs are excluded from these core comparisons).

How we kept claims honest:

  • Security controls stayed on during secure-path measurements (no policy/risk/quota bypasses for speed claims).
  • Raw artifacts are preserved (JSON/trace/time outputs) and called out in the benchmark report.
  • Blockers are explicitly disclosed: when direct legacy reruns were blocked by missing workspace deps, we state that and compare against prior validated legacy artifacts instead of pretending reruns succeeded.
  • Interpretation notes are explicit: the report distinguishes baseline sections vs fresh reruns so readers can see exactly which values came from which run set.
  • Reproducibility over marketing: methodology, caveats, and known limits are included alongside wins.

If you want full details, see:

  • docs/planning/BENCHMARK_COMPARISON_BETWEEN_RUST_VERSION_AND_ORIGINAL__GPT.md (methodology + results + caveats + raw artifact paths)

Why Pi?

Feature Pi (Rust) Typical TS/Python CLI
Startup Native single-binary path (Fresh v0.3.0 measurement pending; pre-v0.3.0 criterion: ~5-7ms p95 version, ~12-15ms help; see tests/perf/reports/budget_summary.json startup_version_p95 and startup_full_agent_p95) Runtime-dependent
Binary size Size-budgeted release profile (LTO, strip, opt-level = "z"); the binary_size_release budget and its latest stripped-artifact measurement are reported in Current Evidence State, never promoted here until claim readiness is ready Runtime-dependent
Memory (idle) Bounded-resource design; the idle_memory_rss budget and its latest release-binary measurement are reported in Current Evidence State (docs/perf-budgets-recipe.md defines the canonical 5-measurement taxonomy) Runtime-dependent
Streaming Native SSE parser Library-dependent
Tool execution Process tree management Basic subprocess
Sessions JSONL with branching Varies
Unsafe code Forbidden N/A

Quick Example

# 1) Start an interactive session
pi

# 2) Ask a codebase question
pi "Summarize the architecture in src/"

# 3) Attach a file inline
pi @src/main.rs "Explain startup flow"

# 4) Run single-shot mode for scripting
pi -p "List likely regression risks for this diff"

# 5) Continue your last project session
pi --continue

# 6) Inspect available models/providers
pi --list-models
pi --list-providers

Foundation Libraries

asupersync

asupersync is a structured concurrency async runtime designed for applications that need predictable resource cleanup. Key features used by pi_agent_rust:

  • Capability-based context (Cx): Async functions receive an explicit context that controls what they can do (HTTP, filesystem, time). This makes testing deterministic.
  • HTTP client with TLS: Built-in HTTP API with rustls, avoiding OpenSSL dependency hell
  • Structured cancellation: When a parent task cancels, all child tasks cancel cleanly. No orphaned futures.

pi_agent_rust runs on asupersync end-to-end today (runtime + HTTP/TLS + cancellation). Provider streaming uses a minimal HTTP client (src/http/client.rs) feeding a custom SSE parser (src/sse.rs).

rich_rust

rich_rust is a Rust port of Will McGugan's Rich Python library. It provides:

  • Markup syntax: [bold red]error[/] renders as bold red text
  • Tables: ASCII/Unicode table rendering with alignment and borders
  • Panels: Boxed content with titles
  • Progress bars: Animated progress indicators
  • Markdown: Terminal-rendered markdown with syntax highlighting
  • Themes: Consistent color schemes across components

The terminal UI uses rich_rust for all output formatting, providing the same visual quality as Rich-based Python tools.


Quick Start

1. Install

# Install latest release binary
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/pi_agent_rust/main/install.sh?$(date +%s)" | bash

If you already have the original TypeScript pi installed, the installer asks whether to make Rust Pi canonical as pi and automatically create legacy-pi for the old command.

2. Configure API Key

export ANTHROPIC_API_KEY="sk-ant-..."

3. Run

# Interactive mode
pi

# With an initial message
pi "Explain this codebase structure"

# Read files as context
pi @src/main.rs "What does this do?"

Using a local model

ollama, llamacpp (llama.cpp's llama-server), mistralrs (mistral.rs), and lmstudio are built-in local providers. The first three need no API key and work out of the box against their default ports:

pi --provider ollama    --model llama3        -p "hi"
pi --provider llamacpp  --model  -p "hi"
pi --provider mistralrs --model default        -p "hi"

To point at any other OpenAI-compatible server (a custom host/port, vLLM, etc.), add it to ~/.pi/agent/models.json:

{
  "providers": {
    "ollama": {
      "api": "openai-completions",
      "baseUrl": "http://127.0.0.1:11000/v1",
      "apiKey": "ollama",
      "models": [{ "id": "Qwen3.6-35B-A3B-4bit", "contextWindow": 32768 }]
    }
  }
}

Then pi --provider ollama --model Qwen3.6-35B-A3B-4bit. See docs/models.md for the full models.json schema, provider aliases, and secret resolution (env vars and !command shell lookups).

Refreshing provider model catalogs

For OpenAI-compatible providers, model discovery is a standalone command: it prints one model ID per stdout line and exits without starting the TUI.

# Use a successful live response or (with a warning on stderr) the static
# registry when live discovery is unavailable
pi --fetch-models openrouter

# Bypass the cache and require a genuinely live response; never fall back
pi --fetch-models openrouter --refresh-models

# Make a successful live/cache catalog available to future --list-models and
# interactive /model pickers
pi --fetch-models openrouter --refresh-models --persist-models

Each standalone CLI invocation starts a new process, so its in-memory cache is fresh. The five-minute cache only avoids repeat discovery calls made within one long-lived process by SDK/library users; --refresh-models bypasses that cache.

Persistence is opt-in. The v2 ~/.pi/agent/models.fetched.json schema stores provider/model IDs, the fetch timestamp, and a non-secret SHA-256 identity binding membership to the provider, API, query-free endpoint, auth-header mode, recognized credential-query ordered name/presence shape, and request-header name/presence shape that produced it. Pi never persists a static fallback, API credentials, URL query values, or request-header values—not even as offline-verifiable hashes. Credential-channel values are rotation-tolerant. Any non-empty query/header value outside a narrowly recognized credential channel may identify a tenant or deployment, so Pi refuses to persist that catalog (and ignores legacy persisted rows for such a current route) rather than reuse unverifiable membership. If a safely bound route shape no longer matches, Pi ignores those generated rows and reports how to refresh them. Credential rotation alone does not invalidate the catalog. The generated membership is likewise not account-bound: switching accounts on the same endpoint/transport shape can retain the prior account's saved model list until you rerun --fetch-models --refresh-models --persist-models. Inference still resolves and sends the current account's credential; only the opt-in model-membership list can be stale across that switch. The generated catalog is loaded first; your hand-written ~/.pi/agent/models.json is loaded afterward and remains authoritative. Pi does not rewrite or merge that user-authored file. Legacy pi.models.fetched.v1 files cannot be rebound safely because they lack this provenance, so Pi preserves them instead of overwriting them. Move the legacy file aside to models.fetched.v1.backup.json, then rerun the verified live refresh and persist command above to create a v2 catalog.


Features

Streaming Responses

Real-time response streaming with extended thinking support:

pi "Write a quicksort implementation"

Watch the response appear incrementally, with thinking blocks shown inline.

36 Built-in Tools

Tools are tiered so the model's live schema stays small while everything remains reachable. The tier table lives in src/xdev.rs; the default --tools list lives in src/cli.rs:

  • Essential (always in the provider schema): read, write, edit, bash, grep, find, ls, hashline_edit, ask, todo, web_search, submit_plan, current_time, xdev
  • Discoverable (registered, hidden from the schema until promoted via xdev list/describe/run/promote): ast_grep, ast_edit, lsp, debug, manage_skill — plus the memory-bank tools (retain, recall, reflect, memory_edit, learn) when memory.backend is local
  • Default-enabled: jobs (background bash job control) and hub (PTY service supervision), alongside the essential tier. The default --tools list names 19 tools; the registry always adds manage_skill and, when any discoverable tool is enabled, the xdev dispatcher
  • --tools opt-in extras: eval, github, security_scan
  • Settings-gated extras (off until enabled in settings.json): browser (browser.enableBrowser), computer (computer.enableComputer), the media trio inspect_image, generate_image, tts (media.enableInspectImage / media.enableGenerateImage / media.enableTts), and read_media (media.enableReadMedia; inline video/audio for Gemini-family models)
  • Opt-in only: subagent (it can start additional coding-agent processes)
Tool Description
read Read files (images and http(s)/skill://-style URLs supported)
write / edit / hashline_edit Create, surgically replace, or hashline-anchored edit files
bash Shell with timeout, command mediation, optional PTY, background jobs
grep / find / ls In-process content search, file discovery, listing
web_search Ranked multi-provider web search with circuit breaking
ask / todo Structured mid-turn option cards; persistent session task list
submit_plan Submit a completed plan for approval when plan mode is active
current_time The host's wall-clock time: UTC and local ISO-8601 timestamps, the UTC offset, Unix epoch seconds, weekday, and ISO week (the local zone follows TZ); takes no arguments. The system prompt carries only the date so the cached prefix stays stable; this tool is the clock
xdev Dispatcher exposing the discoverable tier (list/describe/run/promote)
ast_grep / ast_edit Structural code search and rewrite
lsp / debug Language-server (14 ops) and DAP debugging (29 ops) bridges
eval Persistent Python and JS kernels with cell semantics
jobs / hub Background bash job control; PTY service supervision
github gh-backed PR/issue/run operations
security_scan Rule-pack security scanning to SARIF (plan/run/disposition/compare)
manage_skill + memory tools Managed skills CRUD; opt-in project memory bank
browser Headless Chromium automation over CDP (navigate, snapshot, click, type, screenshot) with a domain allowlist; settings-gated
computer Desktop automation (displays, windows, screenshots, mouse/keyboard, clipboard); mutating actions require approval; settings-gated
inspect_image / generate_image / tts Vision analysis of local images, image generation/editing, and text-to-speech through provider adapters; settings-gated
read_media Attaches a local video/audio file (mp4, webm, mov, mp3, wav, m4a, ogg, flac) as an inline media block. Gemini, Gemini CLI, and Vertex Gemini models receive it natively as inline_data; every other provider sees [media omitted: , , ]. Hard cap 5 MiB per file (media.maxBytes); settings-gated
subagent Delegate isolated work to named Rust Pi child agents

All tools include automatic truncation for large outputs (2000 lines / 1MB), detailed metadata in responses, and process-tree cleanup for bash. Background-job logs live in a dedicated directory with a 256 MiB / 4096-entry budget. When it cannot admit another 16 MiB artifact, Pi deletes the oldest unlocked logs, only as many as needed, and always keeps active logs and at least eight recent ones. Set PI_JOBS_ARTIFACT_RETENTION=preserve to keep every log instead; Pi then refuses new background jobs once the budget is full. Job snapshots report the applied policy, removed-file count, and reclaimed bytes in artifactCleanup. Per-tool exposure is configurable via tools.loadMode. set to essential, discoverable, or off; an explicit --tools list always wins:

pi --tools read,bash,edit,write,grep,find,ls,hashline_edit,subagent \
  "Use the scout agent to inspect the provider implementation."

Native Subagents and Orchestration

Rust Pi includes a native subagent tool; it does not depend on a QuickJS extension and never resolves a child executable by assuming a pi binary on PATH. By default it starts the current Rust Pi executable. Set PI_SUBAGENT_PI_BINARY=/absolute/path/to/rpi only when an explicit binary override is needed.

Agent definitions are Markdown files in $PI_CODING_AGENT_DIR/agents/*.md (normally ~/.pi/agent/agents/*.md) or the nearest .pi/agents/*.md. Project definitions take precedence over same-named user definitions. The process inherits the parent's provider, router, authentication, and model-registry environment, including PI_CODING_AGENT_DIR.

---
name: scout
description: Inspect code and return concise, evidence-backed findings.
model: ai-router/gpt-5.6-sol
reasoning: low
tools: read,grep,find,ls
skills: ../skills/codebase-archaeology/SKILL.md
---

Read relevant code before answering. Do not edit files.

The tool accepts exactly one workflow shape: a single { agent, task }, a bounded parallel tasks array (up to 8, default concurrency 4), or a sequential chain array. In chains, {previous} in a task is replaced by the preceding child's final output. Children run headlessly in isolated ephemeral sessions, stream progress into the parent TUI/JSON output, return structured status and stderr details, and are killed/reaped if the parent is cancelled. Child agents receive the declared tool allowlist; the default child allowlist deliberately excludes subagent to prevent accidental recursive delegation.

In interactive mode, /tan starts the same task-role child machinery without interrupting the main conversation. The child runs in the current directory, appears in hub agent roster with kind=tan, and sends its bounded completion summary through the follow-up queue at the next idle turn boundary. Because /tan inherits the delegation safety gate, it is available only when the opt-in subagent tool is enabled.

Session Management

Sessions persist as JSONL files with full conversation history:

# Continue most recent session
pi --continue

# Open specific session
pi --session ~/.pi/agent/sessions/--home-user-project--/2024-01-15T10-30-00.jsonl

# Ephemeral (no persistence)
pi --no-session

Sessions support:

  • Tree structure for conversation branching
  • Model/thinking level change tracking
  • Automatic compaction for long conversations

Extended Thinking

Enable deep reasoning for complex problems:

pi --thinking high "Design a distributed rate limiter"

Thinking levels: off, minimal, low, medium, high, xhigh, max

Customization (Skills & Prompt Templates)

  • Skills: Drop SKILL.md under ~/.pi/agent/skills/ or .pi/skills/ and invoke with /skill:name.
  • Prompt templates: Markdown files under ~/.pi/agent/prompts/ or .pi/prompts/; invoke via / [args].
  • Packages: Share bundles with pi install npm:@org/pi-packages (skills, prompts, themes, extensions).

Autocomplete

Pi provides context-aware autocomplete in the interactive editor:

  • @ file references: Type @ followed by a path fragment to attach file contents. The completion engine indexes project files (respecting .gitignore) via the ignore crate's WalkBuilder, capping at 5,000 entries.
  • / slash commands: 40+ built-in commands (see the slash-command table below) and user-defined prompt templates and skills all appear as completions.
  • Fuzzy scoring: Prefix matches rank above substring matches. Results are sorted by match quality, then by kind (commands > templates > skills > files > paths).
  • Background refresh: A background thread re-indexes the project file tree every 30 seconds, so completions stay current without blocking the input loop.

Four Execution Modes

Pi runs in four modes, each suited to different workflows:

Mode Invocation Use Case
Interactive pi (default) Full TUI with streaming, tools, session branching, autocomplete
Print pi -p "..." Single response to stdout, no TUI, scriptable
RPC pi --mode rpc Headless JSON protocol over stdin/stdout for IDE integrations
ACP pi --acp JSON-RPC 2.0 Agent Client Protocol over stdin/stdout (e.g. the Zed editor)

Interactive mode provides the full experience: a multi-line text editor with history, scrollable conversation viewport, model selector (Ctrl+L), scoped model cycling (Ctrl+P/Ctrl+Shift+P), session branch navigator (/tree), and real-time token/cost tracking. Since v0.4.0 the default interactive stack is the FrankenTUI (ftui) runtime; pi --inline keeps your shell scrollback by drawing the UI at the bottom of the screen instead of on the alternate screen, and pi --classic (aliases --classic-tui, --charmed, --bubbletea) selects the previous charmed_rust stack until it is removed.

Print mode sends one message, streams the response to stdout, and exits. Useful for shell scripts and one-off queries.

Print mode and tool approval. The approval mode defaults to always-ask on every surface, including -p. The absence of a terminal is not treated as consent, so Pi does not quietly auto-approve a scripted run. Print mode also has no way to prompt, which means a gated tool call in a default -p run is denied. Pass --approval-mode yolo (or --yolo) to auto-approve tool calls, or set approval.mode in settings.json:

echo "list the files here" | pi -p --mode json --yolo

When a run does end with tool calls denied for want of an approval surface, Pi exits 3 and explains why on stderr, rather than exiting 0 on a turn that silently did nothing. In --mode json and --mode rpc the same failure also arrives as the single machine-readable record described under --mode, with "code": "approval.surface_unavailable". Exit 3 means specifically "the model could not use tools"; exit 1 remains an ordinary failure and exit 2 a usage error.

This is a deliberate divergence from the Node Pi CLI, which auto-approves in print mode. Legacy Pi is historical context here, not a compatibility authority, and inheriting its default would silently grant bash and write access to any script that never opted in.

RPC mode exposes a line-delimited JSON protocol for programmatic control. Clients send commands (prompt, steer, follow-up, abort, get-state, compact) and receive streaming events. This is how IDE extensions and custom frontends integrate with Pi. See RPC Protocol for the wire format.

Extensions

Pi supports two extension runtime families with capability-gated host connectors:

  • JS/TS entrypoints run without Node or Bun in an embedded QuickJS runtime.

  • *.native.json descriptors run in the native-rust descriptor runtime.

  • Extension entrypoints are auto-detected:

    • .js/.ts/.mjs/.cjs/.tsx/.mts/.cts run directly in embedded QuickJS (no descriptor conversion).
    • *.native.json loads the native-rust descriptor runtime.
    • One session currently uses one runtime family at a time (JS/TS or native descriptor).
  • Cold/warm extension load paths are instrumented; fresh v0.3.0 measurements are pending

  • Node API shims for fs, path, os, crypto, child_process, url, and more

  • Capability-based security: extensions call explicit connectors (tool/exec/http/session/ui) with audit logging

  • Command-level exec mediation: dangerous shell signatures are classified and blocked before spawn, with redacted denial alerts and mediation ledger entries

  • Trust-state lifecycle and kill-switch controls with audited state transitions (pending/acknowledged/trusted/killed)

  • Workspace trust-on-first-use gate: project-local .pi/settings.json packages and .pi/extensions/ require a one-time interactive approval (keyed to workspace path + content digest; --trust, PI_WORKSPACE_TRUST, or global trustAllWorkspaces for automation; non-interactive launches fail closed)

  • Hostcall reactor mesh with deterministic shard routing, bounded queue backpressure, and optional NUMA-aware telemetry

  • Fresh owner-isolated realms on every reload, with versioned persistent transpile caching instead of mutable realm reuse

Credential-Aware Model Selection

  • /model (or Ctrl+L) opens a selector focused on models that are ready to run with current credentials.
  • Ctrl+P and Ctrl+Shift+P cycle through the scoped model set without opening the overlay.
  • Provider IDs and aliases are matched case-insensitively in model selection and /login.
  • Models that do not require configured credentials can run keyless.

Extensions can register tools, slash commands, event hooks, flags, providers, and shortcuts. See EXTENSIONS.md for the full architecture and docs/extension-catalog.json for the 223-entry catalog with per-extension conformance status and perf budgets.

Extension Validation Pipeline

This project validates extension compatibility with a three-track pipeline:

  • Vendored corpus (223, plus one intentionally excluded negative test fixture): deterministic conformance, compatibility matrix, and scenario suites.
  • Unvendored corpus (777): source acquisition and onboarding prioritization.
  • Release-binary live-provider E2E: real target/release/pi execution against a non-mocked provider/model path.

Why this exists

  • Catch runtime/API regressions in QuickJS host shims and capability policy.
  • Catch dangerous extension shell call patterns with real command mediation on the release binary path.
  • Verify extension behavior against real provider responses, not just fixture/mocked flows.
  • Keep extension support measurable instead of anecdotal.
  • Produce a prioritized queue for onboarding unvendored candidates into vendored conformance.

Pipeline components

  1. Fetch unvendored source corpus

    • Binary: ext_unvendored_fetch_run
    • Execution authority: the registered DSR extension-validation lane; do not invoke the example directly with Cargo
    • Purpose:
      • Clones GitHub repos and unpacks npm tarballs into .tmp-codex-unvendored-cache/
      • Produces machine-readable acquisition status for all unvendored candidates
    • Artifacts:
      • tests/ext_conformance/reports/pipeline/unvendored_fetch_probe_report.json
      • tests/ext_conformance/reports/pipeline/unvendored_fetch_probe_events.jsonl
  2. Run end-to-end validation orchestration

    • Binary: ext_full_validation
    • Execution authority: the registered DSR extension-validation lane
    • Stages (in order):
      1. refresh_onboarding_queue (runs ext_onboarding_queue)
      2. conformance_shard_0..N (runs ext_conformance_generated sharded matrix)
      3. conformance_failure_dossiers
      4. provider_compat_matrix
      5. scenario_conformance_suite
      6. auto_repair_full_corpus
      7. differential_suite (optional, enabled via --run-diff; npm diff via --run-npm-diff)
    • Artifacts:
      • tests/ext_conformance/reports/pipeline/full_validation_report.json
      • tests/ext_conformance/reports/pipeline/full_validation_report.md
      • Plus stage-specific reports under tests/ext_conformance/reports/**
  3. Run dev-firstset live-provider gate (must pass before release build)

    • Binary: ext_release_binary_e2e
    • Execution authority: DSR quality with its configured live-provider inputs
    • Purpose:
      • Proves the current codepath works end-to-end on a representative first-set before paying release-build cost.
      • Serves as the promotion gate to full release-binary validation.
    • Gate:
      • Require pass=20 / total=20 with fail=0.
    • Artifacts:
      • tests/ext_conformance/reports/release_binary_e2e/ollama_firstset_dev_20260219_jobs10_timeout600.json
      • tests/ext_conformance/reports/release_binary_e2e/ollama_firstset_dev_20260219_jobs10_timeout600.md
  4. Run full release-binary live-provider E2E (after step 3 passes)

    • Binary: ext_release_binary_e2e
    • Execution authority: the DSR release-validation lane after DSR quality passes
    • Purpose:
      • Executes target/release/pi directly for each selected extension case.
      • Uses a live provider/model path (default ollama + qwen2.5:0.5b) to exercise non-mocked end-to-end behavior.
      • Emits per-case stdout/stderr captures plus summary artifacts (pi.ext.release_binary_e2e.v1).
    • Artifacts:
      • tests/ext_conformance/reports/release_binary_e2e/ollama_full_release_20260219_jobs10_timeout600.json
      • tests/ext_conformance/reports/release_binary_e2e/ollama_full_release_20260219_jobs10_timeout600.md
      • tests/ext_conformance/reports/release_binary_e2e/cases/*
  5. Aggregate and triage

    • full_validation_report.json combines:
      • Stage-level pass/fail (stageSummary, stageResults)
      • Corpus counts (corpus)
      • Vendored conformance totals (conformance)
      • Provider matrix totals (providerCompat)
      • Scenario totals (scenario)
      • Review queue + verdict classification (reviewQueue, verdictCounts)
    • Important interpretation rule:
      • not_tested_unvendored indicates unvendored candidates not yet in vendored conformance; this is inventory status, not a vendored regression.

Recommended run environment

These runs compile many crates and can be disk-heavy. DSR owns target/temp placement, load admission, and any remote offload. Run the registered recipe:

dsr quality --tool pi_agent_rust

Historical run snapshot (extension gate refresh 2026-05-15)

These artifacts predate v0.3.0 and do not certify the current source revision. The strict drop-in program has been retired; its contract and verdict are no longer release gates and will not be rerun to define product completeness. The figures below are retained only as a dated engineering snapshot.

The current extension must-pass corpus is the exact set selected by docs/extension-inclusion-list.json; the historical counts below do not prove coverage of that authoritative list.

From:

  • tests/ext_conformance/reports/gate/must_pass_gate_verdict.json (generated 2026-05-15T17:03:02.000Z, run local-20260515T170218075Z)

  • tests/ext_conformance/reports/health_delta/health_delta_report.json (generated 2026-05-13T03:37:59.568Z)

  • tests/ext_conformance/reports/journeys/journey_report.json (generated 2026-05-13T02:59:58.302Z)

  • tests/evidence_bundle/index.json (generated 2026-05-12T19:26:21.441Z, run local-20260512T192621Z)

  • tests/full_suite_gate/certification_verdict.json (generated 2026-05-14T19:59:37.227Z)

  • Historical verdict blob at 2fc4b8c0b77ded267cf5e0f517f4b6fa87f45e91:docs/evidence/dropin-certification-verdict.json (generated 2026-05-18T19:37:26Z for source 52e9fbfb24352045985b59df9d7ea63f1f8f2ef8; the live file may contain a later verdict)

  • Historical strict drop-in result: 22/22 certification gates PASS, 16/16 blocking gates PASS - CERTIFIED for source 52e9fbfb24352045985b59df9d7ea63f1f8f2ef8 only (from the Git-pinned historical verdict blob above; it neither describes the live verdict file nor certifies v0.3.0)

  • Unified evidence bundle, as regenerated 2026-08-28: 30 total sections, 18 present, 12 missing, 0 invalid (bundle verdict: partial) (from tests/evidence_bundle/index.json; historical snapshot)

  • Extension must-pass gate, as regenerated 2026-08-17: 206/208 must-pass extensions passed (2 failures); informational stretch set 10/19 passed — the May 2026 123/123 snapshot predates the expanded corpus (from tests/ext_conformance/reports/gate/must_pass_gate_verdict.json; historical snapshot)

  • Extension health delta, as regenerated 2026-08-17: 226 extensions tested at 95.6% pass rate with 0 regressions vs the 2026-02-07 baseline (from tests/ext_conformance/reports/health_delta/health_delta_report.json; historical snapshot)

  • Extension journey coverage, as regenerated 2026-08-17: 125/125 journey scenarios passed (100.0%); command, event-subscriber, multi-capability, passive, and tool-provider categories are green (from tests/ext_conformance/reports/journeys/journey_report.json; historical snapshot)

  • Historical stress-triage evidence is retained under tests/perf/reports/; it is not current enough to support a v0.3.0 performance claim.


Installation

Curl Installer (Recommended)

# Latest release
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/pi_agent_rust/main/install.sh?$(date +%s)" | bash

# Non-interactive + auto PATH update
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/pi_agent_rust/main/install.sh?$(date +%s)" | bash -s -- --yes --easy-mode

# Pin a release tag
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/pi_agent_rust/v0.3.0/install.sh" | bash -s -- --version v0.3.0

# Install from explicit artifact URL + checksum URL
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/pi_agent_rust/v0.3.0/install.sh" | \
  bash -s -- \
    --artifact-url "https://github.com/Dicklesworthstone/pi_agent_rust/releases/download/v0.3.0/pi-linux-amd64.tar.xz" \
    --checksum-url "https://github.com/Dicklesworthstone/pi_agent_rust/releases/download/v0.3.0/pi-linux-amd64.tar.xz.sha256"

# Skip completion setup (CI/non-interactive minimal install)
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/pi_agent_rust/main/install.sh?$(date +%s)" | \
  bash -s -- --yes --no-completions

The installer is idempotent and supports a migration path from TypeScript Pi:

  • Detect existing TS pi command
  • Prompt to install Rust Pi as canonical pi
  • Preserve old CLI behind legacy-pi
  • Record state for clean uninstall/restore

Notable installer flags:

  • --offline [TARBALL]: enforce offline mode; optional local artifact path (.tar.gz, .tar.xz, .zip, or raw binary)
  • --artifact-url: force a specific release artifact URL
  • --checksum / --checksum-url: override checksum source for explicit artifacts
  • --sigstore-bundle-url: override the optional Sigstore bundle URL used by cosign verify-blob; set both COSIGN_IDENTITY_RE and COSIGN_OIDC_ISSUER to the trusted signer policy (there is deliberately no GitHub Actions identity default)
  • --completions auto|off|bash|zsh|fish: force shell completion install target (off is equivalent to --no-completions)
  • --no-completions: disable completion installation
  • --no-agent-skills: skip automatic installation of the pi-agent-rust skill into ~/.claude/skills/ and ~/.codex/skills/
  • --no-verify: skip checksum + signature verification (testing only)
  • --artifact-url without --version uses a synthetic tag for release mode only; if artifact download fails, install exits instead of attempting source fallback
  • Installer honors HTTPS_PROXY / HTTP_PROXY for all network fetches

By default, the installer also installs a pi-agent-rust skill for both Claude Code and Codex CLI:

  • Claude Code: ~/.claude/skills/pi-agent-rust/SKILL.md
  • Codex CLI: ~/.codex/skills/pi-agent-rust/SKILL.md (or $CODEX_HOME/skills/pi-agent-rust/SKILL.md if CODEX_HOME is set)
  • During upgrades, installer-managed legacy pre-tool entries from older versions are removed automatically (idempotent, path-scoped, and non-destructive) when prior installer state is present.

Installer regression harness (options + checksum + signature + completions):

bash tests/installer_regression.sh

Distribution Compatibility Contract (Packaging/Invocation Scope)

For migration adoption, packaging and invocation compatibility follows this contract:

  • This section covers packaging and invocation behavior only. Historical drop-in/parity artifacts do not define the current product or release gate. Current compatibility statements must name the concrete behavior they cover.

  • Canonical executable name is pi across release assets and installer-managed installs.

  • Installer-managed installs also create an rpi compatibility launcher when no conflicting rpi command already exists on your PATH.

  • Existing TypeScript pi installs can be migrated in place; the prior command is preserved as legacy-pi.

  • If you keep TypeScript pi as canonical (--keep-existing-pi), Rust Pi is installed as pi-rust.

  • On Apple Silicon, the installer prefers the native arm64 artifact even when launched from a Rosetta-translated shell.

  • Version-pinned installs are supported via install.sh --version vX.Y.Z for deterministic rollouts.

  • DSR releases ship each platform archive with a same-name .sha256 sidecar for integrity validation. The v0.3.0 release predates that contract and ships an aggregate SHA256SUMS instead; the installer verifies the per-asset sidecar first and falls back to SHA256SUMS only when the sidecar is absent.

Representative smoke checks:

# Canonical command should exist and execute
command -v pi
pi --version
pi --help >/dev/null

# If a TS migration was performed, legacy command remains available
command -v legacy-pi && legacy-pi --version

Source builds

Repository builds are tested with the exact toolchain pinned in rust-toolchain.toml (nightly-2026-08-31). The locked dependency graph requires Rust 1.95 or newer. Project builds are DSR-only:

dsr quality --tool pi_agent_rust
dsr build pi_agent_rust

End users should install a DSR-published archive through install.sh. There is no supported direct-Cargo contributor or installation lane.

Dependencies

Pi has no required external runtime dependencies: the grep and find tools search in-process using the same engines ripgrep is built from (grep-searcher/grep-regex and the ignore walker), so patterns keep ripgrep's default (Rust regex crate) syntax. PCRE2-only features (lookaround, backreferences) were never exposed and remain unsupported.

To shell out to rg/fd instead (debugging escape hatch), set "search_backend": "external" in settings.json and install them via apt install ripgrep fd-find or brew install ripgrep fd.

Uninstall

curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/pi_agent_rust/main/uninstall.sh" | bash

By default, uninstall removes installer-managed Rust binaries/aliases and skill directories, then restores a migrated TypeScript pi if one was preserved.


Commands

Basic Usage

pi [OPTIONS] [MESSAGE]...

# Examples
pi                              # Start interactive session
pi "Hello"                      # Start with message
pi @file.rs "Explain this"      # Include file as context
pi -p "Quick question"          # Print mode (no session)

Interactive file references:

  • Type @relative/path in the editor to attach a file’s contents (autocomplete inserts the @ form).

Options

Option Description
-c, --continue Continue most recent session
-r, --resume Open session picker UI
--session Open specific session file
--session-dir Override session storage directory for this run
`--session-durability strict balanced
--no-session Don't persist conversation
-p, --print Single response, no interaction
`--mode text json
--provider Force provider for this run (aliases supported)
--model Model to use (auto-select fallback: anthropic/claude-sonnet-4-6, then anthropic/claude-opus-4-7, then openai/gpt-5.1-codex)
--thinking Thinking level: off/minimal/low/medium/high/xhigh/max
--tools Comma-separated tool list
--api-key API key (or use provider-specific env vars such as ANTHROPIC_API_KEY, OPENAI_API_KEY, etc.)
`--extension-policy safe balanced
`--repair-policy off suggest
--list-models [PATTERN] List available models (optional fuzzy filter)
--list-providers List canonical provider IDs, aliases, and auth env keys
--fetch-models Print provider model IDs to stdout and exit; warns when using the static fallback
--refresh-models With --fetch-models, bypass the same-process cache and require a successful live response
--persist-models With --fetch-models, atomically save a verified live/same-process-cache catalog for future model pickers
--export Export session file to HTML

Additional high-leverage flags:

  • --no-migrations to skip startup migration checks
  • --explain-extension-policy to print effective capability decisions and exit
  • --explain-repair-policy to print effective repair-policy resolution and exit

Subcommands

# Package management
pi install  [-l|--local]    # Install a package source and add to settings
pi remove  [-l|--local]     # Remove a package source from settings
pi update [source]                 # Update all (or one) non-pinned packages
pi list                            # List user + project packages from settings

# Configuration
pi config                          # Show settings paths + precedence

More utility subcommands:

# Extension catalog index + discovery
pi update-index
pi search "git"
pi info pi-search-agent

# Environment and extension diagnostics
pi doctor
pi doctor --only sessions --format json
pi doctor --only swarm --format json
pi doctor ./path/to/extension --policy safe --fix

# Read-only swarm progress SLO evaluation from normalized evidence
pi swarm-progress --input progress-slo-input.json --format json
pi swarm-progress --input progress-slo-input.json --since HEAD~1 --out-json progress-slo.json

# Session storage migration (JSONL -> v2 sidecar store)
pi migrate ~/.pi/agent/sessions --dry-run
pi migrate ~/.pi/agent/sessions
  • update-index refreshes extension index metadata used by search and info.
  • search and info let you discover and inspect extension metadata without leaving the CLI.
  • doctor checks config, directories, auth, shell setup, sessions, swarm coordination readiness, and extension compatibility. pi doctor --only swarm --format json also reports cgroup CPU quota, cpuset size, NUMA topology, cgroup memory limits, target/tmp headroom, and recommended concurrency budgets before large multi-agent runs.
  • swarm-progress evaluates a normalized progress SLO snapshot and emits advisory JSON/text only; it does not mutate Beads, git, Agent Mail, RCH, validation broker slots, runpacks, or source files. Operator workflow, privacy boundaries, degraded Agent Mail/RCH interpretation, stale-Beads handling, and no-open-work convergence guidance lives in docs/swarm-operations-runbook.md#progress-slo-operator-workflow.
  • migrate validates or creates the v2 session sidecar format for faster resume on larger histories.

Configuration

Pi reads configuration from ~/.pi/agent/settings.json. Keys are shown in their canonical snake_case form below, but every field also accepts the original TypeScript Pi's camelCase spelling (e.g. defaultProvider, shellCommandPrefix) as a serde alias, so an existing pi-mono settings.json is parsed as-is and is never rewritten to snake_case:

{
  "default_provider": "anthropic",
  "default_model": "claude-opus-4-5",
  "default_thinking_level": "medium",

  "compaction": {
    "enabled": true,
    "reserve_tokens": 8192,
    "keep_recent_tokens": 20000
  },

  "retry": {
    "enabled": true,
    "max_retries": 3,
    "base_delay_ms": 1000,
    "max_delay_ms": 30000
  },

  "images": {
    "auto_resize": true,
    "block_images": false
  },

  "terminal": {
    "show_images": true,
    "clear_on_shrink": false
  },

  "http": {
    "proxy": "http://127.0.0.1:2080",
    "no_proxy": ["localhost", "127.0.0.1", ".internal.example"]
  },

  "shell_path": "/bin/bash",
  "shell_command_prefix": "set -e"
}

Network Proxy

Every request Pi makes — provider APIs, OAuth logins, update checks, URL reads, package fetches — goes through the proxy resolved here. https:// targets use a CONNECT tunnel, so TLS stays end-to-end to the origin and the proxy never sees plaintext; http:// targets use an absolute-form request line. Credentials in the proxy URL (http://user:pass@host:port) become Proxy-Authorization, and are never logged or passed to child processes.

Resolution order for a request, first match wins:

  1. A bypass match — http.no_proxy in settings.json, else NO_PROXY / no_proxy. Entries follow the usual convention: * bypasses everything, example.com or .example.com matches the domain and its subdomains, and example.com:8443 restricts the match to that port.
  2. http.https_proxy / http.http_proxy (scheme-specific settings).
  3. http.proxy (both schemes).
  4. PI_HTTPS_PROXY / PI_HTTP_PROXY.
  5. The standard HTTPS_PROXY, HTTP_PROXY, ALL_PROXY (lowercase spellings accepted).

Ambient proxy variables are honored by default, the same as git and curl. Where they are set for some other tool — a capture proxy, a stale VPN helper — turn the inheritance off with "http": { "ignore_env_proxy": true } or PI_HTTP_PROXY=off; explicit settings and PI_*_PROXY still apply. Proxy endpoints may be http://, socks5:// (local DNS) or socks5h:// (DNS on the proxy), with optional username/password authentication; SOCKS credentials are not encrypted on the hop to the proxy. There is no automatic loopback bypass: requests to local model servers (Ollama, LM Studio) go through the proxy too unless localhost and 127.0.0.1 are in NO_PROXY or http.noProxy. An unusable ambient value (e.g. an https:// or socks4:// endpoint) is skipped with a warning rather than failing requests.

The resolved proxy is also injected into every process the bash tool spawns (as HTTP_PROXY/HTTPS_PROXY/NO_PROXY and their lowercase spellings), so git, curl, and npm invocations reach the network the same way Pi does. The injection is a per-child copy: Pi's own environment and the system environment are never modified. Proxy credentials are deliberately left out of that copy, so an authenticated proxy needs its own configuration for those tools (e.g. git config --global http.proxy).

Configuration Precedence

Settings are resolved in priority order (first match wins):

  1. CLI flags (--model, --thinking, --provider, etc.)
  2. Environment variables (ANTHROPIC_API_KEY, PI_CONFIG_PATH, etc.)
  3. Project settings (.pi/settings.json in the working directory)
  4. Global settings (~/.pi/agent/settings.json)
  5. Built-in defaults

This means a CLI flag always overrides a settings.json value, and a project-level setting overrides the global one.

Resource Resolution

Skills, prompt templates, themes, and extensions follow the same resolution order:

  1. CLI-specified paths (--skill, --prompt-template, --theme, -e)
  2. Project directory (.pi/skills/, .pi/prompts/, .pi/themes/, .pi/extensions/)
  3. Global directory (~/.pi/agent/skills/, ~/.pi/agent/prompts/, etc.)
  4. Installed packages (~/.pi/agent/packages/)

When multiple resources share the same name, the first occurrence wins. Collisions are logged as diagnostics.

--no-skills (and its siblings --no-prompt-templates, --no-themes, --no-extensions) disables tiers 2–4 including skills entries listed in settings.json — upstream-pi parity. Explicit CLI paths (tier 1) still load, so pi --no-skills --skill /path/to/skill-a --skill /path/to/skill-b is the way to run with an exact, isolated skill set (e.g. per-profile setups via shell aliases or wrapper scripts).

Project context files are a separate switch. By default pi appends AGENTS.md / CLAUDE.md from ~/.pi/agent/, the working directory, and every ancestor directory to the system prompt (plus imported foreign-format workspace rules). --no-context-files (or PI_NO_CONTEXT_FILES=1) disables that discovery entirely; --no-skills does not cover it. Hosts that compose the whole prompt with --system-prompt typically pass --no-context-files --no-skills together.

Prompt template expansion supports positional arguments: $1, $2, $@ (all args), and slice syntax ${@:start}, ${@:start:length}. For example, a template invoked as /review src/main.rs --strict receives src/main.rs as $1 and --strict as $2.

Environment Variables

Variable Description
ANTHROPIC_API_KEY Anthropic API key
OPENAI_API_KEY OpenAI API key
GOOGLE_API_KEY Google Gemini API key
AZURE_OPENAI_API_KEY Azure OpenAI API key
COHERE_API_KEY Cohere API key
GROQ_API_KEY Groq API key (OpenAI-compatible)
DEEPINFRA_API_KEY DeepInfra API key (OpenAI-compatible)
CEREBRAS_API_KEY Cerebras API key (OpenAI-compatible)
OPENROUTER_API_KEY OpenRouter API key (OpenAI-compatible)
ORCAROUTER_API_KEY OrcaRouter API key (OpenAI-compatible)
MISTRAL_API_KEY Mistral API key (OpenAI-compatible)
MOONSHOT_API_KEY Moonshot/Kimi API key (OpenAI-compatible)
DASHSCOPE_API_KEY DashScope/Qwen API key (OpenAI-compatible)
DEEPSEEK_API_KEY DeepSeek API key (OpenAI-compatible)
FIREWORKS_API_KEY Fireworks API key (OpenAI-compatible)
TOGETHER_API_KEY Together API key (OpenAI-compatible)
PERPLEXITY_API_KEY Perplexity API key (OpenAI-compatible)
XAI_API_KEY xAI API key (OpenAI-compatible)
PI_CONFIG_PATH Custom config file path
PI_CODING_AGENT_DIR Override the global config directory
PI_SUBAGENT_PI_BINARY Explicit Rust Pi executable for native child agents; defaults to the current executable
PI_JOBS_ARTIFACT_RETENTION Background-job artifact policy: rotate (default, deletes the oldest unlocked logs when the budget is full) or preserve (keeps every log, refuses new jobs when full)
PI_PACKAGE_DIR Override the packages directory
PI_SESSIONS_DIR Custom sessions directory

Architecture

┌─────────────────────────────────────────────────────────────────┐
│                           CLI (clap)                            │
│  • Argument parsing    • @file expansion    • Subcommands       │
└─────────────────────────────────┬───────────────────────────────┘
                                  │
┌─────────────────────────────────▼───────────────────────────────┐
│                          Agent Loop                             │
│  • Message history     • Tool iteration    • Event callbacks    │
└────────┬──────────────────────┬──────────────────────┬──────────┘
         │                      │                      │
┌────────▼────────┐  ┌─────────▼──────────┐  ┌───────▼──────────┐
│ Provider Layer  │  │  Tool Registry     │  │  Extension Mgr     │
│ • Anthropic     │  │  • read  • bash    │  │  • QuickJS JS/TS   │
│ • OpenAI (Chat/ │  │  • write • grep    │  │  • Native descriptor│
│   Responses)    │  │  • edit  • find    │  │    runtime          │
│ • Gemini/Cohere │  │  • ls              │  │  • Capability policy│
│ • Azure/Bedrock │  │  • ext-registered  │  │  • Node shims       │
│ • Vertex/Copilot│  │                    │  │  • Event hooks      │
│ • GitLab/Cursor │  │                    │  │  • Runtime risk ctl │
│   /Ext          │  │                    │  │                    │
└────────┬────────┘  └─────────┬──────────┘  └───────┬──────────┘
         │                     │                      │
┌────────▼─────────────────────▼──────────────────────▼──────────┐
│                     Session Persistence                         │
│  • JSONL format (v3)   • Tree structure   • Session index/cache  │
│  • Per-project dirs    • Default-enabled SQLite backend support │
└─────────────────────────────────────────────────────────────────┘

Provider-count rule: Pi has 11 native provider implementation modules. Those modules are anthropic, openai, openai_responses, gemini, cohere, azure, bedrock, vertex, copilot, gitlab, and cursor. User-visible provider IDs, aliases, OpenAI-compatible presets, and extension-provided streamSimple providers are counted separately because several native modules expose multiple routes. The factory module mod.rs and shared live-catalog helper model_fetch.rs are support modules, not provider implementations, and are excluded.

Key Design Decisions

  1. No unsafe code: #![forbid(unsafe_code)] enforced project-wide
  2. Streaming-first: Custom SSE parser, no blocking on responses
  3. Process tree management: sysinfo crate ensures no orphaned processes
  4. Structured errors: thiserror with specific error types per component
  5. Size-budgeted release profile: LTO + strip + opt-level = "z" for budget-compliant shipping artifacts

asupersync Context vs TypeScript Pi (pi-mono)

This Rust port preserves Pi's user experience, but intentionally changes the runtime substrate. The original TypeScript Pi (pi-mono, packages/coding-agent) is built on Node.js + package-level abstractions. pi_agent_rust moves those same behaviors onto asupersync primitives so lifecycle guarantees are explicit in the runtime model.

Concern TypeScript Pi (pi-mono baseline) pi_agent_rust + asupersync
Runtime model Node event loop + Promise/AbortSignal conventions RuntimeBuilder + explicit reactor and runtime handle
Async ownership Task lifetimes coordinated by framework/library code Structured task ownership and explicit cross-thread channels (TUI/RPC bridging)
Cancellation semantics Primarily API- and tool-layer conventions Runtime-aware cancellation checks + bounded timeout handling in tools
I/O capability shape Ambient Node APIs + extension layer policies Capability-scoped context (AgentCx over asupersync::Cx) and explicit hostcall policy
HTTP streaming Provider/client dependent Purpose-built asupersync HTTP/TLS client feeding custom SSE parser
Deterministic test hooks Conventional async test setup asupersync test/runtime hooks used widely in unit/integration tests

Why this is useful in practice:

  • More predictable failure behavior during aborts/timeouts because cancellation is checked in explicit loop boundaries and tool runners.
  • Cleaner resource lifetimes because the runtime, timers, and I/O paths all share one concurrency substrate.
  • Less hidden coupling because the main invariants live in Rust types/algorithms rather than spread across framework conventions.

Runtime Invariants (and Why They Matter)

These are the concrete invariants we rely on in this implementation:

  1. Turn-scoped agent lifecycle

    • The main loop emits AgentStart, TurnStart, TurnEnd, and AgentEnd in a stable order.
    • Tool recursion is bounded by max_tool_iterations (default 50) to avoid unbounded self-tool loops.
    • Benefit: stable event ordering for TUI/RPC consumers and predictable termination behavior.
  2. Abort and timeout behavior is explicit

    • Agent abort checks happen at turn boundaries and around tool execution.
    • bash timeout follows a clear escalation path: terminate process tree, grace period, then hard kill.
    • Benefit: fewer "hung" sessions and reduced orphan-process risk during aggressive tool use.
  3. Session writes are crash-resilient

    • JSONL saves write to a temp file and persist atomically.
    • Session indexing uses SQLite WAL + lock file coordination for concurrent instances.
    • Benefit: better durability and resume reliability under multi-process usage.
  4. Compaction is threshold-driven and boundary-aware

    • Trigger: estimated context tokens exceed context_window - reserve_tokens.
    • Cut-point logic prefers user-turn boundaries and preserves recent context budget.
    • Benefit: compaction recovers context without collapsing near-term task continuity.
  5. Capability policy is fail-closed and precedence-defined

    • Resolution order: per-extension deny -> global deny -> per-extension allow -> default caps -> mode fallback.
    • Benefit: policy outcomes are explainable, deterministic, and auditable.
  6. Streaming parser tolerates real network chunking

    • SSE parser handles CR/LF variants, multi-line data: fields, partial UTF-8 tails, and end-of-stream flush.
    • Benefit: incremental rendering remains robust across providers and network fragmentation.

Design Principles Carried From asupersync Into Pi

The following asupersync principles are reflected directly in pi_agent_rust architecture:

  • Single async substrate: runtime, timers, fs, and HTTP/TLS all run on one coherent foundation.
  • Explicit context threading: AgentCx wraps asupersync::Cx at subsystem boundaries (agent/tools/session/rpc).
  • Bounded operations over best-effort cleanup: timeout paths and compaction thresholds are parameterized and enforceable.
  • Determinism hooks for tests: timer-driver aware sleeps and asupersync test helpers reduce nondeterministic flakiness.

Compared to the original TypeScript implementation, this shifts more correctness responsibility into the runtime and core algorithms themselves, instead of relying primarily on ecosystem conventions.

Additional Major Deltas (Original pi-mono vs Rust Port)

This is a second comparison pass focused on high-impact architectural deltas and rationale.

Area Original pi-mono (packages/coding-agent) pi_agent_rust Why this divergence exists
Distribution model npm package (npm install -g @mariozechner/pi-coding-agent) Single Rust binary (pi) Remove Node runtime dependency and improve startup/deployment portability
Execution surfaces Interactive + print + JSON mode + RPC + SDK Interactive + print + JSON mode + RPC + Rust SDK Rust SDK provides idiomatic companion API for embedding Pi programmatically (documented in docs/sdk.md)
Default built-in tool posture Defaults to read/write/edit/bash (others available) Thirteen Essential-tier built-ins always in the schema (read/write/edit/bash/grep/find/ls/hashline_edit/ask/todo/web_search/submit_plan/xdev), with a discoverable tier behind the xdev dispatcher Keep common code-navigation, shell, and edit workflows available without extra configuration while bounding schema size
Extension trust model Extension/package model documented as full system access Embedded runtime with capability-gated hostcalls and policy profiles Reduce ambient authority and make extension behavior auditable/deny-by-default
Session architecture emphasis JSONL tree session model and branch navigation JSONL v3 tree + derived SQLite metadata index + default-enabled SQLite session backend support Bound eligible resume/lookups and coordinate multi-instance access
Streaming transport stack Node runtime networking stack Purpose-built HTTP/TLS client + custom SSE parser on asupersync Tighter control over chunking, parsing, and failure handling in long streams
Cancellation/timeout mechanics Platform/event-loop cancellation conventions Explicit abort signaling, bounded tool iterations, process-tree termination Minimize hangs/orphans and make stop behavior deterministic under load
Runtime context model Framework-level conventions and extension APIs Explicit AgentCx/asupersync::Cx capability-scoped context threading Make effect boundaries and testability first-class architectural constraints

Practical consequence of these deltas:

  • Extension and package workflows are selected according to current user value, with OMP as the closer product-surface and UX reference. Legacy pi-mono behavior may be adopted selectively when it remains useful.
  • The goal is a coherent Rust-native coding agent, not parity with pi-mono.
  • The Rust SDK provides a companion API for core embedding workflows without requiring TypeScript-specific adaptation patterns.
  • docs/parity-certification.json tracks informational functional-parity progress; it does not authorize strict replacement claims.

Algorithmic Mechanics: pi-mono Baseline vs Rust Implementation

This section compares concrete implementation mechanics for equivalent high-level behavior.

Algorithm pi-mono baseline mechanism Rust implementation mechanism Why the Rust variant exists
Session context rebuild after compaction buildSessionContext() emits compaction summary, then messages from firstKeptEntryId (pre-compaction path), then post-compaction entries to_messages_for_current_path() uses the same ordering and adds a fallback if first_kept_entry_id is missing Avoid silent context loss when compaction anchors are orphaned/corrupted
JSONL persistence Incremental append (appendFileSync) plus full rewrite (writeFileSync) for migrations/rewrites Save via temp file + atomic persist/replace Keep on-disk session state crash-resilient during save operations
Session discovery/resume Directory/file scan and mtime sorting of JSONL files Derived SQLite session metadata index + WAL + lock file + staleness-triggered full reindex Bound resume lookup cost and coordinate concurrent processes
Compaction token accounting Uses assistant usage (totalTokens else input+output+cacheRead+cacheWrite) plus heuristic trailing estimates Uses assistant usage (total_tokens else input+output) plus heuristic trailing estimates; fixed image token estimate Keep accounting stable across providers with uneven cache-token reporting while staying conservative
Cut-point + split-turn handling Valid cut points exclude tool results; split turns are summarized as history + turn-prefix context Same cut-point class and split-turn strategy, implemented in Rust entry/message model Preserve tool-call/result adjacency and turn coherence under budget pressure
Bash timeout/process cleanup Timeout/abort kills process tree (killProcessTree) and returns tail-truncated output Timeout escalation (TERM then grace then KILL) + process-tree walk + shell exit trap + tail truncation Enforce bounded cleanup and reduce descendant-process leaks from background jobs
Streaming event decoding Transport semantics are exposed (sse/websocket/auto); parser details are runtime-internal Explicit SSE parser with BOM stripping, CR/LF normalization, UTF-8 tail buffering, and flush-on-end Make byte-to-event behavior deterministic and provider-SDK-independent

Feature Superset Highlights (Beyond pi-mono Baseline)

The sections above compare mechanics. This section calls out concrete features present in this Rust port that are not part of the pi-mono baseline implementation model.

Rust-port feature Why it is useful/compelling
pi doctor diagnostics command (text/json/markdown, --only, --fix, swarm preflight, extension compatibility checks) Gives actionable environment + compatibility diagnostics, supports CI gating (non-zero on failures), can auto-fix safe issues like missing dirs/permissions, and reports read-only multi-agent readiness before swarm work
Capability-gated extension policy profiles (safe / balanced / permissive) with per-extension overrides Lets operators run shared extensions with explicit capability boundaries instead of ambient full-system access
Secret-aware extension env filtering (pi.env() blocklist for keys/tokens/secrets) Reduces accidental credential exposure from extension code paths
Per-extension trust lifecycle + kill-switch audit trail (pending/acknowledged/trusted/killed, kill_switch, lift_kill_switch) Supports immediate containment, explicit operator provenance, and controlled re-entry after review
Hostcall compatibility-lane emergency controls (global/per-extension forced-compat switches + reason codes) Gives operators a deterministic rollback path for fast-lane incidents without losing extension availability
Runtime risk controller for extension hostcalls (configurable, fail-closed by default) Adds another enforcement layer beyond static policy for suspicious runtime behavior in extension call flows
Argument-aware runtime risk scoring for shell paths (dcg_rule_hit, dcg_heredoc_hit, heredoc AST inspection across Bash/Python/JS/TS/Ruby) Detects destructive intent hidden in multiline scripts and wrapper commands before hostcall execution
Tamper-evident runtime risk ledger tooling (`ext_runtime_risk_ledger verify replay
Unified incident evidence bundle export (risk ledger, security alerts, hostcall telemetry, exec mediation, secret-broker events) Incident response can triage from one structured artifact set instead of stitching ad-hoc logs
Deterministic hostcall reactor mesh with optional NUMA slab pool (shard affinity, global-order drain, bounded SPSC lanes, telemetry) Keeps extension dispatch predictable under load and surfaces queue/backpressure behavior for tuning
Cold realm factory + persistent transpile cache Rebuilds mutable JS state for every load while reusing only versioned compiled artifacts that are safe to share across realms
Extension preflight static analysis (imports/forbidden-pattern scan with policy-aware hints) Catches risky extension patterns before runtime execution
Node/Bun-compatible extension runtime without Node/Bun dependency (embedded QuickJS + shims) Runs legacy extension workflows in a single native binary deployment model
Extension compatibility scanner + conformance harness Makes extension support measurable and auditable instead of anecdotal
Derived SQLite session metadata index (WAL + lock + stale reindex path) Gives fast session resume/list operations at scale without scanning every JSONL file on each query
Session Store V2 rollback and migration ledger (segmented log + checkpoints + rollback events) Long-session recovery can unwind to a known checkpoint with explicit migration/rollback provenance
Default-enabled SQLite session storage support (sqlite-sessions feature) Supports deployments that want database-backed session persistence in addition to JSONL; disable with --no-default-features when building a minimal binary
Crash-resilient session save path (temp file + atomic persist) Improves session-file durability during writes and reduces partial-write failure modes
Unified hostcall dispatcher with typed taxonomy mapping (timeout / denied / io / invalid_request / internal) Produces consistent extension/runtime error semantics and easier client handling
Fail-closed evidence-lineage gates (run_id/correlation_id + cross-artifact lineage checks) Rejects stale or cherry-picked conformance/perf artifacts at release-gate time
Structured auth diagnostics with stable machine codes Improves troubleshooting and operational visibility without leaking sensitive credential material

Deep Dive: Core Algorithms

Math-Driven Decision Systems

Pi deliberately uses advanced math where it improves runtime behavior or benchmark confidence. The goal is not “fancy formulas in docs”; it is safer policy decisions, faster recovery from workload shifts, and more trustworthy performance attribution.

Regime-Shift Detection (CUSUM + BOCPD)

In the extension dispatcher, Pi combines CUSUM and Bayesian online change-point detection to detect load-regime changes early (for example when hostcall traffic suddenly spikes or stalls).

$$ S_t^+ = \max\left(0,;S_{t-1}^+ + (-z_t - k)\right), \quad S_t^- = \max\left(0,;S_{t-1}^- + (z_t - k)\right) $$

$$ H(r)=\frac{1}{\lambda}, \quad P(r_t=0 \mid x_{1:t}) \propto \sum_r P(r_{t-1}=r),H(r),P(x_t \mid r) $$

Intuition: CUSUM catches persistent drift; BOCPD catches sudden regime changes without brittle fixed thresholds.

Conformal Prediction Envelope

Pi tracks nonconformity scores (absolute residuals from the running mean) and treats out-of-interval events as anomalies.

$$ q = \text{score}_{\lceil (n+1)\cdot \text{confidence} \rceil - 1}, \quad \text{anomaly if } |x_t - \mu_t| > q $$

Intuition: thresholds adapt from recent behavior instead of hard-coding one static latency cutoff.

PAC-Bayes Safety Bound

Pi’s safety envelope includes a PAC-Bayes-kl bound over extension outcomes, and can veto aggressive optimization when the bound is too high.

$$ \mathrm{kl}(\hat q ,|, q_{\text{bound}});\le;\frac{\mathrm{KL}(Q|P)+\ln!\left(2\sqrt{n}/\delta\right)}{n} $$

Intuition: this gives an explicit uncertainty-aware ceiling on true error risk before allowing more aggressive runtime behavior.

Off-Policy Evaluation (IPS/WIS/DR + ESS + Regret Gate)

Before approving policy moves, Pi evaluates candidate behavior from trace data:

$$ w_i=\frac{\pi(a_i\mid x_i)}{\mu(a_i\mid x_i)}, \quad \hat V_{\text{IPS}}=\frac{1}{n}\sum_i w_i r_i $$

$$ \hat V_{\text{WIS}}=\frac{\sum_i w_i r_i}{\sum_i w_i}, \quad \hat V_{\text{DR}}=\frac{1}{n}\sum_i\left(\hat r_i + w_i(r_i-\hat r_i)\right) $$

$$ N_{\text{eff}}=\frac{(\sum_i w_i)^2}{\sum_i w_i^2}, \quad \Delta_{\text{regret}}=\bar r_{\text{baseline}}-\hat V_{\text{DR}} $$

Intuition: Pi fails closed if sample support is weak, uncertainty is high, or estimated regret is above threshold.

VOI-Driven Experiment Selection

The VOI planner prioritizes probes that provide the most expected learning under a strict overhead budget.

$$ \text{priority}_i \propto \frac{\text{utility}_i}{\text{overhead}_i} $$

Intuition: run only the experiments that are likely to change decisions; skip stale or low-value probes.

Weighted Bottleneck Attribution (Benchmarking)

For phase-1 matrix benchmarking, Pi computes stage attribution weighted by realistic workload size (session_messages) and reports confidence intervals.

$$ \text{weighted_contribution}_s

\frac{\sum_i w_i,m_{i,s}}{\sum_i w_i,t_i}\cdot 100, \quad w_i=\text{session_messages}_i $$

$$ n_{\text{eff}}=\frac{(\sum_i w_i)^2}{\sum_i w_i^2}, \quad \mathrm{CI}{95}=\mu \pm 1.96\sqrt{\frac{\sigma^2}{n{\text{eff}}}} $$

Intuition: prioritize what dominates real end-to-end latency, not just isolated microbench hotspots.

Online Convex Control + Regret Tracking

Pi also includes an online tuner path for batch/time-slice controls with explicit rollback behavior:

$$ \tau_{t+1}

\mathrm{clip}!\left(\tau_t - \eta\nabla_{\tau}\mathcal{L}t,;\tau{\min},\tau_{\max}\right) $$

Intuition: the system adapts continuously, but if instantaneous loss exceeds a rollback threshold it immediately returns to a safer profile.

Math At a Glance

Technique Where in Pi Why it helps
CUSUM + BOCPD Extension dispatcher regime detector Detects traffic regime shifts early and robustly
Conformal intervals Safety envelope Adaptive anomaly gating without static magic numbers
PAC-Bayes bound Safety envelope veto path Fails closed when uncertainty/risk is too high
IPS/WIS/DR + ESS Off-policy evaluator Approves policy changes only with adequate support
VOI planning Experiment scheduler Uses overhead budget on highest-value probes
Weighted attribution + CI Phase-1 perf matrix reports Ranks optimization work by realistic user impact
OCO + regret rollback Runtime controller Adapts under load while bounding unsafe drift

SSE Streaming Parser

The SSE (Server-Sent Events) parser is a custom implementation that handles Anthropic's streaming response format. Unlike library-based approaches, the parser operates as a state machine that processes bytes incrementally:

Bytes → Line Accumulator → Event Parser → Typed StreamEvent

Key characteristics:

Property Implementation
Buffering Zero-copy where possible; lines accumulated only when incomplete
Event types 12 distinct variants: MessageStart, ContentBlockStart, ContentBlockDelta, ContentBlockStop, MessageDelta, MessageStop, Ping, Error, and thinking-specific events
Error recovery Malformed events logged but don't crash the stream
Memory Fixed-size rolling buffer prevents unbounded growth

The parser handles edge cases like:

  • Multi-line data: fields (concatenated with newlines)
  • Events split across TCP packet boundaries
  • The event: field appearing before or after data:
  • CRLF and LF line endings interchangeably

Truncation Algorithm

Large outputs from tools (file reads, command output, grep results) must be truncated to avoid exhausting the LLM's context window. The truncation algorithm preserves usefulness while staying within limits:

┌─────────────────────────────────────────┐
│           Original Content              │
│         (potentially huge)              │
└─────────────────────────────────────────┘
                    │
                    ▼
┌─────────────────────────────────────────┐
│  HEAD: First N/2 lines                  │
│  ─────────────────────────              │
│  [... X lines truncated ...]            │
│  ─────────────────────────              │
│  TAIL: Last N/2 lines                   │
└─────────────────────────────────────────┘

Constants:

Limit Value Rationale
MAX_LINES 2000 Balances context usage vs. completeness
MAX_BYTES 1MB Prevents binary file accidents
GREP_MAX_LINE_LENGTH 500 chars Truncates minified code

The algorithm:

  1. Splits content into lines
  2. If line count exceeds MAX_LINES, takes first 1000 and last 1000
  3. Inserts a marker showing how many lines were omitted
  4. If byte count still exceeds MAX_BYTES, applies byte-level truncation
  5. Returns metadata indicating truncation occurred, enabling the LLM to request specific ranges

Process Tree Management

The bash tool must handle runaway processes, infinite loops, and fork bombs without leaving orphans. The implementation uses the sysinfo crate to walk the process tree:

// Pseudocode for process cleanup
fn kill_process_tree(root_pid: Pid) {
    let system = System::new();
    let children = find_all_descendants(root_pid, &system);

    // Kill children first (deepest first), then parent
    for child in children.iter().rev() {
        kill(child, SIGKILL);
    }
    kill(root_pid, SIGKILL);
}

Timeout behavior:

  1. Command starts with configurable timeout (default 120s)
  2. Output streams to a rolling buffer in real-time
  3. On timeout: SIGTERM sent, 5s grace period, then SIGKILL
  4. Process tree walked and all descendants killed
  5. Exit code set to indicate timeout vs. normal termination

To avoid orphaned background jobs (e.g. cmd &), the bash script installs an EXIT trap that waits for any remaining child processes and then exits with the original command's status.

This prevents the common failure mode where killing a shell leaves its children running.

Session Tree Structure

Sessions use a tree structure rather than a flat list, enabling conversation branching (useful when exploring different approaches):

                    ┌─────────┐
                    │ Message │ (root)
                    │   #1    │
                    └────┬────┘
                         │
                    ┌────▼────┐
                    │ Message │
                    │   #2    │
                    └────┬────┘
                         │
              ┌──────────┼──────────┐
              │                     │
         ┌────▼────┐          ┌────▼────┐
         │ Message │          │ Message │ (branch)
         │   #3    │          │   #3b   │
         └────┬────┘          └────┬────┘
              │                    │
         ┌────▼────┐          ┌────▼────┐
         │ Message │          │ Message │
         │   #4    │          │   #4b   │
         └─────────┘          └─────────┘

JSONL format (v3):

Each line is a self-contained JSON object with a type discriminator:

{"type":"session","version":3,"cwd":"/project","created":"2024-01-15T10:30:00Z"}
{"type":"message","id":"a1b2c3d4","parent":"root","role":"user","content":[...]}
{"type":"message","id":"e5f6g7h8","parent":"a1b2c3d4","role":"assistant","content":[...]}
{"type":"model_change","id":"i9j0k1l2","parent":"e5f6g7h8","model":"claude-sonnet-4-20250514"}

The parent field creates the tree. Replaying a session walks the tree from root to the current leaf. Branching creates a new message with a different parent than the previous continuation.

Provider Abstraction

The Provider trait abstracts over different LLM backends:

#[async_trait]
pub trait Provider: Send + Sync {
    fn name(&self) -> &str;
    fn api(&self) -> &str;
    fn model_id(&self) -> &str;

    async fn stream(
        &self,
        context: &Context<'_>,
        options: &StreamOptions,
    ) -> Result> + Send>>>;
}

Context structure:

pub struct Context<'a> {
    pub system_prompt: Option>, // System prompt
    pub messages: Cow<'a, [Message]>,        // Conversation history
    pub tools: Cow<'a, [ToolDef]>,           // Available tools with JSON schemas
}

StreamOptions:

pub struct StreamOptions {
    pub temperature: Option,
    pub max_tokens: Option,
    pub api_key: Option,
    pub cache_retention: CacheRetention,
    pub session_id: Option,
    pub headers: HashMap,
    pub thinking_level: Option,
    pub thinking_budgets: Option,
}

This design allows adding new providers (OpenAI, Gemini) without modifying the agent loop. Each provider translates the common types to its wire format and emits a unified StreamEvent stream.

Compaction Algorithm

Long conversations eventually exceed the model's context window. Pi's compaction algorithm reclaims space by summarizing older messages while preserving recent context.

The algorithm runs automatically after each agent turn when estimated token usage exceeds context_window - reserve_tokens:

┌──────────────────────────────────────────────────────────────┐
│                     Full Conversation                         │
│  msg1 → msg2 → msg3 → ... → msgN-5 → msgN-4 → ... → msgN   │
│  ├──── older messages ─────┤ ├─── recent messages ──────────┤ │
│                                                              │
│  Step 1: Find cut point at a valid turn boundary             │
│  Step 2: LLM summarizes msgs 1..N-5 into compact paragraph  │
│  Step 3: Store Compaction entry in session JSONL             │
│  Step 4: Next agent call uses [summary] + msgs N-4..N       │
└──────────────────────────────────────────────────────────────┘

Token estimation counts tokens with real BPE tables (O200k/Cl100k, enabled by the default bpe-tokens feature); when BPE is unavailable it falls back to a conservative chars ÷ 3 heuristic for text, plus a flat 1,200 tokens per image. When an assistant message includes a usage field from the API, that measured value takes precedence over the estimate.

Cut point selection prefers boundaries between complete user-assistant turns. If the budget forces a mid-turn cut, the algorithm includes prefix messages from the split turn so the model retains context about what was being discussed at the boundary.

File operation tracking extracts read, write, and edit tool calls from the messages being summarized. The compaction prompt includes these paths so the summary preserves awareness of which files were examined or modified:


src/main.rs
src/config.rs



src/auth.rs

Configurable parameters:

Parameter Default Purpose
reserve_tokens 8% of context window Safety margin for response generation
keep_recent_tokens 10% of context window Minimum recent context preserved

Compaction can also be triggered manually with /compact in interactive mode or the compact RPC command.

Multi-Provider Routing & Model Registry

Pi routes model requests through a provider factory that resolves the correct backend implementation from a (provider, model, api) tuple.

Resolution flow:

User specifies --provider openai --model gpt-4o
               │
               ▼
  ┌──────────────────────────┐
  │  Provider Metadata Table  │  Maps "openai" → canonical ID,
  │                           │  determines API type (Completions
  │                           │  vs Responses vs custom)
  └────────────┬──────────────┘
               │
  ┌────────────▼──────────────┐
  │  URL Normalization         │  Appends /chat/completions,
  │                            │  /responses, or /chat depending
  │                            │  on detected API type
  └────────────┬──────────────┘
               │
  ┌────────────▼──────────────┐
  │  Compat Config             │  Applies per-model overrides:
  │                            │  system_role_name, max_tokens
  │                            │  field name, feature flags
  └────────────┬──────────────┘
               │
  ┌────────────▼──────────────┐
  │  Provider Instance         │  Anthropic | OpenAI | Gemini
  │                            │  Cohere | Azure | Bedrock | ...
  └───────────────────────────┘

models.json overrides: Users can define custom providers in ~/.pi/agent/models.json or .pi/models.json. Each entry specifies a model ID, base URL, API type, and optional compat flags, letting you route to self-hosted models, proxies, or providers that Pi does not natively support.

Compat config handles the differences between OpenAI-compatible APIs:

Override Example Purpose
system_role_name "developer" o1 models use "developer" instead of "system"
max_tokens_field "max_completion_tokens" Some models require a different field name
supports_tools false Suppress tool definitions for models that reject them
supports_streaming false Fall back to non-streaming for incompatible endpoints
custom_headers {"X-Custom": "val"} Per-provider header injection

Fuzzy matching: When a provider name doesn't match any known provider, Pi computes edit distance against all registered names and suggests the closest match in the error message.

Extension Hostcall Protocol

Extensions run in an embedded QuickJS runtime (rquickjs crate) and communicate with Pi through a structured hostcall protocol. This is the mechanism that lets JavaScript code invoke Pi's built-in tools, make HTTP requests, and interact with the session, all without direct OS access.

Execution model:

┌─────────────────── QuickJS VM ───────────────────┐
│                                                   │
│  extension.js calls:                              │
│    pi.tool("read", {path: "src/main.rs"})         │
│      │                                            │
│      ▼                                            │
│    enqueue HostcallRequest {                      │
│      call_id: "hc-0042",                          │
│      kind: Tool { name: "read" },                 │
│      payload: { path: "src/main.rs" },            │
│    }                                              │
│      │                                            │
│      ▼                                            │
│    return Promise (resolve/reject stored in map)  │
│                                                   │
└────────────────────────┬──────────────────────────┘
                         │
    drain_hostcall_requests()
                         │
                         ▼
┌─────────────── ExtensionDispatcher ──────────────┐
│                                                   │
│  1. Check capability policy:                      │
│     read tool → requires "read" capability        │
│     → Policy says: Allow / Deny / Prompt          │
│                                                   │
│  2. If allowed → dispatch to ToolRegistry         │
│     → Execute read tool                           │
│     → Get ToolOutput                              │
│                                                   │
│  3. complete_hostcall("hc-0042", Ok(result))      │
│     → Resolves the Promise in QuickJS             │
│                                                   │
│  4. runtime.tick()                                │
│     → Drains Promise .then() chains               │
│     → Extension JS continues execution            │
│                                                   │
└───────────────────────────────────────────────────┘

Capability mapping: Each hostcall kind maps to a required capability:

Hostcall Required Capability Dangerous?
pi.tool("read", ...) read No
pi.tool("write", ...) write No
pi.tool("bash", ...) exec Yes
pi.http(request) http No
pi.exec(cmd, args) exec Yes
pi.env(key) env Yes
pi.session(op, ...) session No
pi.ui(op, ...) ui No
pi.log(entry) log No (always allowed)

Deduplication: Each hostcall's parameters are canonicalized (object keys sorted, structure normalized) and SHA-256 hashed. Identical requests within a short window can be deduplicated to avoid redundant tool executions.

Fast lane vs compatibility lane: Pi has two execution lanes for hostcalls:

  • Fast lane is used when the call shape matches known safe patterns (for example common tool and session operations). This avoids extra allocation and parsing work.
  • Compatibility lane is the fallback for uncommon or partially-specified calls.
  • Both lanes still enforce the same capability policy and permission checks.
  • Operators can force compatibility-lane routing globally or per extension as an emergency control path.

For observability, each call is tagged with a stable lane key (for example tool|tool.read|filesystem or tool|fallback|filesystem) so latency and failure trends can be grouped consistently.

Built-in consistency guard (shadow dual execution): Pi can sample a small subset of read-only hostcalls, execute them through both lanes, and compare canonical output fingerprints. If divergence crosses a configured budget, Pi automatically backs off the fast lane for a period. This gives performance wins without silently changing behavior.

Adaptive dispatch mode under load: Pi can switch between:

  • sequential_fast_path for simpler/low-contention workloads
  • interleaved_batching when contention and queue pressure rise

Mode changes are gated by sample coverage and risk checks, so Pi does not switch based on thin or cherry-picked evidence.

Runtime telemetry for debugging and tuning: Pi records structured hostcall telemetry (pi.ext.hostcall_telemetry.v1) with lane choice, fallback reason, dispatch latency share, marshalling path, and optimization hit/miss fields. This is used by perf reports and reliability diagnostics.

Auto-repair pipeline: When an extension fails to load or produces runtime errors, Pi's repair system can automatically fix common issues:

Repair Mode Behavior
Off No repairs
Suggest Log suggestions, don't apply
AutoSafe (default) Apply provably safe fixes (missing file paths, asset references)
AutoStrict Apply aggressive heuristic fixes (pattern-based transforms)

Compatibility scanner: Before loading, Pi statically analyzes extension source code for imports, require() calls, and forbidden patterns (eval, Function(), process.binding, dlopen). The scan produces a capability evidence ledger that informs policy decisions.

Environment variable filtering: Extensions calling pi.env() hit a blocklist that denies access to API keys, credentials, tokens, and private keys. The filter blocks exact matches (ANTHROPIC_API_KEY, AWS_SECRET_ACCESS_KEY), suffix patterns (*_API_KEY, *_SECRET, *_TOKEN), and prefix patterns (AWS_SECRET_*, AWS_SESSION_*). Variables whose names do not match any secret pattern — including most PI_* variables — are served normally; there is no unconditional PI_* exemption, so a name like PI_EXAMPLE_API_KEY is blocked by the suffix rule.

Trust lifecycle and kill switch: Extension trust state is tracked explicitly (pending, acknowledged, trusted, killed). A kill switch demotes an extension to killed, quarantines it in the runtime risk controller, emits a critical alert, and writes an audit record. Lifting the switch requires an explicit operator action and moves the extension back to acknowledged.

Extension Runtime Decision Logic (Plain English)

The extension runtime includes a few small decision engines so behavior stays stable as workload patterns change:

  • Value-of-information planner (VOI): Ranks candidate probes by "expected learning per millisecond" and picks the best set under a strict overhead budget. Stale or low-value candidates are skipped with explicit reasons.
  • Shard load controller: Adjusts routing weights, batch budgets, and backoff/help factors based on queue pressure, latency, and starvation risk. Damping and oscillation guards prevent overreaction.
  • Policy safety evaluator: Replays historical samples with multiple estimators and only approves a policy when sample support is strong, uncertainty is low, and predicted regret stays within limit.

These pieces are intentionally conservative: if confidence is weak, Pi holds steady instead of making an aggressive switch.

Interactive TUI Architecture

The default interactive stack is FrankenTUI (src/interactive_ftui.rs, feature ftui, on by default since the 2026-08-25 cutover). It keeps the Elm Architecture (Model-Update-View): a driver thread owns an asupersync runtime plus an SDK agent session, agent events arrive through an AgentEventSubscription, and PiFtuiModel renders header, markdown conversation, status line, growing editor, and footer regions with tail-follow scrolling, per-entry render caching, inline ask cards, and modal overlays. All agent- and tool-originated text is sanitized before it reaches a frame. pi --inline draws the UI at the bottom of the terminal and preserves shell scrollback.

The previous stack, built on the charmed_rust library family (a Rust port of Go's Bubble Tea), lives in src/interactive.rs and is still selectable with pi --classic until it is deleted. The diagram below describes that classic stack; the FrankenTUI stack keeps the same agent/UI split and the same PiMsg event vocabulary.

Component stack (classic --classic stack):

┌────────────────────────────────────────────────────┐
│                 Terminal (crossterm)                │
│  Raw mode │ Alt screen │ Keyboard/Mouse events      │
└──────────────────────┬─────────────────────────────┘
                       │
┌──────────────────────▼─────────────────────────────┐
│             bubbletea Program Loop                  │
│  Init() → Update(Msg) → View() → render cycle      │
└──────────────────────┬─────────────────────────────┘
                       │
┌──────────────────────▼─────────────────────────────┐
│                  PiApp (Model)                      │
│                                                     │
│  ┌─────────────┐ ┌──────────────┐ ┌─────────────┐  │
│  │  TextArea    │ │  Viewport    │ │  Spinner     │  │
│  │  (editor)    │ │  (convo)     │ │  (status)    │  │
│  └─────────────┘ └──────────────┘ └─────────────┘  │
│                                                     │
│  ┌─────────────────────────────────────────────┐    │
│  │           Overlay Stack                      │    │
│  │  Model Selector │ Session Picker │ /tree     │    │
│  │  Settings UI    │ Theme Picker   │ Branches  │    │
│  │  Capability Prompt (extension UI)            │    │
│  └─────────────────────────────────────────────┘    │
└──────────────────────┬─────────────────────────────┘
                       │
              async channels (mpsc)
                       │
┌──────────────────────▼─────────────────────────────┐
│             Agent Async Task                        │
│  Runs on asupersync runtime                         │
│  Streams provider responses                         │
│  Executes tools                                     │
│  Sends PiMsg events back to TUI thread              │
└────────────────────────────────────────────────────┘

The async/sync bridge: The agent runs on the asupersync async runtime in a separate thread. It communicates with the bubbletea UI thread through mpsc channels. Each streaming event (text delta, tool start, tool update, agent done) becomes a PiMsg variant delivered to PiApp::update(), keeping the UI responsive during API streaming and tool execution.

Viewport scrolling: The conversation viewport tracks whether the user is at the bottom. When new content arrives and the user hasn't scrolled up, the viewport auto-follows the stream tail. Scrolling up disables auto-follow; pressing End or typing a new message re-enables it.

Overlay system: Modal UIs (model selector, session picker, branch navigator, extension capability prompts) stack on top of the main conversation view. Each overlay captures keyboard input until dismissed. Only the topmost active overlay receives events.

Slash commands available in the interactive editor:

Command Action
/help (/h, /?) Show available commands and keybindings
/model (/m) or Ctrl+L Open model selector with fuzzy search; /model roles edits role bindings
Ctrl+P / Ctrl+Shift+P Cycle scoped models forward/backward
/login, /logout Provider OAuth/API-key credential management
/thinking (/t) Change thinking level mid-conversation
/scoped-models (/scoped) Show or set model scope patterns
/tree, /fork Browse the conversation tree; branch from a previous message
/clear (/cls), /new Clear conversation; start a new session
/compact [shake|aggressive] Manual compaction (shake drops oversized tool results with no LLM call)
/checkpoint, /rewind, /fresh, /retry Session restore points and turn control
/undo, /redo Revert/replay file mutations made by tools
/resume (/r), /session (/info), /name, /history Session picker, info, naming, input history
/settings, /theme, /hotkeys (/keys), /changelog Settings UI, themes (incl. auto), keybindings, changelog
/plan, /approval, /advisor Plan mode, approval modes, second-model turn review
/btw , /tan Ephemeral smol-role side question; background task-role tangential work
/mcp, /usage, /rules, /omfg MCP server status, provider credit/quota, stream rules, grievances
/export, /share, /copy (/cp) HTML export, secret/unlisted GitHub Gist share (not private; anyone with the URL can view it), copy last reply
/handoff, /commit, /review Handoff document, dependency-ordered commit splitting, review
/template, /skill: Expand a prompt template; invoke a skill
/reload Reload skills/prompts/themes/extensions from disk
/exit (/quit, /q) or Ctrl+C Exit Pi

RPC Protocol

The RPC mode (pi --mode rpc) exposes a line-delimited JSON protocol over stdin/stdout for programmatic integration. Each line is a self-contained JSON object.

Client → Pi (stdin):

{"type": "prompt", "message": "Explain this function", "id": "req-001"}
{"type": "steer", "message": "Focus on error handling"}
{"type": "follow_up", "message": "Now add tests"}
{"type": "abort"}
{"type": "get_state"}
{"type": "compact", "reserveTokens": 8192, "keepRecentTokens": 20000}

Pi → Client (stdout):

{"type": "agent_start", "sessionId": "..."}
{"type": "message_update", "message": {...}, "assistantMessageEvent": {"type": "text_delta", "delta": "The function", "contentIndex": 0}}
{"type": "tool_execution_start", "toolCallId": "...", "toolName": "read", "args": {}}
{"type": "tool_execution_end", "toolCallId": "...", "toolName": "read", "result": {}, "isError": false}
{"type": "agent_end", "sessionId": "...", "messages": [...]}
{"type": "response", "id": "req-001", "command": "prompt", "success": true, "data": {"status": "ok"}}

Fatal errors (JSON and RPC modes): a failure that ends the process — a bad config file, no credentials for the selected model, an unwritable state directory, a usage error — prints exactly one record on stdout before the non-zero exit, so a host never has to parse stderr prose:

{"type": "error", "phase": "startup", "code": "auth.missing_api_key", "message": "No API key found for provider anthropic. Set env var or use --api-key.", "exit_code": 1}

phase is startup when nothing had been written to stdout yet (no session header, no RPC loop) and run otherwise. code is stable: the auth diagnostic codes (auth.missing_api_key, auth.no_models_available, auth.invalid_api_key, auth.quota_exceeded, auth.oauth.token_refresh_failed, …) when the failure classifies as one, else the family — config, session, provider, auth, tool, usage (argument/validation errors, exit code 2), extension, io, json, state_store, aborted, api, or internal. The human-readable diagnosis with hints still goes to stderr. Text mode prints nothing on stdout.

I/O architecture: Two dedicated threads handle stdin reading and stdout writing, bridged to the async agent runtime via channels. The stdin thread retries on transient errors to prevent dropped input. The stdout thread flushes after every line to prevent buffering delays.

Message queuing: While the agent is streaming a response, incoming messages are routed to one of two queues:

Queue Behavior Use Case
Steering Interrupts current response; processed on next turn Course corrections
Follow-up Queued until current response completes Sequential instructions

Queue modes (All or OneAtATime) control whether multiple queued messages are batched into a single turn or processed individually.

Extension UI over RPC: When an extension requests user input (capability prompt, selection dialog), Pi emits an extension_ui_request event. Response-bearing events include a requestGeneration correlation token, which the client must echo alongside requestId in its extension_ui_response; this prevents a late response from resolving a newer request that reused the same public ID. IDE extensions can then present native UI for capability decisions instead of falling back to terminal prompts.

Session Indexing

Session resume (pi -c or pi -r) needs to find the most recent session for the current project without scanning every JSONL file on disk. Pi maintains a SQLite index (session-index.sqlite) that provides constant-time lookups.

Schema:

CREATE TABLE sessions (
    path            TEXT PRIMARY KEY,
    id              TEXT NOT NULL,
    cwd             TEXT NOT NULL,
    timestamp       TEXT NOT NULL,
    message_count   INTEGER NOT NULL,
    last_modified   INTEGER NOT NULL,
    size_bytes      INTEGER NOT NULL,
    name            TEXT
);

Update lifecycle:

  1. After saving a session JSONL file, Pi upserts its metadata into the index
  2. pi -c queries WHERE cwd = ? ORDER BY last_modified DESC LIMIT 1
  3. pi -r queries the same table and presents a picker sorted by recency

Concurrency: A file-based lock (session-index.lock) serializes writes from concurrent Pi instances. Reads use WAL mode for non-blocking access.

Staleness-based reindexing: If the index is older than a configurable threshold, Pi runs a full re-scan of the sessions directory to catch files created by other instances or manual edits. The re-scan keeps the index accurate without a centralized daemon.

Session Store V2 Sidecar (Large Session Fast-Path)

Pi also supports a v2 sidecar store next to JSONL sessions for faster resume and stronger corruption checks on long histories.

What it adds:

  • Segmented append log files (instead of one ever-growing JSONL file)
  • Offset index rows for direct seeks and fast tail reads
  • Periodic checkpoints and a manifest snapshot
  • Migration ledger entries for auditability
  • Checkpoint-based rollback path with explicit rollback event logging

How resume works:

  1. A verified migration installs a durable clean source-state record, allowing the normal v2 open to use the sidecar index + segments without rescanning the full JSONL file.
  2. Before any JSONL mutation, Pi durably marks the sidecar dirty; dirty, invalid, or otherwise stale sidecars fall back to authoritative JSONL parsing. Filesystem mtimes remain a secondary external-change check, and legacy sidecars are fully verified once before use.
  3. If index data needs repair, Pi marks the sidecar dirty first and adopts the repaired store only after its entry IDs and hash chain match the source JSONL exactly.

Integrity strategy:

  • Segment frames carry payload and chain hashes.
  • Index rows store byte offsets plus CRC32C checksums.
  • Validation checks offset bounds, checksum matches, and frame/index alignment before trusting the sidecar.
  • Truncated trailing frames are recoverable during rebuild; non-EOF frame corruption fails closed instead of silently dropping data.

CLI support:

  • pi migrate --dry-run validates migration without writing.
  • pi migrate performs JSONL-to-v2 migration and verifies parity.

Authentication & Credential Management

Beyond simple API keys, Pi supports OAuth, AWS credential chains, service key exchange, and bearer-token auth. Credentials are stored in ~/.pi/agent/auth.json with file-locked access to prevent corruption from concurrent instances. Stored API keys can be literal strings, $ENV:VAR_NAME references, or $CMD:shell command / $COMMAND:shell command sources that resolve trimmed stdout at request time.

Mechanism Providers Details
API Key Anthropic, OpenAI, Gemini, Cohere, and many OpenAI-compatible providers Static key via env var or settings
OAuth Anthropic, OpenAI Codex, Google Gemini CLI, Google Antigravity, Kimi for Coding, GitHub Copilot, GitLab, and extension-defined OAuth providers PKCE/state-validated flow with automatic refresh; Kimi uses device flow
AWS Credentials Bedrock Access key + secret + optional session token; region-aware
Service Key SAP AI Core Client ID/secret exchange for bearer token
Bearer Token Custom providers Static token in auth storage

OAuth token lifecycle:

  1. User runs pi with an OAuth-configured provider
  2. Pi checks auth.json for an existing token
  3. If missing: opens browser to authorization URL, user authenticates, Pi receives authorization code, exchanges it for access + refresh tokens, stores both with expiry timestamp
  4. If expired but refresh token valid: exchanges refresh token for new access token, updates auth.json
  5. Bearer token attached to API requests

Google CLI-style OAuth providers carry project metadata with the token payload. Pi preserves and refreshes that payload and can resolve project IDs from GOOGLE_CLOUD_PROJECT or local gcloud config when needed.

Credential status reporting: pi config shows the status of each configured provider's credentials: Missing, ApiKey, OAuthValid (with time until expiry), OAuthExpired (with time since expiry), AwsCredentials, or BearerToken.

Diagnostic codes: Auth failures produce specific diagnostic codes (MissingApiKey, InvalidApiKey, QuotaExceeded, OAuthTokenRefreshFailed, MissingAzureDeployment, MissingRegion, etc.) with context-specific error hints rather than generic messages.


Tool Details

read

Read file contents (optionally images):

Input: { "path": "src/main.rs", "offset": 10, "limit": 50 }
  • Supports images (jpg, png, gif, webp) with optional auto-resize
  • Streams file bytes in chunks with hard size limits to reduce peak memory usage
  • Applies defensive image decode limits to block decompression-bomb/OOM inputs
  • Truncates at 2000 lines or 1MB
  • Returns continuation hint if truncated

bash

Execute shell commands with timeout and output capture:

Input: { "command": "cargo test", "timeout": 120 }
  • Default 120s timeout, configurable per-call
  • Set timeout: 0 to disable the default timeout
  • Process tree cleanup on timeout (kills children)
  • Rolling buffer for real-time output
  • Full output saved to temp file if truncated

edit

Surgical string replacement:

Input: { "path": "src/lib.rs", "old": "fn foo()", "new": "fn bar()" }
  • Exact string matching (no regex)
  • Fails if old string not found or ambiguous
  • Returns diff preview

grep

Search file contents:

Input: { "pattern": "TODO", "path": "src/", "context": 2, "limit": 100 }
  • Regex patterns supported
  • Context lines before/after matches
  • Respects .gitignore

find

Discover files by pattern:

Input: { "pattern": "*.rs", "path": "src/", "limit": 1000 }
  • Glob patterns matched in-process (no fd required)
  • Sorted by modification time
  • Respects .gitignore

ls

List directory contents:

Input: { "path": "src/", "limit": 500 }
  • Alphabetically sorted
  • Directories marked with trailing /
  • Truncates at limit

Performance Engineering

Why Rust Matters for CLI Tools

CLI tools have different performance requirements than servers or GUI applications. The critical metric is time-to-first-interaction: how quickly can the user start typing after invoking the command?

Phase Managed-runtime CLI Pi's native binary
Application runtime bootstrap Required Not required
Application module loading Commonly dynamic Ahead-of-time linked Rust code
JIT warmup Runtime-dependent Not used
Current release timing Not measured here Fresh v0.3.0 measurement pending

Current comparative timing numbers are intentionally omitted until a fresh, provenance-matched benchmark run is available.

Optimization Playbook

Pi applies performance-oriented design at multiple layers, not only in the Cargo release profile.

Where performance work is concentrated:

  • Startup path kept minimal: no JS runtime bootstrap, no module graph loading, no JIT warmup.
  • Hot hostcall specialization: common extension hostcalls use typed fast paths; uncommon shapes fall back to compatibility paths.
  • Adaptive dispatch under load: hostcall scheduling can switch modes when contention rises, then switch back when pressure drops.
  • Fast-path safety guardrails: sampled shadow dual-execution checks ensure optimizations do not silently change behavior.
  • Low-allocation rendering: TUI render buffers and markdown render results are cached/reused instead of rebuilt every frame.
  • Fast resume internals: session indexing plus the v2 sidecar layout avoid expensive full-history scans on resume.
  • Bounded growth controls: compaction and truncation keep token/context growth and tool-output payload growth from degrading responsiveness over long sessions.
  • Measurement-first culture: perf artifacts are schema-validated and claim-gated in CI, so optimization work is driven by evidence and regressions are caught early.

These mechanisms are intended to preserve responsiveness under streaming, tool, large-session, and extension workloads. Achieved release-level latency and memory results require fresh, provenance-matched measurements.

Optimization Catalog (Code + Commit History)

This catalog reflects a long sequence of changes across runtime, storage, streaming, and UI.

Concrete engineering work in this codebase includes:

Area What we built Why it matters
Extension dispatch core Typed hostcall opcode fast paths, compatibility fallback lane, zero-copy payload arena, canonical-hash shortcuts, interned operation paths Cuts per-call overhead on the hottest extension operations while preserving correctness fallback paths
Registry/policy lookup Immutable policy snapshots with O(1) capability checks, plus RCU-style metadata snapshots for extension registry/tool metadata Removes repeated dynamic lookup overhead in hot authorization/dispatch paths
Queueing + concurrency Core-pinned SPSC reactor mesh, S3-FIFO-inspired admission with fairness guards, BRAVO-style fallback behavior Improves tail latency under contention instead of only optimizing median latency
Batched execution AMAC-style interleaved batch executor with stall-aware toggling Avoids head-of-line stalls when many independent hostcalls are in flight
IO specialization io_