Open Executive
Open Executive is designed to transform leadership and management. Highly configurable, it can be deployed at any management level. Out of the box it supports spend approval thresholds, integration with corporate knowledge systems, and defined governance for how it interacts with human colleagues and other AI systems. Open Executive can also act as a busy leader's digital twin: with Act as me switched on, it drafts email in their own voice, from their own Gmail, for them to review and send—so leaders can truly scale their impact through AI without giving up the final say.
It meets people where they already work, with integrations for Slack, Discord, Telegram, Google Chat, email, Google Workspace, Microsoft 365 (Outlook mail and calendar, OneDrive), Notion, Confluence and any MCP-compatible AI tool, with more on the roadmap. Open Executive will always be open source.
A managed cloud offering is coming (https://openexecutive.ai), where you can get access without deploying anything yourself.
Demo
A walkthrough of Open Executive in action — watch on YouTube.
What It Does
Developed by sentelabs.ai Open Executive provides a single coherent executive voice backed by nine specialist AI agents:
- Chief Strategy Officer — competitive analysis, M&A, market positioning, OKRs
- Chief Financial Officer — financial modeling, fundraising, unit economics, cash flow
- Chief HR/People Officer — hiring, compensation, performance, culture
- General Counsel — contracts, IP, employment law basics, compliance
- Chief Operating Officer — process design, vendor management, operational scaling
- Chief Marketing Officer — GTM strategy, brand, communications, PR
- Chief Product Officer — roadmap, prioritization, product strategy
- Head of Sales — pipeline and qualification, founder-led sales, pricing conversations, proposals, forecasting
- Board Communications Director — board decks, investor relations, governance
All responses come from one consistent executive voice. The internal agent architecture is never exposed to the user. Beyond Q&A, the system maintains episodic memory of past decisions and initiatives across sessions, and a built-in scheduler can proactively surface follow-ups and time-sensitive actions.
Architecture
User message
↓
Executive Orchestrator (claude-sonnet-5)
↓ tool use → parallel specialist calls
CSO / CFO / CHRO / GC / COO / CMO / CPO / Board
↓ each specialist retrieves relevant context from ChromaDB
Built-in MBA knowledge + Your company documents
↓
Synthesized executive response
Knowledge — Two retrieval layers per specialist call: (1) built-in MBA-level Markdown (knowledge/builtin/, git-tracked) seeded into ChromaDB at startup, and (2) your uploaded company documents chunked and stored in a separate company_docs collection. RAG context is injected into the user turn, never the cached system prompt.
Episodic memory — After every response, a background claude-haiku-4-5 pass extracts key decisions, initiatives, and advice into SQLite. The next session opens with a `` block so the Executive remembers what it recommended last month.
Always in the loop — Each person can turn on "Keep track of what happens" (Settings → Memory). The Executive then keeps private, dated notes of what they said, in chat where it can confirm it's them (the web app signed in, their own Slack or Discord, Telegram with a webhook secret) and in Act as me replies they send. Every note rests on their own words, checked word for word. Only they see it, in Memories → History, where they can correct, pin or forget it. Notes are read back only to them, in a private chat, and expire after 90 days by default (the owner can change this). The owner's morning brief and evening digest include what the owner's own notes say is due, and anyone with notes on gets a short reminder on the day something they promised by email is due. It needs no setup and no extra service.
Scheduler — A built-in job runner claims due actions via UPDATE … RETURNING to prevent double-firing. The API must run as a single instance; do not horizontally scale it without gating the scheduler first.
Prompt caching — The system prompt is structured so the Executive persona, company profile, and knowledge index are cached separately (up to 85% cache hit rate after the first few turns). No dynamic content ever goes in a cached block.
See docs/architecture.md for the full design.
Tech Stack
| Layer | Choice |
|---|---|
| LLM backbone | Anthropic Claude API |
| Default model | claude-sonnet-5 (Executive + most specialists) |
| Deep reasoning | claude-opus-5 (CSO, CFO, GC, Board — with extended thinking) |
| Backend | Python 3.11 + FastAPI |
| Package manager | uv |
| Vector store | ChromaDB (local, embedded) |
| Episodic memory | SQLite |
| Web UI | Next.js 15 (App Router) + Tailwind |
| License | Apache 2.0 |
Repo Layout
openexecutive/
├── packages/
│ ├── core/
│ │ └── openexecutive/
│ │ ├── orchestrator/ # Executive persona + routing loop
│ │ ├── agents/ # 9 specialist agents
│ │ ├── knowledge/ # ChromaDB store + RAG pipeline
│ │ ├── memory/ # Company profile + episodic memory
│ │ ├── onboarding/ # Wizard state machine + profile builder
│ │ ├── prompts/ # Persona + domain prompts + cache manager
│ │ ├── api/ # FastAPI app + routes
│ │ ├── integrations/ # Slack, Email (Gmail or Outlook), Telegram, Google Chat, Discord
│ │ ├── scheduler/ # Background job runner (single-instance)
│ │ ├── alerts/ # Proactive alert system
│ │ ├── audit/ # Audit logging
│ │ ├── architecture/ # Internal architecture utilities
│ │ ├── workflows/ # Multi-step workflow definitions
│ │ └── cli.py # Click CLI
│ └── ui/ # Next.js 15 web UI
├── evals/ # Eval scenarios + LLM-as-judge runner
├── fixtures/ # Demo company fixtures (profiles, docs, rosters)
├── scripts/ # Operator scripts (Google auth)
├── docker/ # Dockerfile(s) + docker-compose.yml
└── docs/ # Architecture + deployment docs
Quick Start
# Clone the repo
git clone https://github.com/SenteLabsAI/OpenExecutive.git
cd OpenExecutive
# Install the backend and web app dependencies
make install
# Set your Anthropic API key
cp .env.example .env
# Edit .env and add ANTHROPIC_API_KEY=sk-ant-...
# Start everything
make dev
Trying it on your own computer needs no sign-in setup. While
AUTH_GOOGLE_ID and AUTH_OIDC_ISSUER are blank, make dev uses local login: the web app
only accepts connections from this computer, you click Open, and you are
the owner. To invite your team, or to run it on a server or with make docker,
set up Google or SSO sign-in by filling in the AUTH_* block (see
docs/auth.md for the Google Cloud Console and SSO steps).
All configuration lives in that repo-root .env — make dev and make docker
both load it for the API and the UI (Auth.js needs AUTH_SECRET /
AUTH_GOOGLE_ID / AUTH_GOOGLE_SECRET at runtime). A
packages/ui/.env.local is also read for UI-only keys, but for keys present
in both files the root .env takes precedence.
Open http://localhost:3000 to start chatting with your executive. The API runs on port 8000 and the UI on 3000.
To check everything is working, open Settings → Setup status. It tests the AI key and each chat app's token, then shows a green, amber or red light for each part of your setup, with what to do about anything that isn't green.
First run: requires Python 3.11+ and Node 22.6+. The initial
uv syncpulls heavy ML dependencies (ChromaDB + sentence-transformers/PyTorch), and the first boot downloads a small embedding model (~90 MB) to build the local vector index — so the firstmake devtakes a few minutes before the app is ready. Subsequent starts are fast.
On Windows: run
makefrom Git Bash or WSL, not PowerShell orcmd. The recipes are POSIX shell (if [ -f .env ]; …), and GNU Make falls back tocmd.exewhen noshis on PATH — which fails with-f was unexpected at this time. If Make still picks the wrong shell, point it at one:make dev SHELL="C:/Program Files/Git/bin/sh.exe". Note thatmake stopuseslsofand has no Windows equivalent; stop the two dev servers from their own terminals instead.
For contributors not using make:
cd packages/core
uv sync
uv run uvicorn openexecutive.api.main:app --reload --port 8000
# In a second terminal
cd packages/ui
npm install
npm run dev
uv run executes inside the project's virtualenv without activating it, so
these commands are the same on macOS, Linux and Windows. (Activating manually
works too, but the path differs per platform: .venv/bin/activate on
macOS/Linux, .venv\Scripts\Activate.ps1 on Windows.) The packages/ui
commands are listed one per line rather than chained with && so that they
run in every shell too: Windows PowerShell 5.1, the version that ships with
Windows, has no && operator and rejects the chained form with The token '&&' is not a valid statement separator in this version.
Set Up Integrations
Every integration is optional and off until you configure it. Each one has a setup guide:
| Integration | What it does | Setup |
|---|---|---|
| Slack | Answers mentions and DMs | Set SLACK_BOT_TOKEN and SLACK_APP_TOKEN (socket mode) |
| Discord | Answers DMs, mentions and slash commands | Run the Discord Bot |
| Telegram | Answers messages to your bot | docs/telegram_setup.md |
| Google Chat | Answers mentions in a space | docs/google_chat_setup.md |
| Email and calendar (Google Workspace) | Reads and sends the Executive's Gmail, books on its calendar | Google Workspace credentials |
| Email and calendar (Microsoft 365) | The same with Outlook mail and calendar | Microsoft 365 credentials |
| Act as me | Drafts replies in the owner's own Gmail for them to send | Your own Gmail |
| Google Drive folders | Syncs shared folders into the knowledge base (read-only) | docs/drive_sync_setup.md |
| OneDrive folders | Syncs chosen OneDrive or SharePoint folders into the knowledge base, as the Executive's Microsoft 365 sign-in | docs/onedrive_sync_setup.md |
| OneDrive in chat | Search, read, upload, move, copy and share OneDrive files (no delete, no anonymous links) | Set MS365_MCP_ONEDRIVE=true and sign in again (docs/onedrive_sync_setup.md) |
| Confluence | Syncs chosen spaces into the knowledge base (read-only, Cloud or Server/Data Center) | docs/confluence_sync_setup.md |
| Notion | Syncs the pages shared with an internal integration into the knowledge base | Set NOTION_SYNC_ENABLED=true and NOTION_API_KEY (see .env.example) |
| MCP tools | Gives the Executive tools from other MCP servers, with deny rules for destructive ones | Copy packages/core/mcp_servers.json.example to company/mcp_servers.json and list your servers under mcpServers; MCP turns on when the file exists (Docker keeps it on /data, see docs/deployment.md) |
| MCP clients | Lets any MCP client talk to the Executive | Connect as an MCP Server |
Synced Drive, OneDrive, Confluence and Notion content is kept apart from your uploaded documents and labelled as unreviewed, so it ranks below them.
Run the Discord Bot
- Create a Discord application at https://discord.com/developers/applications
- Enable the Message Content privileged intent (Bot → Privileged Gateway Intents)
- Invite the bot with
bot+applications.commandsscopes - Set env vars in
.env:DISCORD_BOT_TOKEN,DISCORD_APP_ID,DISCORD_GUILD_IDS - Run the API normally — the bot starts as part of the FastAPI lifespan when
DISCORD_BOT_TOKENis set:
make dev
The bot is embedded in the API process (alongside the email poller, scheduler, and resumer) so it shares the same SQLite database and ChromaDB vector store under /data in production. Skip the token to disable.
For iterating on bot-only code without restarting the API, make discord runs the bot as a standalone process against the same local DB. Use it instead of the embedded bot, not alongside it: stop the API, or start it without DISCORD_BOT_TOKEN, first — two processes on one token each open a gateway connection and every message gets answered twice.
Users can DM the bot, @mention it in a channel (replies in a thread), or use /ask and /today slash commands. Slash commands sync to DISCORD_GUILD_IDS instantly on startup; leave blank for global registration (up to 1-hour propagation delay).
Deploying to production
The bot runs inside the existing API process — no extra service. Set these on the API and restart it:
DISCORD_BOT_TOKEN=...
DISCORD_APP_ID=...
DISCORD_GUILD_IDS=...
Discord user access is managed via the /people UI — add a Person row with discord_user_id set.
The bot starts on the next lifespan boot. To disable it, unset DISCORD_BOT_TOKEN and restart.
Onboarding Your Company
The first time you visit the app, you'll be guided through a wizard to set up your company profile:
- Company basics (name, industry, stage, team size)
- Business model and revenue
- Competitive landscape
- Strategic priorities
- Culture and values
- Optional: financial position, document upload
After onboarding, the Executive will reference your specific company context in every response.
The built-in knowledge base is trusted by default — the Executive can use it straight away, and the Review queue starts empty rather than asking you to sign off on ~81 shipped documents. Items reach the queue when you upload a document, edit an existing file, or deliberately send a domain for review from the Review page. While a domain is being reviewed its content is withheld from the Executive, so the UI confirms the count before it starts.
Interfaces
| Interface | How to Use |
|---|---|
| Web UI | http://localhost:3000 |
| Slack | Mention @OpenExecutive or DM the app |
| CC or email the configured address (Gmail via the Google Workspace MCP, or Outlook via the Microsoft 365 MCP) | |
| Telegram | Message the configured bot |
| Google Chat | Mention the app in a space |
| Discord | DM the bot, @mention it in a channel, or use /ask / /today slash commands |
| CLI | openexecutive chat |
| MCP | Point any MCP client at http://localhost:8000/mcp (see Connect as an MCP Server) |
Connect as an MCP Server
Open Executive exposes its company context and specialist council to external
MCP clients (Claude Code, Claude Desktop, Cursor, or any other agent), so they
can ground themselves in your company without being re-briefed. There is no
separate server process to start: the server is mounted into the FastAPI app at
/mcp over Streamable HTTP, so running the API is running the MCP server.
The endpoint is http://localhost:8000/mcp, or https:///mcp once
deployed. Authentication is the same shared-secret gate as every other route,
so clients send x-api-key: $BACKEND_SHARED_SECRET. With that variable unset
locally the gate is off and the header can be omitted.
On any internet-reachable instance set both BACKEND_SHARED_SECRET and
OE_PUBLIC_DEPLOYMENT=1, which makes the API refuse to start without a
secret rather than serving /mcp unauthenticated. See
Deployment.
{
"mcpServers": {
"open-executive": {
"type": "http",
"url": "http://localhost:8000/mcp",
"headers": {
"x-api-key": "YOUR_BACKEND_SHARED_SECRET"
}
}
}
}
Claude Code can write that entry for you:
claude mcp add --transport http open-executive http://localhost:8000/mcp \
--header "x-api-key: YOUR_BACKEND_SHARED_SECRET"
Connected clients get eight read-only resources (company profile, today's
briefing and activity, people roster, department state, and episodic memory for
decisions, initiatives and advice) and four tools, of which consult_specialist
is the primary one: domain analysis from any of the ten specialists, each
grounded in your company's knowledge base. The full inventory is on the MCP
Server section of the /architecture page. With
signed callers on, an MCP client holds only the
shared secret, so it is never the owner: ask_executive asks as no one.
This is the inverse of the MCP_ENABLED and MCP_SERVERS_CONFIG_PATH settings
under Configuration, which configure the MCP gateway: Open
Executive consuming other servers' tools. The server side needs no configuration
beyond BACKEND_SHARED_SECRET. If a client errors on connect, try /mcp/ with
the trailing slash, since the endpoint 307-redirects and not every client
follows the redirect.
Document Upload
Upload your pitch deck, financial model, strategy docs, or any company documents via the web UI or API. The Executive will reference them when relevant.
# Via CLI
openexecutive upload deck.pdf model.xlsx strategy.md
# Via API
curl -X POST http://localhost:8000/documents \
-F "[email protected]" \
-F "domain=strategy"
Deployment
Two containers — the FastAPI backend and the Next.js UI — plus one persistent
volume at /data. docker/docker-compose.yml is the
reference topology and also what make docker runs locally, except that the
compose UI runs as a dev server rather than the production image.
Versioned images are published to GitHub Container Registry on every release, so a deployment can pull instead of building:
docker pull ghcr.io/sentelabsai/openexecutive-api:
docker pull ghcr.io/sentelabsai/openexecutive-ui:
Tags: X.Y.Z and X.Y for a release, latest for the most recently
published release, and main for the current head of main. The available versions are listed on
each package's page under the repository's Packages. See
docs/deployment.md for how releases are cut.
Upgrading. Settings → About shows the running version and says when a newer
release is out (UPDATE_CHECK_ENABLED=false turns the check off). To upgrade,
back up, then rebuild or pull the new tag; see
docs/deployment.md.
⚠️ Single-instance only: the scheduler claims rows via
UPDATE … RETURNING, which is not safe across processes. A second API replica double-fires every scheduled action. Pin the API to one instance. The UI is stateless.
Set ANTHROPIC_API_KEY, BACKEND_SHARED_SECRET, BACKEND_ALLOWED_ORIGINS and
OE_PUBLIC_DEPLOYMENT=1 on any internet-reachable instance. See
docs/deployment.md for the full guide — persistent state,
health-check timing, resource sizing, operations, and common failure modes.
Access control
The deployed UI is gated behind Google sign-in or SSO sign-in (any OpenID Connect provider, such as Keycloak, Okta or Entra ID) with an email allow-list, and the public API is protected by a shared-secret header between the UI proxy and the FastAPI backend. See docs/auth.md for the full setup (Google Cloud Console and SSO steps, required environment variables, adding/removing users, rotating secrets, and a debugging table).
Configuration
All settings via environment variables. Minimum required: ANTHROPIC_API_KEY —
unless you configure a local or OpenRouter backend instead (see Running on
Local Models). At least one provider must be set or
the app refuses to start.
| Variable | Required | Default | Description |
|---|---|---|---|
ANTHROPIC_API_KEY |
Yes¹ | — | Anthropic API key |
ANTHROPIC_WORKSPACE_ID |
No | — | Required only for an organisation-scoped Anthropic key; sent as the anthropic-workspace-id header. Workspace-scoped keys need no value |
DEFAULT_MODEL |
No | claude-sonnet-5 |
Executive + most specialists |
DEEP_REASONING_MODEL |
No | claude-opus-5 |
CSO, CFO, GC, Board |
VECTOR_STORE_PATH |
No | ./chroma_db |
ChromaDB directory |
EPISODIC_DB_PATH |
No | ./episodic_memory.db |
SQLite for episodic memory |
COMPANY_PROFILE_PATH |
No | ./company/profile.yaml |
Company profile |
ENABLE_CACHING |
No | true |
Anthropic prompt caching |
ROUTING_MODEL |
No | claude-haiku-4-5 |
Model for intent routing |
SLACK_BOT_TOKEN |
No | — | Slack bot OAuth token |
SLACK_APP_TOKEN |
No | — | Slack socket mode token |
EXEC_EMAIL_ADDRESS |
No | — | The Executive's own mailbox (Gmail via Google Workspace MCP, or Outlook via Microsoft 365 MCP) |
EMAIL_POLL_INTERVAL_SECONDS |
No | 60 |
How often to poll for new email |
TELEGRAM_BOT_TOKEN |
No | — | Telegram bot token (from @BotFather) |
TELEGRAM_WEBHOOK_SECRET |
No | — | Random string for webhook validation |
DISCORD_BOT_TOKEN |
No | — | Discord bot token (Developer Portal → Bot tab) |
DISCORD_APP_ID |
No | — | Discord application ID (General Information tab) |
DISCORD_GUILD_IDS |
No | — | Comma-separated guild IDs for dev slash-command registration |
DISCORD_NOTIFY_CHANNEL_ID |
No | — | Default channel ID for outbound notifications |
GOOGLE_CHAT_PROJECT_NUMBER |
No | — | GCP project number for Google Chat |
GOOGLE_CHAT_SERVICE_ACCOUNT_FILE |
No | — | Path to service account JSON key |
GOOGLE_OAUTH_CLIENT_ID |
No | — | Google OAuth client ID (Gmail MCP) |
GOOGLE_OAUTH_CLIENT_SECRET |
No | — | Google OAuth client secret (Gmail MCP) |
MS365_MCP_CLIENT_ID |
No | — | Entra app registration client ID (Microsoft 365 MCP: Outlook mail + calendar) |
MS365_MCP_TENANT_ID |
No | common |
Entra tenant ID for the Microsoft 365 MCP sign-in |
MS365_MCP_CLIENT_SECRET |
No | — | Entra client secret (only for a confidential app registration) |
MS365_MCP_EXPECTED_USERNAME |
No | — | Pin the Microsoft 365 sign-in to this mailbox UPN |
MS365_MCP_ONEDRIVE |
No | false |
Add the OneDrive tools (and the file permissions they need) to the Microsoft 365 sign-in; run --login again after turning it on |
MCP_ENABLED |
No | on when the config file exists | MCP gateway (the Executive using other servers' tools); false keeps it off with the file in place |
MCP_SERVERS_CONFIG_PATH |
No | ./company/mcp_servers.json |
MCP gateway server list and tool filters (see mcp_servers.json.example) |
EMAIL_PROVIDER |
No | google |
Backend for the inbound mailbox poller + alert email: google or microsoft |
CALENDAR_PROVIDER |
No | google |
Backend for the typed calendar booking tools: google or microsoft |
DELEGATION_GOOGLE_CREDENTIALS_DIR |
No | — | Where the owner's own Gmail credential for Act as me lives (see docs/deployment.md) |
DRIVE_SYNC_ENABLED |
No | false |
Sync shared Google Drive folders into the knowledge base |
DRIVE_SYNC_SERVICE_ACCOUNT_FILE |
No | — | Service account key with drive.readonly |
DRIVE_SYNC_FOLDER_IDS |
No | — | Comma-separated Drive folder IDs to sync |
ONEDRIVE_SYNC_ENABLED |
No | false |
Sync OneDrive or SharePoint folders into the knowledge base, as the Microsoft 365 sign-in |
ONEDRIVE_SYNC_FOLDERS |
No | — | Comma-separated / entries (openexecutive onedrive-folder prints one) |
CONFLUENCE_SYNC_ENABLED |
No | false |
Sync Confluence spaces into the knowledge base |
CONFLUENCE_URL |
No | — | Confluence base URL (Cloud or Server/Data Center) |
CONFLUENCE_USERNAME / CONFLUENCE_API_TOKEN |
No | — | Confluence Cloud sign-in (email + API token) |
CONFLUENCE_PERSONAL_TOKEN |
No | — | Confluence Server/Data Center personal access token |
CONFLUENCE_SYNC_SPACE_KEYS |
No | — | Comma-separated space keys to sync |
NOTION_SYNC_ENABLED |
No | false |
Sync the Notion pages shared with your integration into the knowledge base |
NOTION_API_KEY |
No | — | Notion internal integration token |
OPENROUTER_ENABLED |
No | false |
Route Claude calls through OpenRouter and unlock non-Anthropic models per-agent in the Council UI |
OPENROUTER_API_KEY |
No | — | Required when OPENROUTER_ENABLED=true |
OPENROUTER_CATALOG_ENABLED |
No | true |
Fetch OpenRouter's live /models catalog at startup to populate the Council dropdown; falls back to a built-in list on failure |
OPENROUTER_CATALOG_PROVIDERS |
No | openai,google,anthropic,meta-llama,deepseek,x-ai |
Vendor prefixes surfaced from the live catalog |
OPENROUTER_CATALOG_PER_PROVIDER |
No | 6 |
Newest tool-capable paid models per vendor (0 = no cap) |
OPENROUTER_CATALOG_REFRESH_S |
No | 21600 |
Background re-fetch cadence in seconds (0 = startup only) |
LOCAL_MODELS_ENABLED |
No | false |
Route selected slugs to a local OpenAI-compatible server (Ollama, LM Studio, vLLM, llama.cpp) |
LOCAL_BASE_URL |
No | — | Local server URL incl. version path, e.g. http://localhost:11434/v1. Required when LOCAL_MODELS_ENABLED=true |
LOCAL_API_KEY |
No | — | Optional bearer token (vLLM / gateways); Ollama & LM Studio need none |
LOCAL_MODELS |
No | — | Comma-separated local model slugs to surface in the Council UI and route locally, e.g. llama3.3,qwen2.5 |
LOCAL_TIMEOUT_S |
No | 300 |
Per-call timeout for local generation, in seconds |
LOCAL_REASONING_EFFORT |
No | — | reasoning_effort sent on every local request: none, minimal, low, medium or high. Set it for thinking-only models (GLM on Fireworks) that otherwise spend the whole token budget reasoning |
HONCHO_ENABLED |
No | false |
Per-person memory layer (honcho.dev) — a peer card shared across all channels; on the Memories page each person sees only their own |
HONCHO_API_KEY |
No | — | Required when HONCHO_ENABLED=true |
HONCHO_BASE_URL |
No | — | Self-hosted Honcho endpoint |
ENABLE_WEB_SEARCH |
No | true² |
Let the Executive and specialists answer with live web results (news, market data, competitor moves) alongside your uploaded documents |
WEB_SEARCH_MAX_USES |
No | 2 |
Max billed searches per agent per turn |
See .env.example for the full list.
¹
ANTHROPIC_API_KEYis required only when you serve Claude models directly. It can be omitted entirely if you run on local models (LOCAL_MODELS_ENABLED) or route through OpenRouter (OPENROUTER_ENABLED).
² The application default is on, but .env.example ships
ENABLE_WEB_SEARCH=falseso a fresh setup incurs no per-search charges — if the agents tell you they can't search the web or read the news, flip it totruein your.envand restart. Uses Anthropic's server-sideweb_searchtool, so it applies to Claude models (local models can't use it).WEB_SEARCH_ALLOWED_DOMAINS/WEB_SEARCH_BLOCKED_DOMAINSscope where it may look (set at most one).
Running on Local Models
Open Executive can run against any OpenAI-compatible local server — Ollama, LM Studio, vLLM, or llama.cpp — instead of (or alongside) the Anthropic API. Local model slugs route to your server through the same provider abstraction the hosted models use; no agent or orchestrator code changes.
# 1. Pull a capable, tool-use-friendly model (example: Ollama)
ollama pull llama3.3
# 2. In .env — point at the local server and list the slugs to expose
LOCAL_MODELS_ENABLED=true
LOCAL_BASE_URL=http://localhost:11434/v1 # Ollama default
LOCAL_MODELS=llama3.3
# 3. (Optional) run with NO Anthropic key — make local the default everywhere
DEFAULT_MODEL=llama3.3
DEEP_REASONING_MODEL=llama3.3
ROUTING_MODEL=llama3.3
# ...and leave ANTHROPIC_API_KEY unset
The listed slugs appear in the Council UI model dropdown, so you can also run a hybrid setup — keep the Executive on Claude while flipping individual specialists to a local model per-agent.
Caveats. Server-side web search (ENABLE_WEB_SEARCH) and Anthropic prompt
caching / extended thinking have no local equivalent and are automatically
disabled for local models. Multi-agent routing leans heavily on tool use, so
pick a model that's strong at it (e.g. Llama 3.3 70B, Qwen2.5) — small models
may route poorly. LOCAL_API_KEY is only needed if your server (vLLM, or a
gateway) requires a bearer token; Ollama and LM Studio need none.
Using a hosted OpenAI-compatible gateway
The LOCAL_* settings are not limited to localhost — the same recipe works with
any hosted OpenAI-compatible endpoint (an aggregator or inference gateway):
LOCAL_MODELS_ENABLED=true
LOCAL_BASE_URL=https://gateway.example.com/v1
LOCAL_API_KEY=your-gateway-key
LOCAL_MODELS=vendor/model-a,vendor/model-b
The listed slugs are sent to the gateway verbatim and appear in the Council UI
dropdown, exactly like local slugs. Gateways that speak OpenRouter's request
format and model namespace can alternatively be used through the OpenRouter
path (OPENROUTER_ENABLED=true + OPENROUTER_API_KEY +
OPENROUTER_BASE_URL=https://gateway.example.com/v1), which keeps that path's
Claude-name translation and feature handling.
The local-model caveats above apply to the LOCAL_* path unchanged, plus one
that matters more for hosted endpoints: everything the Executive processes —
company profile, documents, conversations — is sent to whichever endpoint you
configure here. Point these settings only at a provider you trust with that
data.
Adding a New Specialist Agent
- Create
packages/core/openexecutive/agents/your_agent.pyextendingBaseAgent - Add a system prompt constant in
packages/core/openexecutive/prompts/domain_prompts.py - Register in
packages/core/openexecutive/orchestrator/router.py— add toSPECIALIST_REGISTRYand thespecialistenum inSPECIALIST_TOOLS - Add domain alias to
DOMAIN_ALIASESinpackages/core/openexecutive/knowledge/retriever.py - Add knowledge docs to
knowledge/builtin/your_domain/ - Add at least 2 eval scenarios to
packages/core/openexecutive/evals/_scenarios/ - Submit a PR — CI requires all of the above
Development
make dev # Start FastAPI + Next.js
make test # Run Python tests
make eval # Run eval suite
make lint # Run ruff + mypy
make docker # Build and run Docker stack
# Unit tests only (no API calls required)
pytest packages/core/tests/unit/ -v
Evaluation System
evals/ contains 29 scenarios covering all 8 domains, scored by claude-opus-4-7 as an LLM-as-judge. Each scenario defines a query, simulated company context, expected topics, required specialist routing, and a domain-specific rubric. Five scoring dimensions (persona coherence, domain accuracy, company context utilization, routing quality, actionability) are each rated 1–5. The CI gate requires ≥ 3.5/5 average; any dimension dropping > 10% vs main fails the PR.
Privacy
Everything in company/ is gitignored — the profile YAML, uploaded documents, and the ChromaDB vector store. None of this leaves your local machine (or your own volume in cloud deployments) except as part of prompts sent to the Anthropic API. Anthropic does not train on API data.
Inside the app, what the Executive keeps about a person is theirs: Always in the loop notes and the peer memory card are shown only to that person, the owner included.
Contributing
See .github/CONTRIBUTING.md. All PRs must include:
- Working implementation (no stubs)
- Tests for new behavior
- Eval scenarios for new agents or prompt changes
License
Apache 2.0 — free to use commercially, requires attribution.
