← Open Source
vxcontrol

pentagi

Fully autonomous AI Agents system capable of performing complex penetration testing tasks

ApplicationsSpecialistsGo
Open on GitHub
Momentum
+18stars in 24 hours+0.1%
25.3k
Stars
3.25k
Forks
+191
This week
15
Contributors
Created 2025-01-06 · Updated 2026-10-05 · #557 today
Top developers
README

PentAGI

Penetration testing Artificial General Intelligence

Join the Community! Connect with security researchers, AI enthusiasts, and fellow ethical hackers. Get support, share insights, and stay updated with the latest PentAGI developments.

Discord⠀Telegram

vxcontrol%2Fpentagi | Trendshift vxcontrol%2Fpentagi | Trendshift

vxcontrol%2Fpentagi | Trendshift vxcontrol%2Fpentagi | Trendshift vxcontrol%2Fpentagi | Trendshift

Table of Contents

Overview

PentAGI is an innovative tool for automated security testing that leverages cutting-edge artificial intelligence technologies. The project is designed for information security professionals, researchers, and enthusiasts who need a powerful and flexible solution for conducting penetration tests.

You can watch the video PentAGI overview: PentAGI Overview Video

Features

  • Secure & Isolated. All operations are performed in a sandboxed Docker environment with complete isolation.
  • Fully Autonomous. AI-powered agent that automatically determines and executes penetration testing steps with optional execution monitoring and intelligent task planning for enhanced reliability.
  • Professional Pentesting Tools. Built-in suite of 20+ professional security tools including nmap, metasploit, sqlmap, and more.
  • Smart Memory System. Long-term storage of research results and successful approaches for future use.
  • Optional Knowledge Graph Integration. Graphiti-powered knowledge graph using Neo4j for semantic relationship tracking and advanced context understanding.
  • Web Intelligence. Built-in browser via scraper for gathering latest information from web sources.
  • External Search Systems. Integration with advanced search APIs including Tavily, Firecrawl, Traversaal, Perplexity, DuckDuckGo, Google Custom Search, Sploitus Search and Searxng for comprehensive information gathering.
  • Team of Specialists. Delegation system with specialized AI agents for research, development, and infrastructure tasks, enhanced with optional execution monitoring and intelligent task planning for optimal performance with smaller models.
  • Comprehensive Monitoring. Detailed logging and integration with Grafana/Prometheus for real-time system observation.
  • Detailed Reporting. Generation of thorough vulnerability reports with exploitation guides.
  • Smart Container Management. Automatic Docker image selection based on specific task requirements.
  • Modern Interface. Clean and intuitive web UI for system management and monitoring.
  • Comprehensive APIs. Full-featured REST and GraphQL APIs with Bearer token authentication for automation and integration.
  • Persistent Storage. All commands and outputs are stored in PostgreSQL with pgvector extension.
  • Scalable Architecture. Microservices-based design supporting horizontal scaling.
  • Self-Hosted Solution. Complete control over your deployment and data.
  • Flexible Authentication. Support for 10+ LLM providers (OpenAI, Anthropic, Google AI/Gemini, AWS Bedrock, Ollama, DeepSeek, GLM, Kimi, Qwen, MiniMax, Mistral, xAI, Custom for any OpenAI-compatible endpoint including Azure OpenAI) plus aggregators (OpenRouter, DeepInfra, Atlas Cloud, OpenCode Go plan). For production local deployments, see our vLLM + Qwen3.5-27B-FP8 guide.
  • API Token Authentication. Secure Bearer token system for programmatic access to REST and GraphQL APIs.
  • Quick Deployment. Easy setup through Docker Compose with comprehensive environment configuration.

Current Capability Boundaries

  • PentAGI today is an autonomous and assistant-guided penetration testing platform, not a CALDERA-style Breach and Attack Simulation (BAS) or adversary emulation product with predefined campaigns or attack plans.
  • BAS-like agent-authored attack scripts should be treated as conceptual or future work, not as a feature that is implemented today.
  • The current flow report UI supports web view, copy to clipboard, Markdown download, and PDF download. JSON flow-report export is not documented as a supported output format today.
  • Provider flexibility is available today through built-in providers and custom/OpenAI-compatible endpoints. See Custom LLM Provider Configuration and the vLLM + Qwen3.5-27B-FP8 guide.

Architecture

System Context

flowchart TB
    classDef person fill:#08427B,stroke:#073B6F,color:#fff
    classDef system fill:#1168BD,stroke:#0B4884,color:#fff
    classDef external fill:#666666,stroke:#0B4884,color:#fff

    pentester["👤 Security Engineer
    (User of the system)"]

    pentagi["✨ PentAGI
    (Autonomous penetration testing system)"]

    target["🎯 target-system
    (System under test)"]
    llm["🧠 llm-provider
    (OpenAI/Anthropic/Ollama/Bedrock/Gemini/Custom)"]
    search["🔍 search-systems
    (Google/DuckDuckGo/Tavily/Firecrawl/Traversaal/Perplexity/Sploitus/Searxng)"]
    langfuse["📊 langfuse-ui
    (LLM Observability Dashboard)"]
    grafana["📈 grafana
    (System Monitoring Dashboard)"]

    pentester --> |Uses HTTPS| pentagi
    pentester --> |Monitors AI HTTPS| langfuse
    pentester --> |Monitors System HTTPS| grafana
    pentagi --> |Tests Various protocols| target
    pentagi --> |Queries HTTPS| llm
    pentagi --> |Searches HTTPS| search
    pentagi --> |Reports HTTPS| langfuse
    pentagi --> |Reports HTTPS| grafana

    class pentester person
    class pentagi system
    class target,llm,search,langfuse,grafana external

    linkStyle default stroke:#ffffff,color:#ffffff

Container Architecture (click to expand)

graph TB
    subgraph Core Services
        UI[Frontend UI  
React + TypeScript]
        API[Backend API  
Go + GraphQL]
        DB[(Vector Store  
PostgreSQL + pgvector)]
        MQ[Task Queue  
Async Processing]
        Agent[AI Agents  
Multi-Agent System]
    end

    subgraph Knowledge Graph
        Graphiti[Graphiti  
Knowledge Graph API]
        Neo4j[(Neo4j  
Graph Database)]
    end

    subgraph Monitoring
        Grafana[Grafana  
Dashboards]
        VictoriaMetrics[VictoriaMetrics  
Time-series DB]
        Jaeger[Jaeger  
Distributed Tracing]
        Loki[Loki  
Log Aggregation]
        OTEL[OpenTelemetry  
Data Collection]
    end

    subgraph Analytics
        Langfuse[Langfuse  
LLM Analytics]
        ClickHouse[ClickHouse  
Analytics DB]
        Redis[Redis  
Cache + Rate Limiter]
        MinIO[MinIO  
S3 Storage]
    end

    subgraph Security Tools
        Scraper[Web Scraper  
Isolated Browser]
        PenTest[Security Tools  
20+ Pro Tools  
Sandboxed Execution]
    end

    UI --> |HTTP/WS| API
    API --> |SQL| DB
    API --> |Events| MQ
    MQ --> |Tasks| Agent
    Agent --> |Commands| PenTest
    Agent --> |Queries| DB
    Agent --> |Knowledge| Graphiti
    Graphiti --> |Graph| Neo4j

    API --> |Telemetry| OTEL
    OTEL --> |Metrics| VictoriaMetrics
    OTEL --> |Traces| Jaeger
    OTEL --> |Logs| Loki

    Grafana --> |Query| VictoriaMetrics
    Grafana --> |Query| Jaeger
    Grafana --> |Query| Loki

    API --> |Analytics| Langfuse
    Langfuse --> |Store| ClickHouse
    Langfuse --> |Cache| Redis
    Langfuse --> |Files| MinIO

    classDef core fill:#f9f,stroke:#333,stroke-width:2px,color:#000
    classDef knowledge fill:#ffa,stroke:#333,stroke-width:2px,color:#000
    classDef monitoring fill:#bbf,stroke:#333,stroke-width:2px,color:#000
    classDef analytics fill:#bfb,stroke:#333,stroke-width:2px,color:#000
    classDef tools fill:#fbb,stroke:#333,stroke-width:2px,color:#000

    class UI,API,DB,MQ,Agent core
    class Graphiti,Neo4j knowledge
    class Grafana,VictoriaMetrics,Jaeger,Loki,OTEL monitoring
    class Langfuse,ClickHouse,Redis,MinIO analytics
    class Scraper,PenTest tools

Entity Relationship (click to expand)

erDiagram
    Flow ||--o{ Task : contains
    Task ||--o{ SubTask : contains
    SubTask ||--o{ Action : contains
    Action ||--o{ Artifact : produces
    Action ||--o{ Memory : stores

    Flow {
        string id PK
        string name "Flow name"
        string description "Flow description"
        string status "active/completed/failed"
        json parameters "Flow parameters"
        timestamp created_at
        timestamp updated_at
    }

    Task {
        string id PK
        string flow_id FK
        string name "Task name"
        string description "Task description"
        string status "pending/running/done/failed"
        json result "Task results"
        timestamp created_at
        timestamp updated_at
    }

    SubTask {
        string id PK
        string task_id FK
        string name "Subtask name"
        string description "Subtask description"
        string status "queued/running/completed/failed"
        string agent_type "researcher/developer/executor"
        json context "Agent context"
        timestamp created_at
        timestamp updated_at
    }

    Action {
        string id PK
        string subtask_id FK
        string type "command/search/analyze/etc"
        string status "success/failure"
        json parameters "Action parameters"
        json result "Action results"
        timestamp created_at
    }

    Artifact {
        string id PK
        string action_id FK
        string type "file/report/log"
        string path "Storage path"
        json metadata "Additional info"
        timestamp created_at
    }

    Memory {
        string id PK
        string action_id FK
        string type "observation/conclusion"
        vector embedding "Vector representation"
        text content "Memory content"
        timestamp created_at
    }

Agent Interaction (click to expand)

sequenceDiagram
    participant O as Orchestrator
    participant R as Researcher
    participant D as Developer
    participant E as Executor
    participant VS as Vector Store
    participant KB as Knowledge Base

    Note over O,KB: Flow Initialization
    O->>VS: Query similar tasks
    VS-->>O: Return experiences
    O->>KB: Load relevant knowledge
    KB-->>O: Return context

    Note over O,R: Research Phase
    O->>R: Analyze target
    R->>VS: Search similar cases
    VS-->>R: Return patterns
    R->>KB: Query vulnerabilities
    KB-->>R: Return known issues
    R->>VS: Store findings
    R-->>O: Research results

    Note over O,D: Planning Phase
    O->>D: Plan attack
    D->>VS: Query exploits
    VS-->>D: Return techniques
    D->>KB: Load tools info
    KB-->>D: Return capabilities
    D-->>O: Attack plan

    Note over O,E: Execution Phase
    O->>E: Execute plan
    E->>KB: Load tool guides
    KB-->>E: Return procedures
    E->>VS: Store results
    E-->>O: Execution status

Memory System (click to expand)

graph TB
    subgraph "Long-term Memory"
        VS[(Vector Store  
Embeddings DB)]
        KB[Knowledge Base  
Domain Expertise]
        Tools[Tools Knowledge  
Usage Patterns]
    end

    subgraph "Working Memory"
        Context[Current Context  
Task State]
        Goals[Active Goals  
Objectives]
        State[System State  
Resources]
    end

    subgraph "Episodic Memory"
        Actions[Past Actions  
Commands History]
        Results[Action Results  
Outcomes]
        Patterns[Success Patterns  
Best Practices]
    end

    Context --> |Query| VS
    VS --> |Retrieve| Context

    Goals --> |Consult| KB
    KB --> |Guide| Goals

    State --> |Record| Actions
    Actions --> |Learn| Patterns
    Patterns --> |Store| VS

    Tools --> |Inform| State
    Results --> |Update| Tools

    VS --> |Enhance| KB
    KB --> |Index| VS

    classDef ltm fill:#f9f,stroke:#333,stroke-width:2px,color:#000
    classDef wm fill:#bbf,stroke:#333,stroke-width:2px,color:#000
    classDef em fill:#bfb,stroke:#333,stroke-width:2px,color:#000

    class VS,KB,Tools ltm
    class Context,Goals,State wm
    class Actions,Results,Patterns em

Chain Summarization (click to expand)

The chain summarization system manages conversation context growth by selectively summarizing older messages. This is critical for preventing token limits from being exceeded while maintaining conversation coherence.

flowchart TD
    A[Input Chain] --> B{Needs Summarization?}
    B -->|No| C[Return Original Chain]
    B -->|Yes| D[Convert to ChainAST]
    D --> E[Apply Section Summarization]
    E --> F[Process Oversized Pairs]
    F --> G[Manage Last Section Size]
    G --> H[Apply QA Summarization]
    H --> I[Rebuild Chain with Summaries]
    I --> J{Is New Chain Smaller?}
    J -->|Yes| K[Return Optimized Chain]
    J -->|No| C

    classDef process fill:#bbf,stroke:#333,stroke-width:2px,color:#000
    classDef decision fill:#bfb,stroke:#333,stroke-width:2px,color:#000
    classDef output fill:#fbb,stroke:#333,stroke-width:2px,color:#000

    class A,D,E,F,G,H,I process
    class B,J decision
    class C,K output

The algorithm operates on a structured representation of conversation chains (ChainAST) that preserves message types including tool calls and their responses. All summarization operations maintain critical conversation flow while reducing context size.

Global Summarizer Configuration Options

Parameter Environment Variable Default Description
Preserve Last SUMMARIZER_PRESERVE_LAST true Whether to keep all messages in the last section intact
Use QA Pairs SUMMARIZER_USE_QA true Whether to use QA pair summarization strategy
Summarize Human in QA SUMMARIZER_SUM_MSG_HUMAN_IN_QA false Whether to summarize human messages in QA pairs
Last Section Size SUMMARIZER_LAST_SEC_BYTES 51200 Maximum byte size for last section (50KB)
Max Body Pair Size SUMMARIZER_MAX_BP_BYTES 16384 Maximum byte size for a single body pair (16KB)
Max QA Sections SUMMARIZER_MAX_QA_SECTIONS 10 Maximum QA pair sections to preserve
Max QA Size SUMMARIZER_MAX_QA_BYTES 65536 Maximum byte size for QA pair sections (64KB)
Keep QA Sections SUMMARIZER_KEEP_QA_SECTIONS 1 Number of recent QA sections to keep without summarization

Assistant Summarizer Configuration Options

Assistant instances can use customized summarization settings to fine-tune context management behavior:

Parameter Environment Variable Default Description
Preserve Last ASSISTANT_SUMMARIZER_PRESERVE_LAST true Whether to preserve all messages in the assistant's last section
Last Section Size ASSISTANT_SUMMARIZER_LAST_SEC_BYTES 76800 Maximum byte size for assistant's last section (75KB)
Max Body Pair Size ASSISTANT_SUMMARIZER_MAX_BP_BYTES 16384 Maximum byte size for a single body pair in assistant context (16KB)
Max QA Sections ASSISTANT_SUMMARIZER_MAX_QA_SECTIONS 7 Maximum QA sections to preserve in assistant context
Max QA Size ASSISTANT_SUMMARIZER_MAX_QA_BYTES 76800 Maximum byte size for assistant's QA sections (75KB)
Keep QA Sections ASSISTANT_SUMMARIZER_KEEP_QA_SECTIONS 3 Number of recent QA sections to preserve without summarization

The assistant summarizer configuration provides more memory for context retention compared to the global settings, preserving more recent conversation history while still ensuring efficient token usage.

Summarizer Environment Configuration

# Default values for global summarizer logic
SUMMARIZER_PRESERVE_LAST=true
SUMMARIZER_USE_QA=true
SUMMARIZER_SUM_MSG_HUMAN_IN_QA=false
SUMMARIZER_LAST_SEC_BYTES=51200
SUMMARIZER_MAX_BP_BYTES=16384
SUMMARIZER_MAX_QA_SECTIONS=10
SUMMARIZER_MAX_QA_BYTES=65536
SUMMARIZER_KEEP_QA_SECTIONS=1

# Default values for assistant summarizer logic
ASSISTANT_SUMMARIZER_PRESERVE_LAST=true
ASSISTANT_SUMMARIZER_LAST_SEC_BYTES=76800
ASSISTANT_SUMMARIZER_MAX_BP_BYTES=16384
ASSISTANT_SUMMARIZER_MAX_QA_SECTIONS=7
ASSISTANT_SUMMARIZER_MAX_QA_BYTES=76800
ASSISTANT_SUMMARIZER_KEEP_QA_SECTIONS=3

Advanced Agent Supervision (click to expand)

PentAGI includes sophisticated multi-layered agent supervision mechanisms to ensure efficient task execution, prevent infinite loops, and provide intelligent recovery from stuck states:

Execution Monitoring (Beta)

  • Automatic Mentor Intervention: Adviser agent (mentor) is automatically invoked when execution patterns indicate potential issues
  • Pattern Detection: Monitors identical tool calls (threshold: 5, configurable) and total tool calls (threshold: 10, configurable)
  • Progress Analysis: Evaluates whether agent advances toward subtask objective, detects loops and inefficiencies
  • Alternative Strategies: Recommends different approaches when current strategy fails
  • Information Retrieval Guidance: Suggests searching for established solutions instead of reinventing
  • Enhanced Response Format: Tool responses include both and sections
  • Configurable: Enable via EXECUTION_MONITOR_ENABLED (default: false), customize thresholds with EXECUTION_MONITOR_SAME_TOOL_LIMIT and EXECUTION_MONITOR_TOTAL_TOOL_LIMIT

Best for: Smaller models (< 32B parameters), complex attack scenarios requiring continuous guidance, preventing agents from getting stuck on single approach

Performance Impact: 2-3x increase in execution time and token usage, but delivers 2x improvement in result quality based on testing with Qwen3.5-27B-FP8

Intelligent Task Planning (Beta)

  • Automated Decomposition: Planner (adviser in planning mode) generates 3-7 specific, actionable steps before specialist agents begin work
  • Context-Aware Plans: Analyzes full execution context via enricher agent to create informed plans
  • Structured Assignment: Original request wrapped in `` structure with execution plan and instructions
  • Scope Management: Prevents scope creep by keeping agents focused on current subtask only
  • Enriched Instructions: Plans highlight critical actions, potential pitfalls, and verification points
  • Configurable: Enable via AGENT_PLANNING_STEP_ENABLED (default: false)

Best for: Models < 32B parameters, complex penetration testing workflows, improving success rates on sophisticated tasks

Enhanced Adviser Configuration: Works exceptionally well when adviser agent uses stronger model or enhanced settings. Example: using same base model with maximum reasoning mode for adviser (see vllm-qwen3.5-27b-fp8.provider.yml) enables comprehensive task analysis and strategic planning from identical model architecture.

Performance Impact: Adds planning overhead but significantly improves completion rates and reduces redundant work

Tool Call Limits (Always Active)

  • Hard Limits: Prevent runaway executions regardless of supervision mode status
  • Differentiated by Agent Type:
    • General agents (Assistant, Primary Agent, Pentester, Coder, Installer): MAX_GENERAL_AGENT_TOOL_CALLS (default: 100)
    • Limited agents (Searcher, Enricher, Memorist, Generator, Reporter, Adviser, Reflector, Planner): MAX_LIMITED_AGENT_TOOL_CALLS (default: 20)
  • Graceful Termination: Reflector guides agents to proper completion when approaching limits
  • Resource Protection: Ensures system stability and prevents resource exhaustion

Reflector Integration (Always Active)

  • Automatic Correction: Invoked when LLM fails to generate tool calls after 3 attempts
  • Strategic Guidance: Analyzes failures and guides agents toward proper tool usage or barrier tools (done, ask)
  • Recovery Mechanism: Provides contextual guidance based on specific failure patterns
  • Limit Enforcement: Coordinates graceful termination when tool call limits are reached

Recommendations for Open Source Models

Must-Have for Models < 32B Parameters: Testing with Qwen3.5-27B-FP8 demonstrates that enabling both Execution Monitoring and Task Planning is essential for smaller open source models:

  • Quality Improvement: 2x better results compared to baseline execution without supervision
  • Loop Prevention: Significantly reduces infinite loops and redundant work
  • Attack Diversity: Encourages exploration of multiple attack vectors instead of fixating on single approach
  • Air-Gapped Deployments: Enables production-grade autonomous pentesting in closed network environments with local LLM inference

Trade-offs:

  • Token consumption: 2-3x increase due to mentor/planner invocations
  • Execution time: 2-3x longer due to analysis and planning steps
  • Result quality: 2x improvement in completeness, accuracy, and attack coverage
  • Model requirements: Works best when adviser uses enhanced configuration (higher reasoning parameters, stronger model variant, or different model)

Configuration Strategy: For optimal performance with smaller models, configure adviser agent with enhanced settings:

  • Use same model with maximum reasoning mode (example: vllm-qwen3.5-27b-fp8.provider.yml)
  • Or use stronger model for adviser while keeping base model for other agents
  • Adjust monitoring thresholds based on task complexity and model capabilities

The architecture of PentAGI is designed to be modular, scalable, and secure. Here are the key components:

  1. Core Services

    • Frontend UI: React-based web interface with TypeScript for type safety
    • Backend API: Go-based REST and GraphQL APIs with Bearer token authentication for programmatic access
    • Vector Store: PostgreSQL with pgvector for semantic search and memory storage
    • Task Queue: Async task processing system for reliable operation
    • AI Agent: Multi-agent system with specialized roles for efficient testing
  2. Optional Knowledge Graph

    • Graphiti: Knowledge graph API for semantic relationship tracking and contextual understanding
    • Neo4j: Graph database for storing and querying relationships between entities, actions, and outcomes
    • When enabled, automatically captures agent responses and tool executions for a flow-scoped knowledge base
  3. Monitoring Stack

    • OpenTelemetry: Unified observability data collection and correlation
    • Grafana: Real-time visualization and alerting dashboards
    • VictoriaMetrics: High-performance time-series metrics storage
    • Jaeger: End-to-end distributed tracing for debugging
    • Loki: Scalable log aggregation and analysis
  4. Analytics Platform

    • Langfuse: Advanced LLM observability and performance analytics
    • ClickHouse: Column-oriented analytics data warehouse
    • Redis: High-speed caching and rate limiting
    • MinIO: S3-compatible object storage for artifacts
  5. Security Tools

    • Web Scraper: Isolated browser environment for safe web interaction
    • Pentesting Tools: Comprehensive suite of 20+ professional security tools
    • Sandboxed Execution: All operations run in isolated containers
  6. Memory Systems

    • Long-term Memory: Persistent storage of knowledge and experiences
    • Working Memory: Active context and goals for current operations
    • Episodic Memory: Historical actions and success patterns
    • Knowledge Base: Structured domain expertise and tool capabilities
    • Context Management: Intelligently manages growing LLM context windows using chain summarization

The system uses Docker containers for isolation and easy deployment, with separate networks for core services, monitoring, and analytics to ensure proper security boundaries. Each component is designed to scale horizontally and can be configured for high availability in production environments.

Quick Start

For a step-by-step walkthrough that connects installation, configuration, LLM and embedding provider testing, and your first login, see the Installing and Configuring PentAGI guide. The sections below remain the detailed reference for each step.

System Requirements

  • Docker and Docker Compose (or Podman - see Podman configuration)
  • Minimum 2 vCPU
  • Minimum 4GB RAM
  • 20GB free disk space
  • Internet access for downloading images and updates

Using Installer (Recommended)

PentAGI provides an interactive installer with a terminal-based UI for streamlined configuration and deployment. The installer guides you through system checks, LLM provider setup, search engine configuration, and security hardening.

Supported Platforms:

macOS security warning: If macOS flags a downloaded installer, use only the official PentAGI links above, choose the archive that matches your CPU architecture, verify the source before continuing, and follow the installer troubleshooting guide before allowing the app to run.

Quick Installation (Linux amd64):

# Create installation directory
mkdir -p pentagi && cd pentagi

# Download installer
wget -O installer.zip https://pentagi.com/downloads/linux/amd64/installer-latest.zip

# Extract
unzip installer.zip

# Run interactive installer
./installer

Prerequisites & Permissions:

The installer requires appropriate privileges to interact with the Docker API for proper operation. By default, it uses the Docker socket (/var/run/docker.sock) which requires either:

  • Option 1 (Recommended for production): Run the installer as root:

    sudo ./installer
    
  • Option 2 (Development environments): Grant your user access to the Docker socket by adding them to the docker group:

    # Add your user to the docker group
    sudo usermod -aG docker $USER
    
    # Log out and log back in, or activate the group immediately
    newgrp docker
    
    # Verify Docker access (should run without sudo)
    docker ps
    

    ⚠️ Security Note: Adding a user to the docker group grants root-equivalent privileges. Only do this for trusted users in controlled environments. For production deployments, consider using rootless Docker mode or running the installer with sudo.

The installer will:

  1. System Checks: Verify Docker, network connectivity, and system requirements
  2. Environment Setup: Create and configure .env file with optimal defaults
  3. Provider Configuration: Set up LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Ollama, DeepSeek, GLM, Kimi, Qwen, MiniMax, Mistral, xAI, Custom)
  4. Search Engines: Configure DuckDuckGo, Google, Tavily, Firecrawl, Traversaal, Perplexity, Sploitus, Searxng, and the optional internal browser-analytics fallback engine
  5. Security Hardening: Generate secure credentials and configure SSL certificates
  6. Deployment: Start PentAGI with docker-compose

Current Web Settings Coverage

The PentAGI web console already manages several settings areas after the server is up and running:

  • Settings -> Providers: Create, edit, delete, and test user-defined provider profiles for supported provider types. These profiles control per-agent model selection, runtime parameters, reasoning options, and pricing metadata.
  • Settings -> Prompts: Manage system, human, and tool prompt templates.
  • Settings -> PentAGI API: Create and manage PentAGI Bearer tokens for REST and GraphQL access.
  • Other UI-managed preferences: Favorite flows are stored as user preferences, and theme selection is handled from the main sidebar/profile controls rather than the Settings pages.

Still Server-Managed

The following configuration areas still need to be set on the server through environment variables, compose files, or mounted config files:

  • LLM credentials and connection details: API keys, endpoints, auth modes, and provider-specific connection settings for OpenAI, Anthropic, Bedrock, Ollama, custom providers, and similar backends; config-path settings apply only where supported, such as OLLAMA_SERVER_CONFIG_PATH, LLM_SERVER_CONFIG_PATH, and BEDROCK_CONFIG_PATH.
  • Search provider credentials and options: Settings such as DUCKDUCKGO_*, GOOGLE_*, TAVILY_API_KEY, FIRECRAWL_API_*, TRAVERSAAL_API_KEY, PERPLEXITY_*, SEARXNG_*, SPLOITUS_ENABLED, and the optional WEB_SEARCH_INTERNAL_* browser-analytics fallback settings.
  • Third-party integrations: Langfuse, Graphiti, and similar external services remain server-side configuration.
  • MCP server management: MCP settings pages are not currently exposed as a live web-console feature.

For Production & Enhanced Security:

For production deployments or security-sensitive environments, we strongly recommend using a distributed two-node architecture where worker operations are isolated on a separate server. This prevents untrusted code execution and network access issues on your main system.

See detailed guide: Worker Node Setup

The two-node setup provides:

  • Isolated Execution: Worker containers run on dedicated hardware
  • Network Isolation: Separate network boundaries for penetration testing
  • Security Boundaries: Docker-in-Docker with TLS authentication
  • OOB Attack Support: Dedicated port ranges for out-of-band techniques

Giving Agents Docker Without Giving Away the Host

Many pentest workflows need docker inside the agent's sandbox. There are two ways to provide it, and they differ sharply in risk.

Recommended — point sandboxes at a hardened dind daemon over TLS. Set DOCKER_INSIDE=true, leave DOCKER_SOCKET empty, and configure the daemon the sandbox may talk to:

DOCKER_INSIDE=true
DOCKER_SOCKET=                                          # mount no socket
DOCKER_INSIDE_HOST=tcp://10.0.0.5:3376                  # hardened dind endpoint
DOCKER_INSIDE_TLS_VERIFY=1
DOCKER_INSIDE_CERT_PATH=/etc/docker/dind/certs/client   # path on the worker node
DOCKER_INSIDE_POLICY_TESTS=true                         # prove it on every worker; off by default

PentAGI injects these into every worker container as DOCKER_HOST, DOCKER_TLS_VERIFY and DOCKER_CERT_PATH (the _INSIDE_ segment is dropped) and bind-mounts the certificate directory read-only at the same path, so docker works inside the sandbox with no further setup.

Not recommended — bind-mounting a Docker socket (DOCKER_SOCKET). This has two failure modes:

  • Boot-order race: a bind-mount source that does not exist yet is created by Docker as a directory. After a worker-node reboot, a worker container can start before dind has recreated its socket — Docker then puts a directory where the socket belongs, and dind cannot start until it is removed by hand.
  • Blast radius: the race is only reliably avoided when the mounted socket is the host daemon's, since that one always exists first. But that grants an autonomous agent the host Docker API: it can start a privileged container, mount /, and compromise the entire node — PentAGI included.

Use DOCKER_SOCKET only on single-node development setups where the host daemon is already trusted.

See: Worker Node Setup for the full dind hardening and TLS configuration, and Worker Docker Access for the exact resolution algorithm.

Running Several Instances (TENANT_ID)

A single PentAGI installation needs none of this — leave TENANT_ID empty (the default) and nothing changes.

Set it when several PentAGI installations share external resources: one PostgreSQL server, one worker node, one Neo4j/Graphiti, one Langfuse. The typical case is a management backend per server with a common worker node and database. Because every instance numbers its flows from 1, they would otherwise collide on container names, database rows, knowledge-graph namespaces and session cookies. TENANT_ID namespaces all of it:

Area Effect when TENANT_ID=acme
PostgreSQL The instance creates and works inside schema acme instead of public; extensions stay shared in DATABASE_EXTENSIONS_SCHEMA (default public, extensions on Supabase)
Worker containers acme-pentagi-terminal- instead of pentagi-terminal-; volumes and hostnames follow, and both carry a pentagi.tenant label
Knowledge graph Graphiti/Neo4j group ids become acme-flow-
Auth Cookie and API token keys are derived from COOKIE_SIGNING_SALT plus the tenant, and the session cookie is renamed
Telemetry Langfuse traces carry the tenant as their environment and a tenant:acme tag; OTel resources gain tenant_id

The value must match ^[a-z][a-z0-9_]{0,31}$ — an invalid one aborts startup rather than being silently normalised.

Some things stay yours to set per instance, because they are host resources rather than names: DATA_DIR (two instances sharing it will overwrite each other's flow data), DOCKER_PORTS_BASE, the published ports, and INSTALLATION_ID. The effective values are printed at startup under Instance identity.

The installer provisions one instance per server. Running several on one server is possible — for example behind a shared nginx — but the stock docker-compose.yml uses fixed container and network names, so it has to be adapted to your own network layout first.

See: Multi-Instance Deployment for validation rules, upgrade notes and the full list of operator responsibilities.

Manual Installation

  1. Create a working directory or clone the repository:
mkdir pentagi && cd pentagi
  1. Copy .env.example to .env or download it:
curl -o .env https://raw.githubusercontent.com/vxcontrol/pentagi/master/.env.example
  1. Touch examples files (example.custom.provider.yml, example.ollama.provider.yml, example.bedrock.provider.yml) or download it. docker-compose.yml mounts each of them into the container, and Docker creates a directory in place of a mount source that does not exist:
curl -o example.custom.provider.yml https://raw.githubusercontent.com/vxcontrol/pentagi/master/examples/configs/custom-openai.provider.yml
curl -o example.ollama.provider.yml https://raw.githubusercontent.com/vxcontrol/pentagi/master/examples/configs/ollama-llama318b.provider.yml
curl -o example.bedrock.provider.yml https://raw.githubusercontent.com/vxcontrol/pentagi/master/examples/configs/bedrock.provider.yml
  1. Fill in the required API keys in .env file.
# Required: At least one of these LLM providers
OPEN_AI_KEY=your_openai_key
ANTHROPIC_API_KEY=your_anthropic_key
GEMINI_API_KEY=your_gemini_key

# Optional: AWS Bedrock provider (enterprise-grade models)
BEDROCK_REGION=us-east-1
# Choose one authentication method:
BEDROCK_DEFAULT_AUTH=true                        # Option 1: Use AWS SDK default credential chain (recommended for EC2/ECS)
# BEDROCK_BEARER_TOKEN=your_bearer_token         # Option 2: Bearer token authentication
# BEDROCK_ACCESS_KEY_ID=your_aws_access_key      # Option 3: Static credentials
# BEDROCK_SECRET_ACCESS_KEY=your_aws_secret_key

# Optional: Ollama provider (local or cloud)
# OLLAMA_SERVER_URL=http://ollama-server:11434   # Local server
# OLLAMA_SERVER_URL=https://ollama.com           # Cloud service
# OLLAMA_SERVER_API_KEY=your_ollama_cloud_key    # Required for cloud, empty for local

# Optional: Chinese AI providers
# DEEPSEEK_API_KEY=your_deepseek_key             # DeepSeek (strong reasoning)
# GLM_API_KEY=your_glm_key                       # GLM (Zhipu AI)
# KIMI_API_KEY=your_kimi_key                     # Kimi (Moonshot AI, ultra-long context)
# QWEN_API_KEY=your_qwen_key                     # Qwen (Alibaba Cloud, multimodal)
# MINIMAX_API_KEY=your_minimax_key               # MiniMax

# Optional: European and US providers
# MISTRAL_API_KEY=your_mistral_key               # Mistral
# XAI_API_KEY=your_xai_key                       # xAI (Grok)

# Optional: Local LLM provider (zero-cost inference)
OLLAMA_SERVER_URL=http://localhost:11434
OLLAMA_SERVER_MODEL=your_model_name

# Optional: Additional search capabilities
DUCKDUCKGO_ENABLED=true
DUCKDUCKGO_REGION=us-en
DUCKDUCKGO_SAFESEARCH=
DUCKDUCKGO_TIME_RANGE=
SPLOITUS_ENABLED=true
GOOGLE_API_KEY=your_google_key
GOOGLE_CX_KEY=your_google_cx
TAVILY_API_KEY=your_tavily_key
FIRECRAWL_API_KEY=your_firecrawl_key
FIRECRAWL_API_URL=
TRAVERSAAL_API_KEY=your_traversaal_key
PERPLEXITY_API_KEY=your_perplexity_key
PERPLEXITY_MODEL=
PERPLEXITY_CONTEXT_SIZE=medium

# Searxng meta search engine (aggregates results from multiple sources)
SEARXNG_URL=http://your-searxng-instance:8080
SEARXNG_CATEGORIES=general
SEARXNG_LANGUAGE=
SEARXNG_SAFESEARCH=0
SEARXNG_TIME_RANGE=
SEARXNG_TIMEOUT=

# Optional: internal browser-analytics fallback engine for web_search (off by default;
# scrapes and summarizes pages instead of calling a paid analytic API)
WEB_SEARCH_INTERNAL_ENABLED=false
WEB_SEARCH_INTERNAL_MAX_SITES=5
WEB_SEARCH_INTERNAL_MAX_SITE_BYTES=10240

## Graphiti knowledge graph settings
GRAPHITI_ENABLED=false
GRAPHITI_TIMEOUT=30
GRAPHITI_URL=

# Neo4j settings (used by Graphiti stack)
NEO4J_USER=neo4j
NEO4J_DATABASE=neo4j
NEO4J_PASSWORD=devpassword
NEO4J_URI=bolt://neo4j:7687

# Assistant configuration
ASSISTANT_USE_AGENTS=false         # Default value for agent usage when creating new assistants
  1. Change all security related environment variables in .env file to improve security.

Security related environment variables

Main Security Settings

  • COOKIE_SIGNING_SALT - Salt for cookie signing, change to random value
  • PUBLIC_URL - Public URL of your server (eg. https://pentagi.example.com)
  • SERVER_SSL_CRT and SERVER_SSL_KEY - Custom paths to your existing SSL certificate and key for HTTPS (these paths should be used in the docker-compose.yml file to mount as volumes)
  • TENANT_ID - Leave empty unless this instance shares external resources with another PentAGI installation. When set, it is mixed into the cookie and API token signing keys and renames the session cookie, so a session minted by one instance is rejected by the others even though they share the same COOKIE_SIGNING_SALT. See Running Several Instances

Scraper Access

  • SCRAPER_PUBLIC_URL - Public URL for scraper if you want to use different scraper server for public URLs
  • SCRAPER_PRIVATE_URL - Private URL for scraper (local scraper server in docker-compose.yml file to access it to local URLs)

Access Credentials

  • PENTAGI_POSTGRES_USER and PENTAGI_POSTGRES_PASSWORD - PostgreSQL credentials
  • NEO4J_USER and NEO4J_PASSWORD - Neo4j credentials (for Graphiti knowledge graph)
  1. Remove all inline comments from .env file if you want to use it in VSCode or other IDEs as a envFile option:
perl -i -pe 's/\s+#.*$//' .env
  1. Run the PentAGI stack:
curl -O https://raw.githubusercontent.com/vxcontrol/pentagi/master/docker-compose.yml
docker compose up -d

Visit localhost:8443 to access PentAGI Web UI (default is [email protected] / admin)

Web UI Accounts

PentAGI does not expose public self-service sign-up from the login page. A fresh installation creates the default local administrator account:

On first login, change the default password before using the instance for real work. If the administrator password is lost later, use the installer maintenance menu to reset the default [email protected] account password.

For multi-user setups, an authenticated administrator can manage local users through the Users REST API (/api/v1/users/). The OpenAPI UI is available at https://localhost:8443/api/v1/swagger/index.html after the instance is running.

[!NOTE] If you caught an error about pentagi-network or observability-network or langfuse-network you need to run docker-compose.yml firstly to create these networks and after that run docker-compose-langfuse.yml, docker-compose-graphiti.yml, and docker-compose-observability.yml to use Langfuse, Graphiti, and Observability services.

You have to set at least one Language Model provider (OpenAI, Anthropic, Gemini, AWS Bedrock, or Ollama) to use PentAGI. AWS Bedrock provides enterprise-grade access to multiple foundation models from leading AI companies, while Ollama provides zero-cost local inference if you have sufficient computational resources. Additional API keys for search engines are optional but recommended for better results.

For fully local deployment with advanced models: See our comprehensive guide on Running PentAGI with vLLM and Qwen3.5-27B-FP8 for a production-grade local LLM setup. This configuration achieves ~13,000 TPS for prompt processing and ~650 TPS for completion on 4× RTX 5090 GPUs, supporting 12+ concurrent flows with complete independence from cloud providers.

LLM_SERVER_* environment variables are experimental feature and will be changed in the future. Right now you can use them to specify custom LLM server URL and one model for all agent types.

PROXY_URL routes the backend's own HTTP calls through a proxy: every request to LLM and embedding providers and to search engines — self-hosted ones such as Ollama or SearXNG included, as there is no NO_PROXY exemption — and the update check. Sandbox containers (where agents run their commands), the scraper, the Graphiti service, OAuth sign-in, and Langfuse/OpenTelemetry export do not use it, so if you need isolation from external networks, restrict their outbound traffic at the network level.

The docker-compose.yml file runs the PentAGI service as root user because it needs access to docker.sock for container management. If you're using TCP/IP network connection to Docker instead of socket file, you can remove root privileges and use the default pentagi user for better security.

Accessing PentAGI from External Networks

By default, the PentAGI web interface and the other compose services bind to 127.0.0.1 (localhost only) for security. To access PentAGI from other machines on your network, you need to configure external access.

[!IMPORTANT] Flow sandboxes are not bound to localhost. Each flow publishes two TCP ports for out-of-band callbacks such as reverse shells, taken from the 2000-port window that starts at DOCKER_PORTS_BASE (28000–29999 by default), on DOCKER_PUBLIC_IP — 0.0.0.0 by default, that is every interface of the Docker host. Docker routes published ports around ufw rules and lets them through firewalld with its own docker zone, so restrict this range with a firewall in front of the host or, with Docker's default iptables backend, with rules in the DOCKER-USER chain. With DOCKER_NETWORK=host, sandboxes share the host's network stack, so their listeners are exposed without any port mapping.

Configuration Steps

  1. Update .env file with your server's IP address:
# Network binding - allow external connections
PENTAGI_LISTEN_IP=0.0.0.0
PENTAGI_LISTEN_PORT=8443

# Public URL - use your actual server IP or hostname
# Replace 192.168.1.100 with your server's IP address
PUBLIC_URL=https://192.168.1.100:8443

# CORS origins - list all URLs that will access PentAGI
# Include localhost for local access AND your server IP for external access
CORS_ORIGINS=https://localhost:8443,https://192.168.1.100:8443

[!IMPORTANT]

  • Replace 192.168.1.100 with your actual server's IP address
  • Do NOT use 0.0.0.0 in PUBLIC_URL or CORS_ORIGINS - use the actual IP address
  • Include both localhost and your server IP in CORS_ORIGINS for flexibility
  1. Recreate containers to apply the changes:
docker compose down
docker compose up -d --force-recreate
  1. Verify port binding:
docker ps | grep pentagi

You should see 0.0.0.0:8443->8443/tcp or :::8443->8443/tcp.

If you see 127.0.0.1:8443->8443/tcp, the environment variable wasn't picked up. In this case, directly edit docker-compose.yml line 31:

ports:
  - "0.0.0.0:8443:8443"

Then recreate containers again.

  1. Configure firewall to allow incoming connections on port 8443:
# Ubuntu/Debian with UFW
sudo ufw allow 8443/tcp
sudo ufw reload

# CentOS/RHEL with firewalld
sudo firewall-cmd --permanent --add-port=8443/tcp
sudo firewall-cmd --reload
  1. Access PentAGI:
  • Local access: https://localhost:8443
  • Network access: https://your-server-ip:8443

[!NOTE] You'll need to accept the self-signed SSL certificate warning in your browser when accessing via IP address.


Running PentAGI with Podman

PentAGI fully supports Podman as a Docker alternative. However, when using Podman in rootless mode, the scraper service requires special configuration because rootless containers cannot bind privileged ports (ports below 1024).

Podman Rootless Configuration

The default scraper configuration uses port 443 (HTTPS), which is a privileged port. For Podman rootless, reconfigure the scraper to use a non-privileged port:

1. Edit docker-compose.yml - modify the scraper service (around line 199):

scraper:
  image: vxcontrol/scraper:latest
  restart: unless-stopped
  container_name: scraper
  hostname: scraper
  expose:
    - 3000/tcp  # Changed from 443 to 3000
  ports:
    - "${SCRAPER_LISTEN_IP:-127.0.0.1}:${SCRAPER_LISTEN_PORT:-9443}:3000"  # Map to port 3000
  environment:
    - MAX_CONCURRENT_SESSIONS=${LOCAL_SCRAPER_MAX_CONCURRENT_SESSIONS:-10}
    - USERNAME=${LOCAL_SCRAPER_USERNAME:-someuser}
    - PASSWORD=${LOCAL_SCRAPER_PASSWORD:-somepass}
  logging:
    options:
      max-size: 50m
      max-file: "7"
  volumes:
    - scraper-ssl:/usr/src/app/ssl
  networks:
    - pentagi-network
  shm_size: 2g

2. Update .env file - change the scraper URL to use HTTP and port 3000:

# Scraper configuration for Podman rootless
SCRAPER_PRIVATE_URL=http://someuser:somepass@scraper:3000/
LOCAL_SCRAPER_USERNAME=someuser
LOCAL_SCRAPER_PASSWORD=somepass

[!IMPORTANT] Key changes for Podman:

  • Use HTTP instead of HTTPS for SCRAPER_PRIVATE_URL
  • Use port 3000 instead of 443
  • Change internal expose to 3000/tcp
  • Update port mapping to target 3000 instead of 443

3. Recreate containers:

podman-compose down
podman-compose up -d --force-recreate

4. Test scraper connectivity:

# Test from within the pentagi container
podman exec -it pentagi wget -O- "http://someuser:somepass@scraper:3000/html?url=http://example.com"

If you see HTML output, the scraper is working correctly.

Podman Rootful Mode

If you're running Podman in rootful mode (with sudo), you can use the default configuration without modifications. The scraper will work on port 443 as intended.

Docker Compatibility

All Podman configurations remain fully compatible with Docker. The non-privileged port approach works identically on both container runtimes.

Assistant Configuration

PentAGI allows you to configure default behavior for assistants:

Variable Default Description
ASSISTANT_USE_AGENTS false Controls the default value for agent usage when creating new assistants

The ASSISTANT_USE_AGENTS setting affects the initial state of the "Use Agents" toggle when creating a new assistant in the UI:

  • false (default): New assistants are created with agent delegation disabled by default
  • true: New assistants are created with agent delegation enabled by default

Note that users can always override this setting by toggling the "Use Agents" button in the UI when creating or editing an assistant. This environment variable only controls the initial default state.

How to Use PentAGI After Login

Once the stack is running and you can sign in to the web UI, the fastest way to start is through the Flows workflow.

1. Create your first flow

  1. Open Flows in the sidebar.
  2. Click New Flow.
  3. Choose the mode that fits your goal:
    • Automation: fully autonomous execution for a testing goal you want PentAGI to carry out end-to-end
    • Assistant: interactive back-and-forth help when you want to steer the investigation step by step. In this mode you can also enable the Use Agents toggle to let PentAGI delegate subtasks to specialized sub-agents for more complex investigations.
  4. Select the LLM provider you want to use for this flow.
  5. Describe the target and the objective in natural language in the message box.

Good first prompts usually include:

  • the target system or URL
  • the type of assessment you want
  • any scope limitations or rules of engagement
  • the result you expect, such as a vulnerability report or validation of a hypothesis

Example:

Assess https://target.example for common web application vulnerabilities. Focus on authentication, file handling, and injection issues. Stay within the provided target only and summarize confirmed findings with reproduction steps.

Only test systems you own or are explicitly authorized to assess. See EULA.md for the acceptable use requirements.

2. Use templates for repeatable workflows

The new flow form includes a template picker, which can prefill the message box with a saved flow template. This is useful when you run similar assessments repeatedly.

  • Use an existing template if you already have one saved in Templates
  • Start from the example prompt in examples/prompts/base_web_pentest.md if you need a practical baseline for web testing
  • Adjust the target, scope, and constraints before starting the flow

Templates are starting points. You do not need special syntax to use PentAGI: plain natural-language instructions work well as long as the target and goal are clear.

3. Monitor execution and review output

After submitting the flow, PentAGI opens the flow page automatically.

  • Use the main flow view to follow messages, agent activity, and task progress
  • Inspect tool activity and terminal output as the flow runs
  • Review generated tasks and subtasks to understand what PentAGI is doing

Once the flow has enough results, use the Report menu on the flow page to:

  • open the report in a web view
  • copy the generated report to the clipboard
  • download the report as Markdown
  • download the report as PDF

4. Use the Assistant view to steer an active flow

Each flow also includes an Assistant view for interactive guidance. This is useful when the autonomous run uncovers something that needs human direction instead of a hard restart.

  • Open the Assistant view for the same flow when you want to inspect the current state before changing anything.
  • Use the assistant to check flow status, stop the current task, submit follow-up instructions, or patch the remaining planned subtasks before the next step runs.
  • Treat this as an explicit control path for the current flow, not as an invisible background queue. If you want to change direction, say so clearly and keep the new instruction tied to the current engagement scope.
  • This works best for clarifying scope, redirecting priorities after intermediate findings, or answering an automation checkpoint without losing the rest of the flow context.

5. Manage flow-scoped files

Each flow has its own Files tab in the flow page. Files are scoped to the parent flow: they live in {dataDir}/flow-{id}-data/ on the host and never leak into other flows.

The tab exposes three sources of files:

  • Uploads (uploads/): files you provide from the web UI. Use the Upload files action, or drag and drop directly onto the Files tab. While the agent container is running, uploaded files are also pushed into it at /work/uploads/ so the agent can read them with normal shell tools.
  • Resources (resources/): files attached from your saved user resources library via Attach resources from library. Attached resources are copied into the flow and pushed into the running container at /work/resources/.
  • Container (container/): snapshots pulled from the running agent container via Pull file or directory from container. These are read-only on the flow side and are never sent back to the container.

Per-file actions in the Files tab include Download, Copy path, Save as resource (promote a flow file into your reusable resources library), and Delete. The Pull action is disabled when the container is not running, with the tooltip "Container is not running".

Uploaded files and attached resources are listed automatically in the agent's system prompts via the {{.UserFiles}} template variable, which renders a compact XML block (with nested and `` sections), so the assistant and automation agents can reference them by path without you pasting the contents into chat. Container snapshots are visible in the UI only and are not auto-injected back into the prompt.

Current limits and limitations to be aware of:

  • Maximum upload file size is 300 MB; per upload request up to 1000 files and 2 GB total. File names are capped at 255 bytes (roughly 255 ASCII characters; non-ASCII names use multiple bytes per character).
  • Uploads and resources are mirrored into the running container at the fixed paths /work/uploads/ and /work/resources/; files written to other container paths are not auto-mirrored back into the flow file model. Container snapshots can originate from any container path you pull (for example /etc/...) and are cached on the flow side under container/; they are not pushed back into the container.
  • Container snapshots are point-in-time pulls. Editing a snapshot in the UI does not write back into the running container.
  • Deleting a flow today removes the flow record and its long-term memory entries, but does not yet archive or remove the flow's flow-{id}-data/ directory on disk. Operators are still expected to clean up the data directory manually if they want to reclaim the space.

For early testing, start with a narrow target and a single clear objective. This makes the output easier to review and helps you refine your prompts before running larger assessments.

API Access

PentAGI provides comprehensive programmatic access through both REST and GraphQL APIs, allowing you to integrate penetration testing workflows into your automation pipelines, CI/CD processes, and custom applications.

Generating API Tokens

API tokens are managed through the PentAGI web interface:

  1. Navigate to Settings → API Tokens in the web UI
  2. Click Create Token to generate a new API token
  3. Configure token properties:
    • Name (optional): A descriptive name for the token
    • Expiration Date: When the token will expire (minimum 1 minute, maximum 3 years)
  4. Click Create and copy the token immediately - it will only be shown once for security reasons
  5. Use the token as a Bearer token in your API requests

Each token is associated with your user account and inherits your role's permissions.

Using API Tokens

Include the API token in the Authorization header of your HTTP requests:

# GraphQL API example
curl -X POST https://your-pentagi-instance:8443/api/v1/graphql \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"query": "{ flows { id title status } }"}'

# REST API example
curl https://your-pentagi-instance:8443/api/v1/flows \
  -H "Authorization: Bearer YOUR_API_TOKEN"

API Exploration and Testing

PentAGI provides interactive documentation for exploring and testing API endpoints:

GraphQL Playground

Access the GraphQL Playground at https://your-pentagi-instance:8443/api/v1/graphql/playground

  1. Click the HTTP Headers tab at the bottom
  2. Add your authorization header:
    {
      "Authorization": "Bearer YOUR_API_TOKEN"
    }
    
  3. Explore the schema, run queries, and test mutations interactively

Swagger UI

Access the REST API documentation at https://your-pentagi-instance:8443/api/v1/swagger/index.html

  1. Click the Authorize button
  2. Enter your token in the format: Bearer YOUR_API_TOKEN
  3. Click Authorize to apply
  4. Test endpoints directly from the Swagger UI

Generating API Clients

You can generate type-safe API clients for your preferred programming language using the schema files included with PentAGI:

GraphQL Clients

The GraphQL schema is available at:

  • Web UI: Navigate to Settings to download schema.graphqls
  • Direct file: backend/pkg/graph/schema.graphqls in the repository

Generate clients using tools like:

REST API Clients

The OpenAPI specification is available at:

  • Swagger JSON: https://your-pentagi-instance:8443/api/v1/swagger/doc.json
  • Swagger YAML: Available in backend/pkg/server/docs/swagger.yaml

Generate clients using:

API Usage Examples

Creating a New Flow (GraphQL)

mutation CreateFlow {
  createFlow(
    modelProvider: "openai"
    input: "Test the security of https://example.com"
  ) {
    id
    title
    status
    createdAt
  }
}

Listing Flows (REST API)

curl https://your-pentagi-instance:8443/api/v1/flows \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  | jq '.flows[] | {id, title, status}'

Python Client Example

import requests

class PentAGIClient:
    def __init__(self, base_url, api_token):
        self.base_url = base_url
        self.headers = {
            "Authorization": f"Bearer {api_token}",
            "Content-Type": "application/json"
        }
    
    def create_flow(self, provider, target):
        query = """
        mutation CreateFlow($provider: String!, $input: String!) {
          createFlow(modelProvider: $provider, input: $input) {
            id
            title
            status
          }
        }
        """
        response = requests.post(
            f"{self.base_url}/api/v1/graphql",
            json={
                "query": query,
                "variables": {
                    "provider": provider,
                    "input": target
                }
            },
            headers=self.headers
        )
        return response.json()
    
    def get_flows(self):
        response = requests.get(
            f"{self.base_url}/api/v1/flows",
            headers=self.headers
        )
        return response.json()

# Usage
client = PentAGIClient(
    "https://your-pentagi-instance:8443",
    "your_api_token_here"
)

# Create a new flow
flow = client.create_flow("openai", "Scan https://example.com for vulnerabilities")
print(f"Created flow: {flow}")

# List all flows
flows = client.get_flows()
print(f"Total flows: {len(flows['flows'])}")

TypeScript Client Example

import axios, { AxiosInstance } from 'axios';

interface Flow {
  id: string;
  title: string;
  status: string;
  createdAt: string;
}

class PentAGIClient {
  private client: AxiosInstance;

  constructor(baseURL: string, apiToken: string) {
    this.client = axios.create({
      baseURL: `${baseURL}/api/v1`,
      headers: {
        'Authorization': `Bearer ${apiToken}`,
        'Content-Type': 'application/json',
      },
    });
  }

  async createFlow(provider: string, input: string): Promise {
    const query = `
      mutation CreateFlow($provider: String!, $input: String!) {
        createFlow(modelProvider: $provider, input: $input) {
          id
          title
          status
          createdAt
        }
      }
    `;

    const response = await this.client.post('/graphql', {
      query,
      variables: { provider, input },
    });

    return response.data.data.createFlow;
  }

  async getFlows(): Promise {
    const response = await this.client.get('/flows');
    return response.data.flows;
  }

  async getFlow(flowId: string): Promise {
    const response = await this.client.get(`/flows/${flowId}`);
    return response.data;
  }
}

// Usage
const client = new PentAGIClient(
  'https://your-pentagi-instance:8443',
  'your_api_token_here'
);

// Create a new flow
const flow = await client.createFlow(
  'openai',
  'Perform penetration test on https://example.com'
);
console.log('Created flow:', flow);

// List all flows
const flows = await client.getFlows();
console.log(`Total flows: ${flows.length}`);

Security Best Practices

When working with API tokens:

  • Never commit tokens to version control - use environment variables or secrets management
  • Rotate tokens regularly - set appropriate expiration dates and create new tokens periodically
  • Use separate tokens for different applications - makes it easier to revoke access if needed
  • Monitor token usage - review API token activity in the Settings page
  • Revoke unused tokens - disable or delete tokens that are no longer needed
  • Use HTTPS only - never send API tokens over unencrypted connections

Token Management

  • View tokens: See all your active tokens in Settings → API Tokens
  • Edit tokens: Update token names or revoke tokens
  • Delete tokens: Permanently remove tokens (this action cannot be undone)
  • Token ID: Each token has a unique ID that can be copied for reference

The token list shows:

  • Token name (if provided)
  • Token ID (unique identifier)
  • Status (active/revoked/expired)
  • Creation date
  • Expiration date

Custom LLM Provider Configuration

When using custom LLM providers with the LLM_SERVER_* variables, you can fine-tune the reasoning format used in requests.

[!TIP] For production-grade local deployments, consider using vLLM with Qwen3.5-27B-FP8 for optimal performance. See our comprehensive deployment guide which includes hardware requirements, configuration templates (thinking mode and non-thinking mode), and performance benchmarks showing 13K TPS prompt processing on 4× RTX 5090 GPUs.

Variable Default Description
LLM_SERVER_URL Base URL for the custom LLM API endpoint
LLM_SERVER_KEY API key for the custom LLM provider
LLM_SERVER_MODEL Default model to use (can be overridden in provider config)
LLM_SERVER_CONFIG_PATH Path to the YAML configuration file for agent-specific models
LLM_SERVER_PROVIDER Provider name prefix for model names (e.g., openrouter, deepseek for LiteLLM proxy)
LLM_SERVER_PRESERVE_REASONING false Preserve reasoning content in multi-turn conversations (required by some providers)
LLM_SERVER_API_TYPE azure or azure_ad for an Azure OpenAI deployment, empty for a plain endpoint; see Using Azure OpenAI
LLM_SERVER_API_VERSION 2024-10-21 The api-version Azure requires; ignored by a plain endpoint

The LLM_SERVER_PROVIDER setting is particularly useful when using LiteLLM proxy, which adds a provider prefix to model names. For example, when connecting to Moonshot API through LiteLLM, models like kimi-2.5 become moonshot/kimi-2.5. By setting LLM_SERVER_PROVIDER=moonshot, you can use the same provider configuration file for both direct API access and LiteLLM proxy access without modifications.

The LLM_SERVER_PRESERVE_REASONING setting controls whether reasoning content is preserved in multi-turn conversations:

  • false (default): Reasoning content is not preserved in conversation history
  • true: Reasoning content is preserved and sent in subsequent API calls

This setting is required by some LLM providers (e.g., Moonshot) that return errors like "thinking is enabled but reasoning_content is missing in assistant tool call message" when reasoning content is not included in multi-turn conversations. Enable this setting if your provider requires reasoning content to be preserved.

Using Azure OpenAI

The custom provider calls Azure OpenAI deployments directly:

LLM_SERVER_URL=https://.openai.azure.com  # the Endpoint from the resource's Keys and Endpoint page, without an /openai path
LLM_SERVER_KEY=your_azure_openai_api_key            # KEY 1 or KEY 2 from the same page
LLM_SERVER_API_TYPE=azure                           # the api-key header and deployment URLs
LLM_SERVER_API_VERSION=2024-10-21                   # the default; a newer version works as well
LLM_SERVER_MODEL=                                   # Leave empty, models are specified in the config
LLM_SERVER_CONFIG_PATH=/opt/pentagi/conf/azure-openai.provider.yml

Azure addresses a deployment, not a model: PentAGI calls /openai/deployments//chat/completions?api-version=, so every model: in the provider config must be the name of a deployment in your resource, and LLM_SERVER_PROVIDER must stay empty, because its prefix would become part of the deployment name. The bundled azure-openai.provider.yml expects deployments named after their models: gpt-4.1, gpt-4.1-mini and o4-mini. If yours are named differently, copy the file, change the model: values, and mount your copy:

PENTAGI_LLM_SERVER_CONFIG_PATH=/path/on/host/my-azure.provider.yml  # mounted at /opt/pentagi/conf/custom.provider.yml
LLM_SERVER_CONFIG_PATH=/opt/pentagi/conf/custom.provider.yml

After docker compose up -d, check the deployments before running a flow. Without -config, ctester tests the file LLM_SERVER_CONFIG_PATH names, the bundled one or your copy; the bundled configuration's own run is in examples/tests/azure-openai-report.md:

docker exec -it pentagi /opt/pentagi/bin/ctester
  • PentAGI does not read a model catalogue from Azure, so it knows no context window for a deployment and does not compact agent chains to fit one.
  • LLM_SERVER_API_TYPE=azure_ad sends LLM_SERVER_KEY as a Microsoft Entra ID bearer token instead of an API key. PentAGI does not refresh it, and Entra tokens expire after about an hour, so use an API key for anything longer.
  • Azure's OpenAI-compatible v1 API works as a plain endpoint too: set LLM_SERVER_URL=https://.openai.azure.com/openai/v1 and leave LLM_SERVER_API_TYPE empty; model: is still the deployment name.
  • The installer's custom provider form has API Type and API Version fields for the same two settings.
  • Embeddings are configured separately through EMBEDDING_* and have no Azure mode; LLM_SERVER_API_TYPE does not apply to them.

Troubleshooting: tool-call (function-call) parser errors

PentAGI drives its agents with tool calls (also called function calls), so any custom OpenAI-compatible backend configured through LLM_SERVER_* must return valid tool-call JSON in the format the OpenAI Chat Completions API defines. When the backend emits malformed, truncated, or non-conforming tool-call arguments, the agent chain cannot continue.

Self-hosted engines such as llama.cpp, SGLang, and vLLM usually require a specific tool-call parser and a matching chat template to produce correct tool-call output. If the parser is missing or mismatched for the model you are serving, tool-call arguments can come back corrupted. Compatibility therefore depends on the backend's tool-call/function-call behavior and configuration, not on PentAGI alone; not every llama.cpp or SGLang setup produces valid tool calls out of the box.

Typical symptoms:

  • Backend or proxy errors such as Failed to parse tool call arguments as JSON (often surfaced through a LiteLLM proxy as an HTTP 500), or other unexpected 5xx/4xx responses from the LLM endpoint.
  • A flow that runs for a few steps and then stops responding to new input in the UI.
  • Repeated or looping tool calls that never converge.
  • A flow that fails right at the start with failed to select primary docker image via llm call, because the first action in a flow is an LLM tool call to choose the container image; a backend that cannot return a valid tool call fails at this step too.

How to investigate:

  1. Check both sides of the connection: the PentAGI logs (docker compose logs -f pentagi) and the inference backend or proxy logs (llama.cpp, SGLang, vLLM, or LiteLLM). The backend log usually shows the same parse error when it produced the malformed tool call.
  2. Validate the provider before running a full flow with the ctester utility, which exercises tool-calling agent types directly. See Testing LLM Agents.
  3. Confirm the backend's tool-call parser and chat template are the ones recommended for the model you are serving, and that the model itself supports tool calling.
  4. Update PentAGI to the latest build. Recent versions sanitize malformed function-call arguments returned by the model so a single bad response no longer stalls the whole flow; older builds forwarded the corrupted arguments and could get stuck.

Ollama Provider Configuration

PentAGI supports Ollama for both local LLM inference (zero-cost, enhanced privacy) and Ollama Cloud (managed service with free tier).

Configuration Variables

Variable Default Description
OLLAMA_SERVER_URL URL of your Ollama server or Ollama Cloud
OLLAMA_SERVER_API_KEY API key for Ollama Cloud authentication
OLLAMA_SERVER_MODEL Default model for inference
OLLAMA_SERVER_CONFIG_PATH Path to custom agent configuration file
OLLAMA_SERVER_PULL_MODELS_TIMEOUT 600 Timeout for model downloads (seconds)
OLLAMA_SERVER_PULL_MODELS_ENABLED false Auto-download models on startup
OLLAMA_SERVER_LOAD_MODELS_ENABLED false Query server for available models

Ollama Cloud Configuration

Ollama Cloud provides managed inference with a free tier and paid plans. Paid usage is billed from included or purchased credits at each model's published per-token rate; off-peak discounts apply to selected DeepSeek models.

Free Tier Setup (Single Model)

# Free tier allows one model at a time
OLLAMA_SERVER_URL=https://ollama.com
OLLAMA_SERVER_API_KEY=your_ollama_cloud_api_key
OLLAMA_SERVER_MODEL=gpt-oss:120b  # Example: OpenAI OSS 120B model

Paid Tier Setup (Multi-Model with Pre-built Configuration)

For paid tiers supporting concurrent requests and pay-as-you-go credits, use the pre-built Ollama Cloud configuration:

# Using pre-built Ollama Cloud configuration (included in Docker image)
OLLAMA_SERVER_URL=https://ollama.com
OLLAMA_SERVER_API_KEY=your_ollama_cloud_api_key
OLLAMA_SERVER_CONFIG_PATH=/opt/pentagi/conf/ollama-cloud.provider.yml

The pre-built ollama-cloud.provider.yml configuration includes optimized model assignments for all agent types:

  • Simple/Reflector/Enricher: minimax-m2.7:cloud - Reliable low-latency utility model
  • Simple JSON/Primary Agent/Assistant/Pentester/Searcher: deepseek-v4.1-flash:cloud - Efficient long-context model at the same input/output rate as MiniMax M2.7
  • Generator/Adviser: kimi-k3:cloud - Flagship long-horizon planning and knowledge-work model
  • Refiner: glm-5.3:cloud - Flagship coding and agentic model
  • Coder/Installer: kimi-k2.7-code:cloud - Coding-specialized long-context model

Custom Configuration (Advanced)

To create your own agent configuration, mount a custom file from your host filesystem:

# Using custom provider configuration
OLLAMA_SERVER_URL=https://ollama.com
OLLAMA_SERVER_API_KEY=your_ollama_cloud_api_key
OLLAMA_SERVER_CONFIG_PATH=/opt/pentagi/conf/ollama.provider.yml

# Mount custom configuration from host filesystem (in .env or docker-compose override)
PENTAGI_OLLAMA_SERVER_CONFIG_PATH=/path/on/host/my-ollama-config.yml

The PENTAGI_OLLAMA_SERVER_CONFIG_PATH environment variable maps your host configuration file to /opt/pentagi/conf/ollama.provider.yml inside the container.

Example custom configuration (my-ollama-config.yml):

primary_agent:
  model: "deepseek-v4-flash:cloud"
  temperature: 1.0
  top_p: 0.95
  max_tokens: 32768

coder:
  model: "kimi-k2.7-code:cloud"
  temperature: 1.0
  max_tokens: 20480

Local Ollama Configuration

For self-hosted Ollama instances:

# Basic local Ollama setup
OLLAMA_SERVER_URL=http://localhost:11434
OLLAMA_SERVER_MODEL=llama3.1:8b-instruct-q8_0

# Production setup with auto-pull and model discovery
OLLAMA_SERVER_URL=http://ollama-server:11434
OLLAMA_SERVER_PULL_MODELS_ENABLED=true
OLLAMA_SERVER_PULL_MODELS_TIMEOUT=900
OLLAMA_SERVER_LOAD_MODELS_ENABLED=true

# Using pre-built configurations from Docker image
OLLAMA_SERVER_CONFIG_PATH=/opt/pentagi/conf/ollama-llama318b.provider.yml
# or
OLLAMA_SERVER_CONFIG_PATH=/opt/pentagi/conf/ollama-qwen332b-fp16-tc.provider.yml
# or
OLLAMA_SERVER_CONFIG_PATH=/opt/pentagi/conf/ollama-qwq32b-fp16-tc.provider.yml

Performance Considerations:

  • Model Discovery (OLLAMA_SERVER_LOAD_MODELS_ENABLED=true): Adds 1-2s startup latency querying Ollama API
  • Auto-pull (OLLAMA_SERVER_PULL_MODELS_ENABLED=true): First startup may take several minutes downloading models
  • Pull timeout (OLLAMA_SERVER_PULL_MODELS_TIMEOUT=900): 15 minutes in seconds
  • Static Config: Disable both flags and specify models in config file for fastest startup

Creating Custom Ollama Models with Extended Context

PentAGI requires models with larger context windows than the default Ollama configurations. You need to create custom models with increased num_ctx parameter through Modelfiles. While typical agent workflows consume around 64K tokens, PentAGI uses 110K context size for safety margin and handling complex penetration testing scenarios.

Important: The num_ctx parameter can only be set during model creation via Modelfile - it cannot be changed after model creation or overridden at runtime.

Example: Qwen3 32B FP16 with Extended Context

Create a Modelfile named Modelfile_qwen3_32b_fp16_tc:

FROM qwen3:32b-fp16
PARAMETER num_ctx 110000
PARAMETER temperature 0.3
PARAMETER top_p 0.8
PARAMETER min_p 0.0
PARAMETER top_k 20
PARAMETER repeat_penalty 1.1

Build the custom model:

ollama create qwen3:32b-fp16-tc -f Modelfile_qwen3_32b_fp16_tc
Example: QwQ 32B FP16 with Extended Context

Create a Modelfile named Modelfile_qwq_32b_fp16_tc:

FROM qwq:32b-fp16
PARAMETER num_ctx 110000
PARAMETER temperature 0.2
PARAMETER top_p 0.7
PARAMETER min_p 0.0
PARAMETER top_k 40
PARAMETER repeat_penalty 1.2

Build the custom model:

ollama create qwq:32b-fp16-tc -f Modelfile_qwq_32b_fp16_tc

Note: The QwQ 32B FP16 model requires approximately 71.3 GB VRAM for inference. Ensure your system has sufficient GPU memory before attempting to use this model.

These custom models are referenced in the pre-built provider configuration files (ollama-qwen332b-fp16-tc.provider.yml and ollama-qwq32b-fp16-tc.provider.yml) that are included in the Docker image at /opt/pentagi/conf/.

OpenAI Provider Configuration

PentAGI integrates with OpenAI's comprehensive model lineup, featuring advanced reasoning capabilities with extended chain-of-thought, agentic models with enhanced tool integration, and specialized code models for security engineering.

Configuration Variables

Variable Default Description
OPEN_AI_KEY API key for OpenAI services
OPEN_AI_SERVER_URL https://api.openai.com/v1 OpenAI API endpoint

Configuration Examples

# Basic OpenAI setup
OPEN_AI_KEY=your_openai_api_key
OPEN_AI_SERVER_URL=https://api.openai.com/v1

# Using with proxy for enhanced security
OPEN_AI_KEY=your_openai_api_key
PROXY_URL=http://your-proxy:8080

Supported Models

PentAGI supports 32 OpenAI models with tool calling, streaming, reasoning modes, and prompt caching. Models marked with * are used in default configuration. Models marked ⚠️ are deprecated by OpenAI and kept only for backward compatibility with agent configs already pinned to those names — avoid them for new assignments.

GPT-5.6 Series - Latest Frontier (Feb 2026 knowledge cutoff, 1.05M context, 128K max output)

Model ID Thinking Reasoning Effort Price (Input/Output/Cache) Use Case
gpt-5.6-sol ✅ low/medium/high/xhigh $5.00/$30.00/$0.50 Frontier model for complex professional work, most demanding autonomous pentesting, sophisticated exploit chain development, deep multi-stage attack simulation
gpt-5.6-terra* ✅ low/medium/high/xhigh $2.50/$15.00/$0.25 Balances intelligence and cost; multi-phase security assessments, coordinated multi-tool pentesting (generator/refiner/adviser/coder default)
gpt-5.6-luna ✅ low/medium/high/xhigh $1.00/$6.00/$0.10 Optimized for cost-sensitive, high-volume workloads; rapid reconnaissance, bulk vulnerability scanning, real-time monitoring

GPT-5.5 Series - Frontier (Dec 2025 knowledge cutoff, 1.05M context, 128K max output)

Model ID Thinking Reasoning Effort Price (Input/Output/Cache) Use Case
gpt-5.5 ✅ none/low/medium/high/xhigh $5.00/$30.00/$0.50 New class of intelligence for coding and professional work; complex security research, advanced autonomous pentesting
gpt-5.5-pro ✅ medium/high/xhigh $30.00/$180.00/$0.00 Uses more compute for smarter, more precise responses; no cached-input discount; mission-critical security research, zero-day discovery

GPT-5.4 Series - Advanced Reasoning at Scale (1M context)

Model ID Thinking Reasoning Effort Price (Input/Output/Cache) Use Case
gpt-5.4 ✅ low/medium/high/xhigh $2.50/$15.00/$0.25 Best intelligence at scale for agentic, coding, and professional workflows; maximum cognitive depth for pentesting
gpt-5.4-mini* ✅ low/medium/high/xhigh $0.75/$4.50/$0.075 Strongest mini model for coding, computer use, subagents (primary_agent/assistant/reflector/installer/pentester default)
gpt-5.4-nano* ✅ low/medium/high/xhigh $0.20/$1.25/$0.02 Cheapest GPT-5.4-class model for simple, high-volume tasks (simple/simple_json/searcher/enricher default)

GPT-5.2 Series - Previous Flagship Agentic

Model ID Thinking Reasoning Effort Price (Input/Output/Cache) Use Case
gpt-5.2 ✅ low/medium/high/xhigh $1.75/$14.00/$0.175 Superseded by 5.4/5.6; autonomous security research, complex exploit chain development
gpt-5.2-pro ✅ medium/high/xhigh $21.00/$168.00/$0.00 Superior agentic coding and long-context performance, mission-critical security research, zero-day discovery

GPT-5/5.1 Series - Advanced Agentic Models

Model ID Thinking Price (Input/Output/Cache) Use Case
gpt-5 ✅ $1.25/$10.00/$0.125 Autonomous security research, exploit chain development, coordinating multi-tool pentesting workflows
gpt-5.1 ✅ $1.25/$10.00/$0.125 Bridges GPT-5 and GPT-5.2 with faster responses; balanced penetration testing with strong tool coordination
gpt-5-pro ✅ (high) $15.00/$120.00/$0.00 Reduced hallucinations, exceptional accuracy, critical security operations
gpt-5-mini ✅ $0.25/$2.00/$0.025 Automated vulnerability analysis, exploit generation with strong function calling
gpt-5-nano ✅ $0.05/$0.40/$0.005 High-throughput security scanning, reconnaissance, real-time monitoring

GPT-4.1 Series - Enhanced Intelligence (Non-Reasoning)

Model ID Thinking Price (Input/Output/Cache) Use Case
gpt-4.1 ❌ $2.00/$8.00/$0.50 Superior function calling, complex threat analysis, sophisticated exploit development
gpt-4.1-mini ❌ $0.40/$1.60/$0.10 Routine security assessments, automated code analysis (no longer used in default configuration)

GPT-4o Series - Multimodal (Non-Reasoning)

Model ID Thinking Price (Input/Output/Cache) Use Case
gpt-4o-mini ❌ $0.15/$0.60/$0.075 Compact multimodal with strong function calling, high-frequency scanning, cost-effective bulk operations

o-Series - Advanced Reasoning Models (Current)

Model ID Thinking Price (Input/Output/Cache) Use Case
o3 ✅ $2.00/$8.00/$0.50 Succeeded by GPT-5; multi-stage attack chains, deep vulnerability analysis
o3-pro ✅ $20.00/$80.00/$0.00 More compute for better responses; zero-day research, critical security investigations

Deprecated Models - Kept for Backward Compatibility ⚠️

These models were marked deprecated by OpenAI. PentAGI keeps them defined only so that pre-existing agent configs pinned to these names keep working; do not assign them to new agents.

Model ID Thinking Price (Input/Output/Cache) Notes
gpt-5.2-codex ✅ $1.75/$14.00/$0.175 Superseded code-specialized model; use gpt-5.6-terra/gpt-5.4-mini instead
gpt-5.1-codex-max ✅ $1.25/$10.00/$0.125 Superseded; enhanced reasoning for coding workflows
gpt-5.1-codex ✅ $1.25/$10.00/$0.125 Superseded standard code-optimized model
gpt-5-codex ✅ $1.25/$10.00/$0.125 Superseded foundational code-specialized model
gpt-5.1-codex-mini ✅ $0.25/$2.00/$0.025 Superseded compact code model
codex-mini-latest ✅ $1.50/$6.00/$0.375 Superseded compact code model
gpt-4o ❌ $2.50/$10.00/$1.25 Superseded by GPT-5.x/5.6 series multimodal flagship
gpt-4.1-nano ❌ $0.10/$0.40/$0.025 Superseded ultra-fast lightweight model
o3-mini ✅ $1.10/$4.40/$0.55 Superseded compact reasoning model
o4-mini ✅ $1.10/$4.40/$0.275 Succeeded by gpt-5-mini
o1 ✅ $15.00/$60.00/$7.50 Superseded premier reasoning model
o1-pro ✅ $150.00/$600.00/$0.00 Superseded, highest cost point of the o-series

Prices: Per 1M tokens. Reasoning models include thinking tokens in output pricing.

[!WARNING] GPT-5/5.1/5.2 Models - Trusted Access Required

The original GPT-5, GPT-5.1, and GPT-5.2 models (gpt-5, gpt-5.1, gpt-5.2, gpt-5-pro, gpt-5.2-pro, and all deprecated Codex variants) work unstably with PentAGI and may trigger OpenAI's cybersecurity safety mechanisms without verified access. This does not affect the newer GPT-5.4/5.5/5.6 series used in PentAGI's default configuration below.

To use these models reliably:

  1. Individual users: Verify your identity at chatgpt.com/cyber
  2. Enterprise teams: Request trusted access through your OpenAI representative
  3. Security researchers: Apply for the Cybersecurity Grant Program (includes $10M in API credits)

Recommended alternatives without verification:

  • Use PentAGI's defaults — gpt-5.4-mini/gpt-5.4-nano/gpt-5.6-terra — which work out of the box
  • Use o3/o3-pro for reasoning tasks
  • Use gpt-4.1 series for general intelligence and function calling without reasoning

Reasoning Configuration:

  • Reasoning forced off by default: every default agent assigned gpt-5.4-mini or gpt-5.6-terra (primary_agent, assistant, generator, refiner, adviser, reflector, coder, installer, pentester) sets reasoning: {mode: off} — this genuinely disables reasoning, it is not simply "low effort". PentAGI calls OpenAI exclusively through /v1/chat/completions (never /v1/responses), and this endpoint rejects requests that combine function tools with these models' default-on thinking; forcing thinking off is required for tool calls to work reliably (see the investigation notes in backend/pkg/providers/openai/config.yml).
  • No override needed for gpt-5.4-nano: used for simple, simple_json, searcher, and enricher, this tier does not default to thinking on, so tools attach without conflict and no reasoning override is required.
  • Manual tuning available: outside the default assignments, GPT-5.6/5.5/5.4/5.2 series models expose explicit reasoning effort levels (low/medium/high/xhigh, plus none on GPT-5.5) for custom agent configs that need variable reasoning depth with tool calling disabled or via /v1/responses.

Key Features:

  • Extended Reasoning: GPT-5.4/5.5/5.6 and o-series models with chain-of-thought for complex security analysis
  • Agentic Intelligence: GPT-5.4/5.5/5.6 series with enhanced tool integration, million-token context windows, and autonomous capabilities
  • Prompt Caching: Cost reduction on repeated context (10-50% of input price)
  • Code Specialization: Legacy Codex models remain available (deprecated) for vulnerability discovery and exploit development in pinned configs
  • Multimodal Support: gpt-4o-mini for vision-based security assessments
  • Tool Calling: Robust function calling across all models for pentesting tool orchestration
  • Streaming: Real-time response streaming for interactive workflows
  • Proven Track Record: Industry-leading models with CVE discoveries and real-world security applications

Anthropic Provider Configuration

PentAGI integrates with Anthropic's Claude models, featuring advanced extended thinking capabilities, exceptional safety mechanisms, and sophisticated understanding of complex security contexts with prompt caching.

Configuration Variables

Variable Default Description
ANTHROPIC_API_KEY API key for Anthropic services
ANTHROPIC_SERVER_URL https://api.anthropic.com/v1 Anthropic API endpoint

Configuration Examples

# Basic Anthropic setup
ANTHROPIC_API_KEY=your_anthropic_api_key
ANTHROPIC_SERVER_URL=https://api.anthropic.com/v1

# Using with proxy for secure environments
ANTHROPIC_API_KEY=your_anthropic_api_key
PROXY_URL=http://your-proxy:8080

# Workload Identity Federation instead of a long-lived key (leave ANTHROPIC_API_KEY empty)
ANTHROPIC_FEDERATION_RULE_ID=fdrl_...
ANTHROPIC_ORGANIZATION_ID=00000000-0000-0000-0000-000000000000
ANTHROPIC_SERVICE_ACCOUNT_ID=svac_...
ANTHROPIC_WORKSPACE_ID=wrkspc_...
ANTHROPIC_IDENTITY_TOKEN_FILE=/var/run/secrets/anthropic.com/token

With Workload Identity Federation, PentAGI exchanges the identity token your platform issues (Kubernetes, GitHub Actions, cloud IAM, any OIDC issuer) for a short-lived Claude API token and refreshes it on its own. The token file is re-read on every exchange and must be mounted into the pentagi container. A token carrying a jti claim, as Kubernetes and GitHub Actions tokens do, can be exchanged only once, so the file must hold a new token before each refresh: rotate it well within the lifetime of the minted token. ANTHROPIC_IDENTITY_TOKEN takes the token itself where the platform injects it as a variable, but it is read once and cannot be rotated, so it suits only runs shorter than the identity token's lifetime. A set ANTHROPIC_API_KEY wins over federation. See backend/docs/config.md for every variable.

[!NOTE] Google Vertex AI for Claude models

PentAGI does not currently expose a dedicated Google Vertex AI configuration path for Anthropic Claude in .env. There is no separate Vertex AI API key field at this time, and the existing Anthropic variables (ANTHROPIC_API_KEY, ANTHROPIC_SERVER_URL) target the direct Anthropic API. Supported routes for Claude are:

If you need to use Vertex AI today, the safest supported workaround is to expose Vertex AI through an OpenAI-compatible proxy or gateway that translates Vertex AI calls into the Chat Completions format while preserving the chat and tool-call behavior PentAGI relies on, then point the Custom LLM provider at that gateway via LLM_SERVER_URL, LLM_SERVER_KEY, and LLM_SERVER_MODEL. This path is only as reliable as the gateway you choose.

Supported Models

PentAGI supports 9 Claude models with tool calling, streaming, extended thinking, adaptive thinking, and prompt caching. Models marked with * are used in default configuration.

Claude 5 Series - Newest Models (2026)

Model ID Thinking Release Date Price (Input/Output/Cache R/W) Use Case
claude-sonnet-5* ✅ Jul 2026 $3.00/$15.00/$0.30/$3.75 Best combination of speed and intelligence for coding, agents, and professional work at scale. Adaptive thinking only (manual budget thinking rejected); sampling parameters not supported. Default model for primary agent, assistant, coder, adviser, installer, pentester
claude-fable-5 ✅ Jun 2026 $10.00/$50.00/$1.00/$12.50 Anthropic's most capable widely released model for long-running agents and the most demanding reasoning workloads. Adaptive thinking always on (budget thinking and an explicit disable are rejected); sampling parameters not supported

Claude 4 Series

Model ID Thinking Release Date Price (Input/Output/Cache R/W) Use Case
claude-opus-4-8* ✅ May 2026 $5.00/$25.00/$0.50/$6.25 Flagship for coding, agents, and deep reasoning in enterprise security workflows. Adaptive thinking only — budget thinking and sampling params (temperature/top_p/top_k) are rejected. Default model for generator and refiner; most demanding exploit development and multi-stage attack simulation
claude-opus-4-7 ✅ Apr 2026 $5.00/$25.00/$0.50/$6.25 Advanced software engineering and long-running agentic security analysis. Adaptive thinking only (manual budget thinking rejected)
claude-sonnet-4-6 ✅ Feb 2026 $3.00/$15.00/$0.30/$3.75 Best speed/intelligence balance with adaptive thinking. Multi-phase security assessments, intelligent vulnerability analysis, real-time threat hunting
claude-opus-4-6 ✅ Feb 2026 $5.00/$25.00/$0.50/$6.25 Most intelligent model for autonomous agents and coding. Extended + adaptive thinking for complex exploit development, multi-stage attack simulation
claude-haiku-4-5* ❌ Oct 2025 $1.00/$5.00/$0.10/$1.25 Fast and efficient model with exceptional function calling and low latency, no thinking support. Default model for simple, simple_json, reflector, searcher, enricher; high-frequency scanning, real-time monitoring, bulk automated testing

Legacy Models - Still Supported

Model ID Thinking Release Date Price (Input/Output/Cache R/W) Use Case
claude-sonnet-4-5 ✅ Sep 2025 $3.00/$15.00/$0.30/$3.75 State-of-the-art reasoning (superseded by sonnet-4-6/sonnet-5). Sophisticated penetration testing, advanced threat analysis
claude-opus-4-5 ✅ Nov 2025 $5.00/$25.00/$0.50/$6.25 Ultimate reasoning (superseded by opus-4-6/4-7/4-8). Critical security research, zero-day discovery, red team operations

Prices: Per 1M tokens. Cache pricing includes both Read and Write costs.

Extended Thinking Configuration (default agent config, see backend/pkg/providers/anthropic/config.yml):

  • Generator / Refiner (claude-opus-4-8): adaptive reasoning at xhigh/high effort for maximum reasoning depth on complex exploit development
  • Primary agent, assistant, coder, adviser, installer, pentester (claude-sonnet-5): adaptive reasoning (adviser at xhigh effort) for balanced code analysis and vulnerability research
  • Reflector, searcher (claude-haiku-4-5): fixed reasoning budget of 1024 tokens for focused reasoning on specific tasks
  • Simple, simple_json, enricher (claude-haiku-4-5): no thinking, optimized for speed

Key Features:

  • Extended Thinking: All Claude 4.5+ models support configurable chain-of-thought reasoning depths for complex security analysis
  • Adaptive Thinking: Claude 4.6 series (Opus/Sonnet) dynamically adjusts reasoning depth based on task complexity; Claude Opus 4.7/4.8 and the Claude 5 series (Sonnet/Fable) are adaptive-thinking-only (manual budget thinking and sampling parameters are rejected with HTTP 400)
  • Prompt Caching: Significant cost reduction with separate read/write pricing (10% read, 125% write of input)
  • Extended Context Window: 200K tokens standard, up to 1M tokens (beta) for Claude Opus/Sonnet 4.6 for comprehensive codebase analysis
  • Tool Calling: Robust function calling with exceptional accuracy for security tool orchestration
  • Streaming: Real-time response streaming for interactive penetration testing workflows
  • Safety-First Design: Built-in safety mechanisms ensuring responsible security testing practices
  • Multimodal Support: Vision capabilities in latest models for screenshot analysis and UI security assessment
  • Constitutional AI: Advanced safety training providing reliable and ethical security guidance

Google AI (Gemini) Provider Configuration

PentAGI integrates with Google's Gemini models through the Google AI API, offering state-of-the-art multimodal reasoning capabilities with extended thinking and context caching.

Configuration Variables

Variable Default Description
GEMINI_API_KEY API key for Google AI services
GEMINI_SERVER_URL https://generativelanguage.googleapis.com Google AI API endpoint

Configuration Examples

# Basic Gemini setup
GEMINI_API_KEY=your_gemini_api_key
GEMINI_SERVER_URL=https://generativelanguage.googleapis.com

# Using with proxy
GEMINI_API_KEY=your_gemini_api_key
PROXY_URL=http://your-proxy:8080

Supported Models

PentAGI lists 14 Gemini model IDs that responded through the tested API endpoint. Availability in the catalogue does not mean a model is suitable for every agent role: the saved-chain results below determine the bundled defaults. Models marked with * are used in config.yml; the full replay matrix describes role and outcome limits.

Model ID Thinking Context Price (Input/Output/Cache) Saved-chain result
gemini-3.8-flash ✅ 1M $0.75/$3.75/$0.075 API responds; 10/10 saved generator calls were content-filtered
gemini-3.7-flash ✅ 1M $0.75/$3.75/$0.075 API responds; 10/10 saved generator calls were content-filtered; former default
gemini-3.6-flash ✅ 1M $0.75/$3.75/$0.075 API responds; 10/10 saved generator calls returned text without a tool
gemini-3.5-flash ✅ 1M $1.5/$9/$0.15 API responds; 10/10 saved generator calls returned text without a tool
gemini-3.5-flash-lite* ✅ 1M $0.3/$2.5/$0.03 Complex-role default; eight direct plans and two after a safe memorist continuation
gemini-3.1-pro-preview ✅ 1M $2/$12/$0.2 API responds; 10/10 saved generator calls returned text without a tool
gemini-3.1-pro-preview-customtools ✅ 1M $2/$12/$0.2 API responds; 10/10 saved generator calls returned text without a tool
gemini-3.1-flash-lite* ✅ 1M $0.25/$1.5/$0.025 Simple-role default; direct plan on all ten saved generator chains
gemini-3-flash-preview ✅ 1M $0.5/$3/$0.05 API responds; three direct plans and five generator timeouts
gemini-2.5-pro ✅ 1M $1.25/$10/$0.125 API responds; ten direct plans, but four role-smoke provider errors
gemini-2.5-flash ✅ 1M $0.3/$2.5/$0.03 Ten direct plans; access is limited for new API projects
gemini-2.5-flash-lite ✅ 1M $0.1/$0.4/$0.01 API responds; five of ten generator answers were empty
gemma-4-31b-it ✅ 256K Free/Free/Free API responds; seven direct plans and three generator timeouts
gemma-4-26b-a4b-it ✅ 256K Free/Free/Free Direct plan on all ten saved generator chains

Prices: Per 1M tokens. Google limits Gemini 2.5 access for new projects, so PentAGI uses the tested 3.1 Flash-Lite as its fallback and simple-role default. Existing custom 2.5 configurations remain usable where the project has access.

Default Model Assignments (config.yml):

  • gemini-3.1-flash-lite - simple, simple_json, reflector, searcher, enricher
  • gemini-3.5-flash-lite - primary_agent, assistant, generator, refiner, adviser, coder, installer, pentester

Both default models passed a basic API/configuration smoke for every role. Google reports stronger agentic and coding performance for 3.5 Flash-Lite than 3.1 Flash-Lite; PentAGI's saved production chains cover generator and reflector only, so the smoke does not establish semantic quality or multi-turn reliability for other roles. On two saved generator chains, 3.5 Flash-Lite first chose the allowed memorist tool and called subtask_list after a synthetic tool response; no real tool was executed in that continuation.

Reasoning Effort Levels:

  • High: Deeper reasoning for planning and review (primary_agent, generator, refiner, adviser)
  • Medium: Balanced reasoning for execution roles (assistant, coder, installer, pentester)
  • Not set: No thinking setting is sent and the model's own default applies (simple, simple_json, reflector, searcher, enricher)

AWS Bedrock Provider Configuration

PentAGI integrates with Amazon Bedrock, offering access to 20+ foundation models from leading AI companies including Anthropic, Amazon, Cohere, DeepSeek, OpenAI, Qwen, Mistral, and Moonshot.

Configuration Variables

Variable Default Description
BEDROCK_REGION us-east-1 AWS region for Bedrock service
BEDROCK_DEFAULT_AUTH false Use AWS SDK default credential chain (environment, EC2 role, ~/.aws/credentials) - highest priority
BEDROCK_BEARER_TOKEN Bearer token authentication - priority over static credentials
BEDROCK_ACCESS_KEY_ID AWS access key ID for static credentials
BEDROCK_SECRET_ACCESS_KEY AWS secret access key for static credentials
BEDROCK_SESSION_TOKEN AWS session token for temporary credentials (optional, used with static credentials)
BEDROCK_SERVER_URL Custom Bedrock endpoint (VPC endpoints, local testing)
BEDROCK_CONFIG_PATH Path to a custom YAML provider config file (overrides the built-in default config for model/pricing definitions)

Authentication Priority: BEDROCK_DEFAULT_AUTH → BEDROCK_BEARER_TOKEN → BEDROCK_ACCESS_KEY_ID+BEDROCK_SECRET_ACCESS_KEY

Configuration Examples

# Recommended: Default AWS SDK authentication (EC2/ECS/Lambda roles)
BEDROCK_REGION=us-east-1
BEDROCK_DEFAULT_AUTH=true

# Bearer token authentication (AWS STS, custom auth)
BEDROCK_REGION=us-east-1
BEDROCK_BEARER_TOKEN=your_bearer_token

# Static credentials (development, testing)
BEDROCK_REGION=us-east-1
BEDROCK_ACCESS_KEY_ID=your_aws_access_key
BEDROCK_SECRET_ACCESS_KEY=your_aws_secret_key

# With proxy and custom endpoint
BEDROCK_REGION=us-east-1
BEDROCK_DEFAULT_AUTH=true
BEDROCK_SERVER_URL=https://bedrock-runtime.us-east-1.vpce-xxx.amazonaws.com
PROXY_URL=http://your-proxy:8080

Custom Provider Config and Models (advanced)

By default the Bedrock provider uses a per-agent config and model catalog compiled into the binary. One optional path overrides them without rebuilding:

This is useful to expose a Bedrock model newer than the compiled-in catalog — for example Z.AI's zai.glm-4.7-flash. Use the exact Model ID from the model's AWS Bedrock detail page; add a us./eu./apac. inference-profile prefix only when that page marks the model as requiring cross-region inference (zai.glm-4.7-flash is In-Region, so it is used as-is, with no prefix).

With Docker Compose the file is mounted from the host at /opt/pentagi/conf/bedrock.provider.yml. The default source is ./example.bedrock.provider.yml beside docker-compose.yml, a copy of the built-in config (examples/configs/bedrock.provider.yml) that the installer extracts, so editing it and pointing the backend at the mounted path is enough:

# tell the backend to read the mounted file
BEDROCK_CONFIG_PATH=/opt/pentagi/conf/bedrock.provider.yml

# optional: mount another file from the host instead of ./example.bedrock.provider.yml
PENTAGI_BEDROCK_CONFIG_PATH=/path/on/host/my-bedrock.provider.yml

Supported Models

PentAGI supports 24 AWS Bedrock models with tool calling, streaming, and multimodal capabilities. Models marked with * are used in default configuration.

Model ID Provider Thinking Multimodal Price (Input/Output) Use Case
us.amazon.nova-2-lite-v1:0 Amazon Nova ❌ ✅ $0.33/$2.75 Adaptive reasoning, efficient thinking
us.amazon.nova-premier-v1:0 Amazon Nova ❌ ✅ $2.50/$12.50 Complex reasoning, advanced analysis
us.amazon.nova-pro-v1:0 Amazon Nova ❌ ✅ $0.80/$3.20 Balanced accuracy, speed, cost
us.amazon.nova-lite-v1:0 Amazon Nova ❌ ✅ $0.06/$0.24 Fast processing, high-volume operations
us.amazon.nova-micro-v1:0 Amazon Nova ❌ ❌ $0.035/$0.14 Ultra-low latency, real-time monitoring
us.anthropic.claude-opus-4-8 Anthropic ✅ ✅ $5.00/$25.00 Flagship coding/agents/deep reasoning; adaptive thinking only (sampling params rejected)
us.anthropic.claude-opus-4-7 Anthropic ✅ ✅ $5.00/$25.00 Advanced engineering, long-running agents; adaptive thinking only
us.anthropic.claude-opus-4-6-v1* Anthropic ✅ ✅ $5.00/$25.00 World-class coding, enterprise agents
us.anthropic.claude-sonnet-4-6 Anthropic ✅ ✅ $3.00/$15.00 Frontier intelligence, enterprise scale
us.anthropic.claude-opus-4-5-20251101-v1:0 Anthropic ✅ ✅ $5.00/$25.00 Multi-day software development
us.anthropic.claude-haiku-4-5-20251001-v1:0* Anthropic ✅ ✅ $1.00/$5.00 Near-frontier performance, high speed
us.anthropic.claude-sonnet-4-5-20250929-v1:0* Anthropic ✅ ✅ $3.00/$15.00 Real-world agents, coding excellence
us.anthropic.claude-sonnet-4-20250514-v1:0 Anthropic ✅ ✅ $3.00/$15.00 Balanced performance, production-ready
us.anthropic.claude-3-5-haiku-20241022-v1:0 Anthropic ❌ ❌ $0.80/$4.00 Fastest model, cost-effective scanning
cohere.command-r-plus-v1:0 Cohere ❌ ❌ $3.00/$15.00 Large-scale operations, superior RAG
deepseek.v3.2 DeepSeek ❌ ❌ $0.58/$1.68 Long-context reasoning, efficiency
openai.gpt-oss-120b-1:0* OpenAI (OSS) ✅ ❌ $0.15/$0.60 Strong reasoning, scientific analysis
openai.gpt-oss-20b-1:0 OpenAI (OSS) ✅ ❌ $0.07/$0.30 Efficient coding, software development
qwen.qwen3-next-80b-a3b Qwen ❌ ❌ $0.15/$1.20 Ultra-long context, flagship reasoning
qwen.qwen3-32b-v1:0 Qwen ❌ ❌ $0.15/$0.60 Balanced reasoning, research use cases
qwen.qwen3-coder-30b-a3b-v1:0 Qwen ❌ ❌ $0.15/$0.60 Vibe coding, natural-language first
qwen.qwen3-coder-next Qwen ❌ ❌ $0.45/$1.80 Tool use, function calling optimized
mistral.mistral-large-3-675b-instruct Mistral ❌ ✅ $4.00/$12.00 Advanced multimodal, long-context
moonshotai.kimi-k2.5 Moonshot ❌ ✅ $0.60/$3.00 Vision, language, code in one model

Prices: Per 1M tokens. Models with thinking/reasoning support additional compute costs during reasoning phase.

Tested but Incompatible Models

Some AWS Bedrock models were tested but are not supported due to technical limitations:

Model Family Reason for Incompatibility
GLM (Z.AI) Tool calling format incompatible with Converse API (expects string instead of JSON)
AI21 Jamba Severe rate limits (1-2 req/min) prevent reliable testing and production use
Meta Llama 3.3/3.1 Unstable tool call result processing, causes unexpected failures in multi-turn workflows
Mistral Magistral