Awesome AI Security Tools 
A curated list of public-source, research, and commercial tools for AI security and AI-assisted cybersecurity — autotriage, agent security, AI/ML supply chain, pentest agents, AI SAST, LLM-driven fuzzing, threat intelligence, SOC/SIEM triage, reverse engineering, LLM red-teaming, and more.
Type legend: 🟢 public source / open-source · 🔬 research (paper / benchmark / dataset / framework) · 🟠 commercial with open components · ⚠️ restrictive, non-commercial, or unclear/no license — check before use.
GitHub-hosted entries show static ★ stars and last-commit snapshots; refresh them with python3 scripts/update_github_metrics.py before release. Most recently refreshed entry: 2026-10-05. Hugging Face model entries show license, access, and artifact metadata. Ordering within a section favors flagship and actively maintained projects.
Contents
- Autotriage of Security Findings
- AI Agent & Coding-Agent Security
- AI/ML Supply Chain & Model Security
- Pentest & Red-Team Agents
- AI-Powered Recon & Narrow ML Tools
- AI-Powered SAST & Secure Code Review
- AI-Powered Threat Modeling
- LLM-Driven Fuzzing
- Threat Intelligence
- Log Analysis / SIEM / SOC Triage
- Reverse Engineering
- LLM Red-Teaming & Guardrails
- LLM Honeypots & Deception
- CTF / Exploit / Bug-Bounty Agents & Benchmarks
- Cloud / IaC / DFIR / OSINT / Phishing
- Related Awesome Lists
- Contributing
- Contact
- License
Autotriage of Security Findings
AI/LLM tools that triage, deduplicate, prioritize, or validate the output of scanners and finding sources.
- Metis 🟢 — Security-review framework that triages external SARIF findings using source-navigation evidence, optional C/C++ code graphs, and structured LLM decisions. (Arm) Caveat: hosted providers receive selected source and evidence and require credentials; local providers need separately provisioned weights, and indexing may use a separate embedding provider. Debug traces can contain sensitive content. Model verdicts are not proof of exploitability or grounds for automatic suppression. (★ 872 · updated 2026-09-30)
- nuclei-autotriage 🟢⚠️ — Two-stage LLM triage (falsifier + red-team pass) of Nuclei JSONL findings via OpenAI-compatible endpoints (vLLM/Ollama). (CyberOK) — note: restrictive personal/non-commercial EULA, not a permissive OSS license. (★ 1 · updated 2026-05-25)
- Related: agent-audit · asamm
- seclab-taskflow-agent 🟢 — YAML-driven taskflow agent framework for triaging CodeQL/SAST alerts and filtering false positives. (GitHub Security Lab) (★ 221 · updated 2026-08-17)
- Related: SigmaOptimizer
- honeyslop 🟢 — Code-canary decoys to triage AI-hallucinated ("slop") vulnerability reports flooding bug-bounty programs. (★ 97 · updated 2026-05-20)
- nano-analyzer 🟢🔬 — Minimal three-stage LLM pipeline (context → scan → skeptical triage) for zero-day discovery in C/C++. (AISLE) (★ 307 · updated 2026-04-14)
- SigmaOptimizer 🟢 — Generates, tests, and refines Sigma rules from real logs with false-positive checking. (★ 11 · updated 2025-08-01)
- Related: soctalk · seclab-taskflow-agent
- ai-soc-triage-assistant 🟢⚠️ — SOC alert triage assistant with prompt-injection guardrails, output validation, and MITRE ATT&CK mapping. (★ 0 · updated 2026-02-23)
See also: OpenAI's Aardvark research preview — public references exist, but there is no standalone installable repository to badge here.
AI Agent & Coding-Agent Security
Securing the AI agents themselves — auditing coding agents (Claude Code, Codex, OpenClaw), scanning skills / plugins / MCP manifests, and governance for agentic development. A fast-moving 2026 category, split below by role.
Scanners & Auditors
- Geiger 🟢 — Local CLI that inventories AI-agent installations, MCP configurations, hooks, and extensions, and reports credential-shaped values and configuration drift. Caveat: configuration inventory, not runtime monitoring or enforcement; exposure labels are heuristics, and coverage depends on known locations and readable files. Detector errors can leave gaps even when the strict-mode exit code is clean; explicit report options write local files. (★ 158 · updated 2026-09-21)
- Sources: Reviewed scan engine
- Guardana 🟢 — Verifies AI artifacts, live-system checks, and recorded agent traces, including evidence of unapproved side effects, outside the request path. Caveat: beta; trace checks depend on complete, trustworthy evidence and do not enforce tool calls. Built-in artifact checks are local, while active probes contact configured targets and optional collectors receive reports. Not a general SAST scanner or compliance certification. (★ 145 · updated 2026-10-05)
- agent-audit 🟢 — Forensic auditor for local AI coding agents (Claude Code, Codex CLI, OpenClaw) and project-surface scanner for repos shipping skills, plugins, and MCP manifests; 296 bundled rules across native + imported detector families, with optional LLM cross-verification. (CyberOK / S. Gordeychik) (★ 15 · updated 2026-07-15)
- Sources: asamm · ATR – Agent Threat Rules · aguara · Cisco AI Defense – skill-scanner
- Related: asamm · aguara · agentguard · agentic-radar · nuclei-autotriage
- AI-Infra-Guard 🟢 — Full-stack AI red-teaming platform covering OpenClaw security scan, agent scan, skills scan, MCP scan, AI-infra vulnerability scan, and LLM jailbreak evaluation. (Tencent Zhuque Lab) (★ 6,226 · updated 2026-09-10)
- SkillSpector 🟢 — Security scanner for AI-agent skills used by Claude Code, Codex CLI, Gemini CLI, and similar ecosystems; combines static analysis, AST/YARA/taint checks, optional LLM semantic review, MCP least-privilege/tool-poisoning checks, risk scoring, and SARIF/JSON/Markdown output. (NVIDIA) (★ 14,696 · updated 2026-08-15)
- Ramparts 🟢 — Rust scanner for MCP servers and agent-skill bundles with YARA rules, optional LLM analysis, OSV/CVE lookups, OWASP MCP Top 10 mapping, and SARIF/JSON/Markdown reports. (★ 96 · updated 2026-08-07)
- lintlang 🟢 — Local deterministic linter for AI-agent instructions, tool definitions, and system prompts that flags ambiguous descriptions, missing limits, conflicting directives, and schema gaps, with CLI and SARIF output. — note: heuristic static instruction/config review only; not runtime enforcement, prompt-injection detection, semantic-correctness proof, or a security guarantee. (★ 134 · updated 2026-10-02)
- Related: little-canary · rule-audit
- mcp-armor 🟢 — Local MCP security scanner with auto-discovery for agentic IDE configs, tool/resource/prompt inventory, prompt-injection checks, rug-pull and tool-poisoning detection, baseline drift monitoring, and JSON/Markdown reports. (Aira Security) (★ 120 · updated 2026-03-27)
- Related: SkillSpector · Ramparts · Cisco AI Defense – mcp-scanner
- aguara 🟢 — Single-binary static scanner (Go, no LLM) for AI-agent skills and MCP servers; multi-layer engine (pattern + NLP + taint tracking + rug-pull detection). Companion aguara-mcp exposes scanning as an MCP tool. (★ 86 · updated 2026-08-12)
- Related: aguara-mcp · agent-audit · Snyk Agent Scan · Cisco AI Defense – skill-scanner
- agent-scan 🟢 — Security scanner for AI agents, MCP servers, and agent skills; the successor path for the original Invariant Labs mcp-scan work. (Snyk) (★ 2,913 · updated 2026-08-13)
- inkog 🟠 — Commercial-backed static security scanner for AI agents across LangChain, LangGraph, CrewAI, AutoGen, and no-code workflows; Apache-2.0 CLI with proprietary deep-scan engine. (Inkog) (★ 28 · updated 2026-06-07)
- Related: Snyk Agent Scan · agentic-radar
- AgentShield 🟢 — Security scanner for AI-agent configurations, MCP servers, hooks, and tool permissions with CLI, GitHub Action, and app workflows. (★ 1,070 · updated 2026-07-22)
- Related: agent-audit · Snyk Agent Scan
- repo-forensics 🟢⚠️ — Offline scanner for AI-agent repos, skills, plugins, and MCP servers; license is PolyForm Noncommercial. (★ 155 · updated 2026-08-08)
- Related: agent-audit · aguara
- skill-scanner 🟠 — Scanner for agent skills combining YAML + YARA patterns, LLM-as-a-judge, and behavioral dataflow analysis (Codex / Cursor skill formats). (Cisco AI Defense) (★ 2,440 · updated 2026-08-04)
- Related: defenseclaw · aguara · Cisco AI Defense – mcp-scanner
- mcp-scanner 🟢⚠️ — Scanner for MCP servers and agentic tool surfaces, covering tools, prompts, resources, package risk, malware indicators, and deployment readiness. (Cisco AI Defense) (★ 1,037 · updated 2026-08-07)
- Related: Cisco AI Defense – skill-scanner · Snyk Agent Scan · aguara
- mcp-guardian 🟢 — JS/TS library and CLI for detecting prompt injection in MCP tool descriptions and pinning tool definitions. (★ 6 · updated 2026-07-29)
- Related: Cisco AI Defense – mcp-scanner · Snyk Agent Scan
- MCP Observatory 🟢🟠 — CI-native MCP-server testing tool for schema drift, safe attack simulation, record/replay verification, health scoring, and SARIF evidence before agents depend on a server. (KryptosAI) — note: the local evidence engine is open source; hosted telemetry intelligence, fleet workflows, and commercial ranking remain outside the package. (★ 176 · updated 2026-08-15)
- Related: Cisco AI Defense – mcp-scanner · skilltotal
- agentic-radar 🟠 — CLI security scanner for agentic workflows (LangGraph, CrewAI, n8n, etc.) — maps tools/data flows and flags risks. (SplxAI) (★ 1,036 · updated 2025-11-27)
- skilltotal 🟢 — Offline deterministic static scanner (regex + AST, no LLM, no account) for AI components — agent skills/plugins, MCP servers, npm & PyPI packages, and git repos; flags supply-chain risk, dangerous capabilities, prompt-injection surfaces, MCP tool poisoning/shadowing, and data-exfiltration paths, maps to the OWASP Agentic Skills Top 10, and emits JSON + SARIF 2.1.0. (skilltotal.ai) (★ 1 · updated 2026-08-17)
- Sunglasses 🟢 — Local input/content scanner for AI agents that checks prompts, files, media metadata, skills, and tool descriptions against pattern and mechanism-based prompt-injection, exfiltration, command-injection, and agent-threat rules. — note: early-stage project; published precision/recall benchmark is self-reported by the project. (★ 4 · updated 2026-08-14)
- Related: skilltotal · Armorer Guard
- trentclaw 🟠 — Client-side security auditor for OpenClaw deployments: applies pattern-based secret redaction locally, then uploads config/skill metadata and confirm-gated skill archives to Trent AI's API, which identifies misconfigurations, risky skills (prompt injection, permission escalation, data exfiltration), and chained attack paths. (Trent AI) — note: core detection runs server-side via the Trent AI API (requires an API key); the Apache-2.0 client collects OpenClaw config/skill metadata, applies pattern-based secret redaction locally, and uploads skill archives only after an explicit in-terminal confirmation. (★ 23 · updated 2026-07-27)
- A2A Security Scanner 🟢 — CLI and PyPI scanner for Agent-to-Agent (A2A) agent cards, source code, registries, and live endpoints using specification validation, YARA rules, heuristics, endpoint testing, and an optional LLM analyzer. (Cisco AI Defense) (★ 163 · updated 2026-04-16)
- Related: Cisco AI Defense – mcp-scanner · Agent Threat Rules
- Ship Safe 🟢🟠 — Local security CLI for application code, AI-agent and MCP configuration, secrets, dependencies, CI, and cloud/IaC surfaces, with deterministic checks, optional AI-assisted analysis, and SARIF output. — note: the MIT CLI works without an account for its core checks; optional AI modes can send selected context to the configured provider, while hosted dashboards and organization workflows are separate commercial features. Coverage and benchmark figures are project-reported. (★ 827 · updated 2026-08-24)
- Related: SkillTotal · agent-audit
- AgentSeal 🟠⚠️ — Agent-security CLI and runtime library for scanning MCP servers, skills, prompts, and configuration, plus local guard and monitoring workflows with optional model-assisted red teaming. — note: source-available under FSL-1.1-Apache-2.0 rather than an OSI-approved open-source license; offline guard and MCP/skill scanning are available locally, while red-team modes require local Ollama or configured provider credentials. Catalog and coverage figures are project-reported. (★ 343 · updated 2026-06-11)
- Related: agent-scan · AgentLock
- SlowMist Agent Security 🟢 — Security-review skill and workflow for auditing agent skills, MCP servers, repositories, URLs, and documents before installation or use. (SlowMist) — note: Markdown-based workflow skill executed by a compatible host agent, not a deterministic standalone scanner; findings depend on the selected model and its available tools. (★ 502 · updated 2026-04-17)
- Related: MCP-Security-Checklist · sast-skills
- Sandbox Probe 🟢 — Static Go probe that measures the effective filesystem, network, process, credential, and runtime capabilities exposed inside an AI-agent sandbox, then compares sandbox and host baselines. (ControlPlane) — note: boundary-measurement auditor, not an enforcement layer; some integration scripts can ask real agents to execute the probe, while deterministic model-free stubs are available for CI. (★ 25 · updated 2026-08-25)
- Related: Sandlock · AIO Sandbox
Frameworks, Rule Standards & Benchmarks
- RepoGuardBench 🟢🔬 — Repository prompt-injection benchmark with coding tasks, attack carriers, defenses, and separate scoring of attack attempts, observed outcomes, and task utility. Caveat: its subprocess workspace is not an OS sandbox: allowing Python or pytest does not prevent host-file access or networking. Use a separate disposable environment without secrets; marker-based outcomes and model behavior are not proof of production isolation. (★ 70 · updated 2026-10-03)
- Project CodeGuard 🟢⚠️ — Model-agnostic secure-coding rules and skills framework with translators for popular coding agents, validators, release artifacts, and an MCP server for centrally distributing the rules. (CoSAI / OASIS) — note: framework and ruleset, not a deterministic scanner or runtime enforcement boundary; repository content uses CC BY 4.0 rather than a conventional software license. (★ 338 · updated 2026-09-18)
- asamm 🔬 — Agentic SAMM — an OWASP SAMM extension for AI-driven development: an entry-point-based threat taxonomy plus 17 controls across 5 SAMM functions (Governance, Design, Implementation, Verification, Operations) with L1/L2/L3 maturity. License: CC BY-SA 4.0. (CyberOK / S. Gordeychik) (★ 17 · updated 2026-07-26)
- agent-threat-rules (ATR) 🟢 — Open, versioned, machine-readable detection rules for AI-agent threats (prompt injection, tool poisoning, MCP attacks, and skill compromise) — "Sigma for agents"; 768 rules across 10 categories with integrations for Microsoft AGT, Cisco AI Defense, MISP, OWASP, FINOS, and SigmaHQ. (★ 371 · updated 2026-08-17)
- Agent Governance Toolkit 🟢 — Multi-language toolkit for policy-enforced agent tool calls and audit records, with optional identity, MCP-gateway, sandboxing, reliability, and compliance components. (Microsoft) — note: official public preview; APIs and deployment patterns may change before general availability. (★ 5,962 · updated 2026-08-12)
- Related: ATR – Agent Threat Rules · ToolHive
- MCP-Security-Checklist 🟢 — Security checklist for MCP clients, servers, multi-MCP deployments, lifecycle controls, authz/authn, isolation, and crypto-specific MCP integrations. (SlowMist) (★ 835 · updated 2025-04-28)
- Anthropic-Cybersecurity-Skills 🟢 — Large community cybersecurity skill library for AI agents, mapped to MITRE ATT&CK, NIST CSF, MITRE ATLAS, D3FEND, and NIST AI RMF. — note: independent community project, not affiliated with Anthropic. (★ 28,060 · updated 2026-08-08)
- Related: sast-skills · Cisco AI Defense – skill-scanner
- Claude-BugHunter 🟢 — Claude Code / agent-skill bundle for authorized bug hunting and external red-team workflows across web, API, identity, cloud, recon, reporting, Burp MCP, slash commands, and the cbh CLI. — note: skill bundle and workflow knowledge base, not a standalone scanner. (★ 3,630 · updated 2026-08-17)
- Related: sast-skills · Anthropic-Cybersecurity-Skills
- AgentDojo 🟢🔬 — Benchmark environment for prompt-injection attacks and defenses in tool-using LLM agents. (★ 753 · updated 2026-06-02)
- Related: agent-audit · ATR – Agent Threat Rules
- Agent3Sigma-Canary 🟢🔬 — Sandboxed research framework for evaluating AI-agent security over complete execution trajectories, covering direct/indirect injection, skill and memory poisoning, and practical risk outcomes. (Ant Group) — note: research framework that requires Docker plus target and auxiliary LLM configuration; use only in controlled environments. (★ 35 · updated 2026-08-09)
- Agent Security Bench (ASB) 🟢🔬 — Official ICLR 2025 benchmark for evaluating attacks and defenses in LLM-based agents across ten scenarios, including direct and indirect prompt injection, memory poisoning, and defensive strategies. — note: research benchmark rather than a production control; reproducing evaluations requires configured target and evaluator models. (★ 285 · updated 2026-04-16)
- Related: AgentDojo · Agent3Sigma-Canary
- Skill-Inject 🟢🔬 — Benchmark for measuring prompt-injection vulnerabilities carried by agent skill files across Claude Code, Codex CLI, and Gemini CLI under multiple safety-policy conditions. — note: benchmark artifact that executes controlled malicious skill scenarios; run only in an isolated test environment with synthetic data and accounts. (★ 91 · updated 2026-07-01)
- Related: SkillSpector · AgentDojo
- AI Security Verification Standard (AISVS) 🔬⚠️ — Stable verification standard defining testable security requirements for AI applications across model lifecycle, supply chain, data handling, agentic systems, and MCP integrations. (OWASP) — note: security standard and checklist, not an executable scanner; share-alike terms apply to adapted material. (★ 429 · updated 2026-07-30)
- OWASP Agent Security Regression Harness 🟢 — Vendor-neutral harness for running repeatable agent and MCP abuse scenarios, evaluating policy assertions over execution traces, and emitting machine-readable regression results for local development and CI. (OWASP) — note: early OWASP Incubator project; it is a regression harness for known abuse cases, not a scanner, leaderboard, general benchmark, or guarantee of agent security. (★ 49 · updated 2026-07-27)
- Related: Agent Security Bench (ASB) · AgentDojo
Runtime Protection & Enforcement
- OpenAPPA 🟢 — Rust information-flow policy engine and agent-tool runtime that tracks trajectory labels and effects and binds approvals to actions using recorded events. (Archestra) Caveat: preview/RFC with evolving interfaces. Policy evaluation is deterministic for fixed policy and recorded evidence, not for whole-agent behavior; protection requires correct tool contracts, trusted labels, and complete mediation. Configured annotation authorities may use LLMs, and project-run benchmarks do not establish universal prompt-injection or exfiltration prevention. (★ 1,427 · updated 2026-10-05)
- Sources: Reviewed policy checks · Recorded-state projection
- Vetto 🟢 — Rust command wrapper that applies native OS containment to coding-agent subprocesses through Linux Landlock/namespaces and platform-specific sandbox backends. Caveat: guarantees vary by OS and selected tier: Linux FULL requires kernel support, while FS-ONLY lacks mount/PID/network namespaces and has descendant-cleanup limits. Allowed paths, environment values, and network destinations remain agent capabilities; this is not semantic prompt-injection protection. Optional telemetry is opt-in, and release binaries and escape resistance were not independently verified. (★ 45 · updated 2026-10-05)
- Sources: Reviewed threat model · Telemetry configuration
- Spring MCP Security 🟢 — Spring security components for MCP authentication, JWT audience validation, principal-bound sessions, and transport Origin/Host checks. (Spring AI Community) Caveat: match library and Spring AI versions and the supported transport/stack combination. Operators must configure issuers, scopes, and application authorization; these controls do not detect prompt injection. (★ 117 · updated 2026-09-20)
- yoloAI 🟢 — Coding-agent sandbox runner with multiple backends, credential-brokering integrations, and a copy/diff/apply workflow for reviewing workspace changes. Caveat: public beta; isolation depends on the backend and configuration. Select network isolation explicitly, avoid privileged modes, and review mounts; agents without broker support receive mounted credentials, and hosted agents still send context to their providers. (★ 216 · updated 2026-08-21)
- Jailoc 🟢 — Container launcher for OpenCode with configurable workspace mounts, network/DNS policy, resource limits, and secret references. (Seznam) Caveat: Docker and the agent image remain trusted dependencies; extra mounts, SSH-agent forwarding, optional Docker access, and supplied credentials expand the boundary. Container configuration is not a guarantee against escape or data disclosure to model providers. (★ 32 · updated 2026-09-24)
- Agent Sandbox (agentbox) 🟢 — Container-based coding-agent environment with proxy-mediated egress policies and server-side credential injection. Caveat: the reviewed firewall permits the detected Docker host subnet, not only the proxy, and has no IPv6 rules; proxy enforcement is inactive outside enforce mode. Review IPv6, neighboring services, writable mounts, and the IDE/control plane rather than assuming complete network or host isolation. (★ 208 · updated 2026-09-26)
- Crust 🟢⚠️ — Local Go gateway that extracts tool calls from supported model API responses and applies rules before forwarding them to an agent. Caveat: source-available under Elastic License 2.0, including managed-service restrictions, not a permissive OSS license. Only intercepted supported paths are covered; parse failures can pass the original body through. The sample forwards to OpenRouter, disables local telemetry, and leaves the storage encryption key empty; review provider disclosure and any enabled local records. (★ 440 · updated 2026-03-30)
- Agent Safehouse 🟢 — macOS launcher that assembles Seatbelt sandbox-exec profiles to scope coding agents' filesystem and integration access. Caveat: macOS-only hardening, not a VM or prompt-injection detector. The default profile permits inbound and outbound networking; exfiltration prevention is explicitly out of scope. Granted workspaces and integrations remain accessible, and launched agents retain their own credential, billing, and network behavior. (★ 2,092 · updated 2026-09-30)
- OpenShell 🟢 — Policy-governed runtime for autonomous and coding agents with container or microVM-backed sandboxes, filesystem/process/network controls, endpoint-bound credential injection, and audit logs. (NVIDIA) — note: pre-release runtime whose effective isolation depends on the selected compute driver and policy; Kubernetes and GPU paths are experimental. Anonymous operational telemetry is enabled by default but can be disabled at deployment time or compiled out. (★ 8,715 · updated 2026-09-21)
- Sources: NVIDIA agent-stack security model
- Numbat 🟢 — Endpoint-local visibility and detection for AI-agent activity across hooks, plugins, OTLP, and on-disk artifacts, with CEL rules, multi-step sequence detections, and forensic reconstruction. (Perplexity AI) — note: monitoring is the default posture; blocking is opt-in, limited to supported synchronous pre-action hooks, and the shipped rules are monitor-only until operators explicitly configure enforcement. (★ 1,069 · updated 2026-09-15)
- agentsh 🟢 — Execution-layer policy shell for AI agents that intercepts file, network, process, signal, and selected database activity and emits structured audit events. (Canyon Road) — note: the shell shim bypasses policy for non-TTY stdin unless
--forceis used, which is critical for headless agents and CI. Linux is the production target; native macOS enforcement is alpha and native Windows drivers are not yet production-ready. (★ 385 · updated 2026-09-15) - brood-box 🟢 — Experimental runner for coding agents in hardware-isolated microVMs with copy-on-write workspace snapshots, egress profiles, selective secret forwarding, and file-by-file review before applying changes. (Stacklok) — note: APIs and behavior are explicitly experimental;
workspace-mode=directbypasses snapshot isolation and writes directly to the workspace, so it is suitable only for already-trusted interactive work. (★ 72 · updated 2026-09-18) - Greywall 🟢 — Kernel-enforced filesystem, network, syscall, and command-policy wrapper for coding agents on Linux and macOS, with a separate traffic-observability mode and least-privilege profile generation. (Greyhaven) — note: deny-by-default applies to
greywall;greywatchis intentionally permissive and records rather than blocks activity. Verify the effective platform backend and selected mode before treating it as an enforcement boundary. (★ 298 · updated 2026-08-13) - nono 🟢 — Least-privilege sandbox for AI coding agents that isolates the agent and delegated tools with composable filesystem, network, credential-proxy, and command policies. (NoLabs) — note: APIs are still stabilizing ahead of the 1.0 release; review every pulled profile before use. (★ 3,687 · updated 2026-08-17)
- Related: microsandbox · ToolHive
- Coi 🟢 — Incus-based session manager for coding agents with protected workspace mounts, configurable network policy, optional host-side monitoring, and persistent or ephemeral Linux containers. — note: shared-kernel containers are not a microVM boundary; workspace writes reach the host and configured credentials can be copied into the container. Default
restrictednetworking allows public-internet egress and monitoring is off. The hardened profile enables monitoring and reduces kernel attack surface, but is not a general exfiltration or hostile-agent containment guarantee. (★ 733 · updated 2026-10-02)- Sources: Containment threat model · Reviewed defaults
- cplt 🟢 — Kernel-backed sandbox wrapper for AI coding agents that applies Seatbelt on macOS or Landlock and seccomp on Linux, content-pins approvals for repository policy, filters environment and resource access, and gates selected git and GitHub commands. (NAV (Norwegian Labour and Welfare Administration)) — note: no native Windows backend; the standard posture permits outbound HTTPS on port 443 and warns rather than blocks
git push, while stricter egress and push blocking require explicit configuration. Linux has documented limitations around Git metadata, localhost, and SSH-agent isolation, so review the threat model and effective policy for the target platform. (★ 114 · updated 2026-09-06)- Related: nono · microsandbox
- Arcjet Guard 🟢🟠 — JavaScript runtime guard for AI-agent tool calls and MCP handlers, with prompt-injection detection, sensitive-data detection/redaction, and custom local policy rules. (Arcjet) — note: open SDK packages integrate with Arcjet's hosted platform; assess the service, account, and data-processing requirements for the protections you enable. (★ 681 · updated 2026-08-15)
- ToolHive 🟢 — Platform for running MCP servers in isolated containers with per-request identity/access policy, registry and gateway workflows, audit logs, Kubernetes operator support, and observability hooks. (Stacklok) (★ 2,019 · updated 2026-08-14)
- Related: microsandbox · defenseclaw
- Pipelock 🟢 — AI-agent firewall and verifiable egress-control layer mediating HTTP, WebSocket, CONNECT, MCP, and A2A traffic to detect prompt injection, secret exfiltration, SSRF, and suspicious outbound actions. — note: open-source core is Apache-2.0; commercial reporting/features are also advertised. (★ 796 · updated 2026-08-17)
- Related: ToolHive · mcp-context-protector
- node9 🟢 — Local policy and human-approval gate for supported coding-agent tool calls routed through agent hooks or an MCP wrapper, with credential-path checks, secret detection, and audit logging. (Node9) — note: cooperative hooks/MCP enforcement, not process isolation or a complete credential boundary; scripts are judged by their command line rather than their contents, and the egress allowlist is opt-in. Local use needs no account; optional hosted team features can transmit operational data.
node9 initasks before sending an install ping, with Yes preselected; the payload includes a persistent machine ID, detected agents, OS, and version, so it is pseudonymous rather than anonymous. (★ 216 · updated 2026-09-27) - SourceryKit 🟠⚠️ — Python SDK for agent guardrails that intercepts outbound HTTP calls, enforces trusted-endpoint policies, logs requests, and checks agent handoff claims against a Provably backend before propagation. (ProvablyAI) — note: BSL-1.1 licensed; requires Provably backend/API credentials and database setup. (★ 18 · updated 2026-08-12)
- Related: Pipelock · mcp-context-protector
- emisar 🟠⚠️ — Agent-infrastructure control plane that exposes declared, typed actions through MCP, applies policy and approval gates before dispatch, revalidates calls on an outbound-only host runner, and records separate control-plane and host audit trails. — note: runner, MCP bridge, and packs are Apache-2.0; the hosted portal/control plane is BSL-1.1 and converts to Apache-2.0 on 2029-07-26. Connecting a runner requires an emisar account and outbound HTTPS to the control plane. (★ 416 · updated 2026-08-17)
- mcp-context-protector 🟢 — MCP security wrapper that sits in front of downstream MCP servers, scans tool responses with guardrail providers, and supports quarantine/review workflows for desktop and coding-agent MCP configs. (Trail of Bits) (★ 222 · updated 2026-02-13)
- Related: Cisco AI Defense – mcp-scanner · ToolHive · Pipelock
- MCP Defender 🟢⚠️ — Desktop app that proxies MCP tool-call requests and responses for Cursor, Claude, VS Code, and Windsurf, checks intercepted traffic against signatures, and prompts users to allow or block suspicious calls. — note: AGPL-3.0 licensed; project has been acquired by Docker. (★ 255 · updated 2026-06-05)
- Related: ToolHive · mcp-context-protector
- MCP Gateway 🟢 — Plugin-based MCP gateway that proxies configured MCP servers, sanitizes sensitive request/response data, supports guardrail plugins such as basic masking and Presidio, and runs a server reputation/risk check before loading MCP servers. (Lasso Security) (★ 385 · updated 2026-01-22)
- Parallax 🟢 — Rust runtime policy engine for AI agents: evaluates lifecycle events with regex, keyword, Sigma, CEL, and SQL rules to block or redact prompt injection, data exfiltration, dangerous tool calls, and secret leakage. — note: early-stage project with limited adoption signal. (★ 35 · updated 2026-06-05)
- Related: Pipelock · Armorer Guard
- Armorer Guard 🟢 — Local Rust scanner and MCP proxy for AI-agent prompt injection, credential leakage, exfiltration, and risky tool-call arguments, with structured reasons and no scanner network calls. — note: young project with limited independent adoption signal. (★ 42 · updated 2026-08-09)
- Related: agentguard · ATR – Agent Threat Rules
- onecli 🟢 — Credential gateway and encrypted vault for AI agents; injects real API credentials at the gateway so agents only see placeholder keys. (★ 3,104 · updated 2026-07-31)
- Related: agentguard · defenseclaw
- microsandbox 🟢 — Local-first, microVM-backed programmable sandboxes for AI agents with SDKs, CLI, MCP support, and rootless hardware isolation. (★ 7,569 · updated 2026-08-17)
- agentguard 🟢 — Real-time security layer for coding agents: hooks scan every new skill, block dangerous actions before execution, run daily posture patrols, and track which skill triggered each action (incl. Web3-specific checks). (★ 456 · updated 2026-06-25)
- Related: agent-audit · defenseclaw
- AgentAegis 🟢 — OpenClaw security plugin that observes or blocks selected prompt, tool-call, memory, exfiltration, and protected-path risks through configurable agent lifecycle hooks. (Ant Group) — note: the code defaults blocking defenses to
enforce; the README'sobserveconfiguration is an optional staged-rollout example. Coverage depends on enabled hooks, modes, and protected assets; this is OpenClaw-specific defense in depth, not a general sandbox or injection-proof boundary, and published effectiveness demonstrations are project-operated. (★ 199 · updated 2026-06-26)- Sources: Configuration defaults
- defenseclaw 🟠 — Enforcement and evidence layer for agentic deployments: static CodeGuard checks, sandboxing, registry ingestion with SSRF guards, and audit/observability. (Cisco AI Defense) (★ 819 · updated 2026-08-17)
- Related: Cisco AI Defense – skill-scanner · agentguard
- clawsec 🟢⚠️ — Security skill suite for OpenClaw-family agents; AGPL-3.0 licensed. (Prompt Security) (★ 1,083 · updated 2026-08-05)
- Related: agentguard · Cisco AI Defense – skill-scanner
- AgentLock 🟢🔬⚠️ — Pre-action authorization gate for LLM agent tool calls that decides from session provenance rather than content, with deny-by-default tool permissions, parameter lineage, Ed25519 signed receipts, and a hash-chained audit log; AGPL-3.0 licensed with commercial options. — note: evaluated on AgentDojo with predictions pre-registered before the runs; the published results include a suite where the defense costs more utility than the attack it prevents. (★ 19 · updated 2026-08-12)
- h5i 🟢 — Local Rust CLI for auditable coding-agent workspaces: per-agent worktrees with sandbox policies, provenance capture, peer review, neutral verification, secret/prompt-injection audit signals, and refs/h5i/* run metadata. — note: security-adjacent agent-workspace governance tool, not a vulnerability scanner or VM-equivalent sandbox. (★ 531 · updated 2026-08-16)
- Related: microsandbox · defenseclaw
- DvalinCode 🟢 — Local-first AI coding agent with runtime governance controls: org/repo policy gates for tools, models, MCP servers, paths, and commands, plus provider/shell/MCP egress controls and hash-chained audit logs. — note: young project with limited independent adoption signal. (★ 112 · updated 2026-08-17)
- Related: h5i · Pipelock · Armorer Guard
- TAP 🟢🟠 — Credential-isolation proxy and MCP server for AI agents: agents send placeholder credentials, TAP injects real secrets server-side after per-action policy checks, with optional human approval on sensitive calls. (human.tech) — note: Apache-2.0 runtime is self-hostable, but the hosted dashboard and managed-service deployment glue are proprietary; self-hosting puts credential/key isolation and policy-engine hardening on the operator. (★ 12 · updated 2026-07-23)
- Agent Memory Guard 🟢 — Runtime middleware for AI-agent memory reads and writes, screening prompt injection, memory poisoning, secret/PII leakage, protected-key tampering, and size anomalies before persisted memory is reused. (OWASP) — note: OWASP Incubator project; published benchmark numbers are project-reported and should be independently reproduced before production enforcement. (★ 125 · updated 2026-08-16)
- AIO Sandbox 🟢⚠️ — All-in-one Docker workspace for AI agents with browser, shell, file, code-execution, MCP, and VSCode interfaces, plus API-key/JWT controls and private-deployment guidance. — note: the public repository ships SDKs, integrations, and docs rather than the core runtime service; official Chromium-enabled deployments use
seccomp=unconfined. Treat it as a trusted execution environment, not a hardened isolation boundary; use separate VMs or a hardened runtime plus network policy for hostile workloads. (★ 5,727 · updated 2026-08-17) - Agentgateway 🟢 — Agent-native proxy and gateway for MCP and A2A traffic with OAuth/JWT/API-key authentication, CEL-based RBAC policies, TLS, rate limiting, and OpenTelemetry observability. (★ 4,385 · updated 2026-08-14)
- Related: MCP Gateway · Pipelock
- Kubernetes Agent Sandbox 🟢 — Kubernetes CRDs and controllers for isolated, stateful singleton agent workloads, delegating low-level isolation to configured runtimes such as gVisor or Kata Containers. (Kubernetes SIG Apps) — note: sandbox orchestrator, not an isolation runtime itself; security depends on the selected RuntimeClass, network policy, and workload configuration. (★ 3,544 · updated 2026-08-16)
- Sources: Kubernetes announcement
- Related: AIO Sandbox
- Prismor 🟢 — Self-hosted runtime control plane for coding agents with pre-tool-call hooks, policy-driven observe/approve/block decisions, an MCP gateway, secret and egress controls, and tamper-evident audit evidence. (★ 291 · updated 2026-08-16)
- Related: Armorer Guard · AgentLock
- Gram 🟢🟠⚠️ — Open-source stack behind Speakeasy's AI control plane that centrally manages MCPs, Skills, and Assistants with granular permissions, policy enforcement, threat detection, and observability. (Speakeasy) — note: AGPL-3.0 copyleft license; Gram is also available as Speakeasy's hosted commercial control plane, while local development uses a full Docker/Mise stack. (★ 266 · updated 2026-09-09)
- tirith 🟢⚠️ — Terminal guard for developers and AI coding agents that intercepts homograph and terminal-injection tricks, obfuscated execution chains, pipe-to-shell patterns, credential exfiltration, and malicious skill/config files. — note: AGPL-3.0 with a separate commercial license; shell interception is a host-side guard, not a substitute for sandboxing or least-privilege tool access. (★ 2,664 · updated 2026-08-13)
- ADR 🟢🔬 — Agentic AI Detection and Response system combining cross-client agent telemetry, ADR-Bench security scenarios, and a dual-agent detector for suspicious intent, tool use, and execution traces. (Uber) — note: deployed at Uber and published with an MLSys 2026 paper; the open release includes the Sensor, benchmark, and Detector, but not ADR Prevention or the offline ADR Explorer. Default detector configurations require model-provider credentials. (★ 1,442 · updated 2026-08-16)
- xaidr 🟢 — In-process runtime security sensor for AI agents that inspects input, tool calls, output, and agent-to-agent envelopes inside the agent process, with structured shell-command classification, YAML policy, privilege tiers, an opt-in circuit breaker, and OpenTelemetry export. (Delphi Security) — note: early-stage in-process sensor, not an isolation boundary; monitor mode is the default and unexpected scan failures fail open while emitting degraded telemetry. Published detection and false-positive figures are project-reported on its committed corpus. (★ 22 · updated 2026-08-17)
- Related: agentguard · defenseclaw
- Agentmetry 🟢 — Local-first flight recorder for AI coding agents and MCP servers that writes a hash-chained JSONL trail with RFC 6962 Merkle roots, applies MITRE-mapped sequence detection across a session, attests every 300s which agent surfaces are covered, uncovered, absent or unknown, and forwards to Splunk HEC, Elastic ECS, Google SecOps UDM or CloudEvents. — note: young project with limited independent adoption signal; it records and detects rather than blocks, and its README states that hooks are cooperative and that tamper-evident is not the same as attributable. (★ 12 · updated 2026-08-24)
- piighost 🟢 — Local runtime pseudonymization layer that replaces detected PII with stable placeholders before model calls and restores the original values in responses and selected tool arguments, with integrations for LangChain, Pydantic AI, LlamaIndex, and OpenAI-compatible clients. — note: reversible de-identification, not anonymization; the placeholder mapping remains sensitive, restored tool arguments can contain the original PII, and optional LLM-backed detectors may send content to the configured model provider. (★ 11 · updated 2026-08-23)
- HOL Guard 🟢🟠 — Local-first runtime security layer for coding agents that evaluates commands, package installs, skills, MCP configuration, and sensitive actions through policy, approval, evidence, and audit workflows. (Hashgraph Online) — note: integrations are cooperative hooks rather than a hard isolation boundary, and documented hook failures can fail open; synchronized evidence, team policy, and fleet workflows use the optional hosted service. (★ 466 · updated 2026-08-24)
- StackOne Defender 🟢 — Offline TypeScript runtime guard for indirect prompt injection in tool results, using bundled ONNX classifiers, deterministic checks, sanitization, and allow/block verdicts. (StackOne) — note: blocking high-risk results is opt-in by default; missing Tier-2 dependencies can fall back to simpler Tier-1 detection unless strict mode is enabled, so configure fail behavior explicitly for enforcement use. (★ 117 · updated 2026-08-19)
- Related: LLM Guard · little-canary
- Earl 🟢 — Capability proxy for AI agents that exposes approved operation names while keeping request templates and credentials outside the model, with HCL policy, audit, and egress controls. (Mathematic) — note: operation templates and policy configuration form part of the trusted boundary and require review; the proxy reduces credential and request-construction exposure but is not a sandbox for the agent process. (★ 113 · updated 2026-08-23)
- Related: onecli · agentgateway
- Bifrost 🟢🟠 — Apache-licensed AI and MCP gateway with multi-provider routing, virtual-key access controls, budgets, rate limits, MCP aggregation, OAuth, automatic fallbacks, and load balancing. (Maxim) — note: model-provider credentials and proxied request/response data are sensitive; restrict and authenticate the gateway and admin interface, use TLS, and load only trusted plugins. Content guardrails, RBAC/SSO, clustering, and other advanced governance controls require Bifrost Enterprise. (★ 7,548 · updated 2026-08-25)
- Related: Portkey AI Gateway · Agentgateway
- Portkey AI Gateway 🟢🟠 — Open AI gateway with provider routing, fallback and retry controls, guardrail integrations, observability, and MCP traffic support for model and agent applications. (Portkey) — note: the gateway is general infrastructure rather than a standalone security scanner; model-provider credentials are required, and some RBAC, analytics, and managed control-plane capabilities belong to Portkey's hosted or enterprise products. (★ 12,815 · updated 2026-05-25)
- Related: agentgateway · Pipelock
- Casbin AI Gateway 🟢 — Local gateway and policy layer for model-provider and MCP traffic, combining provider-key mediation, access control, prompt records, and configurable request handling. (Apache Casbin) — note: binds to localhost by default because its local UI can exercise administrative operations and the relay can access stored provider keys; do not expose it beyond a trusted host without adding authentication and reviewing prompt-recording settings. (★ 570 · updated 2026-08-24)
- Related: agentgateway · Portkey AI Gateway
- Adrian 🟢🟠 — Runtime monitoring and intervention layer that correlates agent actions with available reasoning traces through Python and TypeScript SDKs, a Claude Code integration, and managed or self-hosted deployments. (Secure Agentics) — note: the bundled self-hosted stack requires Docker, substantial disk space, and an NVIDIA GPU for the recommended local classifier; the README's
+35%and4xheadline extrapolates cited third-party research rather than an Adrian benchmark. (★ 552 · updated 2026-08-20)- Related: Agentic Radar · Armorer Guard
- Sandlock 🟢 — Unprivileged Linux process sandbox using Landlock, seccomp-BPF, and seccomp user notification to apply per-process filesystem, network, syscall, and execution policies without a container or VM. (Multikernel) — note: Linux-kernel process isolation rather than prompt-injection detection; protection depends on available Landlock and seccomp features and is not a multi-tenant VM boundary. (★ 381 · updated 2026-08-23)
- Related: Sandbox Probe · AIO Sandbox
- Lunar 🟢🟠 — API and MCP gateway combining outbound-traffic visibility, policy enforcement, rate limits, retries, circuit breakers, and centralized MCP server aggregation for agent workloads. (Lunar.dev) — note: the repository is active but its latest formal GitHub release is from 2024; the README positions the open core for non-production or personal use and directs production deployments to commercial platform tiers. (★ 485 · updated 2026-08-28)
- Related: agentgateway · Bifrost
- Norviq 🟢 — Kubernetes policy enforcement point for LLM agent tool calls that content-hash pins each tool definition at discovery and evaluates every tools/call against OPA/Rego policy before the upstream MCP server receives it. — note: ships in audit mode — the installed baseline records what it would have refused and lets the call proceed, and all shipped policy presets default to allow, so enforcement starts with the first rule an operator writes. The discovery-time description scanner is a documented heuristic and the project publishes the payloads that defeat it. Young project with limited independent adoption signal. (★ 24 · updated 2026-09-01)
- Related: Prismor · mcp-context-protector
- sofagent 🟢 — Commit-time audit and governance suite for AI coding agents that scans git diffs against deterministic rules, records local audit history, and exposes MCP tools for governance aggregation. — note: HMAC signing is optional, while local hooks, configuration, and key material remain accessible to same-user agents; the default setup is not fail-closed and Git hooks can be bypassed, so treat the history as local audit evidence rather than a hardened tamper-proof boundary. (★ 42 · updated 2026-09-03)
- Related: Pipelock
- Tenuo 🟢🟠 — Capability-based authorization for agent tool calls using signed, holder-bound warrants that constrain tools and arguments, verify locally, and narrow authority across delegation. — note: the Rust core, SDKs, and authorizer sidecar are Apache-2.0; the managed control plane is commercial. Enforcement requires trusted verification on every effecting tool path, outside agent control when bypass is possible; this is not a sandbox. The stateless core permits identical-call replay within its proof-of-possession window, so state-changing operations need application-level nonce or idempotency controls. (★ 100 · updated 2026-10-03)
- Related: Agentgateway · AgentLock
AI/ML Supply Chain & Model Security
Tools for securing model artifacts, serialized ML files, AI/ML supply-chain surfaces, and malicious-package detection datasets/benchmarks.
- MindArmour 🟢🔬 — MindSpore toolkit for adversarial robustness and model-privacy research, including gradient-based targeted and untargeted attacks. (MindSpore) Caveat: requires a compatible MindSpore, hardware, model, and dataset stack; package metadata points to Gitee for source and issue tracking. Model-specific evaluations are not robustness guarantees. (★ 92 · updated 2026-02-25)
- FedREDefense 🟢🔬 — ICML 2024 research implementation of federated-learning model-poisoning defense using reconstruction error of client updates. Caveat: an experimental training/evaluation pipeline, not a production federation security layer; GPU, datasets, and attack configuration affect results. Consult the published erratum for the FLTrust baseline rather than copying the original comparison figures. (★ 31 · updated 2026-09-16)
- Related: FLDetector poisoning-detection research (historical; no root license, legacy MXNet 1.9.1) · FL-WBC client-side poisoning defense (historical; no root license, legacy PyTorch 1.2) · NDSS21 model-poisoning research (historical; no root license, legacy datasets/frameworks) · ModelPoisoning FL research (historical; no root license, legacy TensorFlow 1.8)
- ModelAudit 🟢 — Model-artifact scanner with format-specific static parsers, a Rust-backed pickle-analysis companion, metadata checks, and JSON/SARIF reporting. (Promptfoo) Caveat: installed usage enables telemetry by default; disable with
PROMPTFOO_DISABLE_TELEMETRY=1orNO_ANALYTICS=1. Local scans need no hosted LLM, while remote sources need network access and sometimes credentials. Reports can retain raw secrets; optionalmetadata --trust-loaderscan deserialize artifacts. A clean static scan does not establish model safety. (★ 76 · updated 2026-10-04) - AI BOM 🟢 — Inventories models, agents, tools, MCP clients and servers, datasets, prompts, guardrails, secrets, and cloud AI resources, with CycloneDX 1.6 output and policy workflows. (Cisco AI Defense) — note: core inventory is distinct from the narrower model-file AIsbom already listed; the extended
analyzepipeline requires an external LLM provider and can send analysis context off-host. Optional sanitized Galileo telemetry is separately configurable. (★ 110 · updated 2026-09-17)- Sources: Cisco announcement
- Related: AIsbom
- Fraim 🟢 — Framework for AI-powered security workflows including LLM SAST and IaC analysis with SARIF/HTML output. (★ 160 · updated 2025-12-01)
- Related: sast-skills
- Adversarial Robustness Toolbox (ART) 🟢 — Flagship machine-learning security library for evaluating and defending models against evasion, poisoning, extraction, and inference attacks across major ML frameworks. (LF AI & Data / IBM) (★ 6,179 · updated 2025-11-13)
- Foolbox 🟢 — Classic Python toolbox for generating adversarial examples and benchmarking robustness of PyTorch, TensorFlow, and JAX models. (★ 2,972 · updated 2024-03-04)
- Related: Adversarial Robustness Toolbox
- modelscan 🟢 — Scans ML model files for unsafe serialization patterns and embedded code, with a focus on model serialization attacks. (Protect AI) (★ 762 · updated 2026-02-18)
- Related: Fickling · picklescan · ai-exploits
- Fickling 🟢 — Python pickle decompiler, rewriter, and static analyzer for inspecting and detecting malicious pickle/PyTorch payloads. (Trail of Bits) (★ 662 · updated 2026-08-13)
- Related: modelscan · picklescan
- picklescan 🟢 — Lightweight CLI/library for detecting suspicious Python pickle operations in ML and model artifacts. (★ 418 · updated 2026-07-01)
- AIsbom 🟢 — AI software bill of materials tooling for AI/ML supply-chain inventory and provenance metadata. (★ 76 · updated 2026-08-17)
- Related: modelscan · model-provenance-kit
- model-provenance-kit 🟢 — Toolkit for model-family provenance and fingerprinting across model weights, tokenizers, and architecture signals. (Cisco AI Defense) (★ 101 · updated 2026-08-12)
- Related: AIsbom
- pickle-fuzzer 🟢 — Structure-aware fuzzer for pickle scanners, useful for hardening tools such as modelscan, Fickling, and picklescan. (Cisco AI Defense) (★ 17 · updated 2026-08-03)
- Related: modelscan · Fickling · picklescan
- Medusa 🟢⚠️ — AI-first security scanner for AI/ML repos, agents, and MCP surfaces; AGPL-3.0 licensed. (Pantheon Security) (★ 964 · updated 2026-06-24)
- Related: agent-audit · modelscan
- PrivacyRaven 🟢🔬 — Privacy-testing library for deep-learning systems, covering model extraction and membership-inference style attacks. (Trail of Bits) — note: archived/hiatus project, but still a useful reference implementation. (★ 214 · updated 2025-09-05)
- Related: Adversarial Robustness Toolbox
- gym-malware 🟢🔬 — OpenAI Gym environment for reinforcement-learning agents that mutate PE malware to evade static ML malware detectors. (★ 636 · updated 2018-06-15)
- deep-pwning 🟢🔬 — Historical "Metasploit for machine learning" framework for experimenting with adversarial robustness of ML models. (★ 571 · updated 2022-05-17)
- open-malicious-code-benchmark 🟢🔬 — OMCBench benchmark suite for malicious-code/package detection: labeled Python and JavaScript package archives, common runners, and published precision/recall/F1 metrics. (False Positive Community) — note: evaluates an unreleased commercial ML detector (MOLOT / PT Application Inspector) alongside open-source baselines. (★ 17 · updated 2026-06-09)
- Related: GuardDog · OSSGadget · malicious-code-ruleset · bandit4mal
- malicious-software-packages-dataset 🟢🔬 — Human-vetted dataset of malicious software packages across npm, PyPI, IDE extensions, and AI Skills, useful for detector training and evaluation. (Datadog Security Labs) — note: contains real malware samples; Datadog notes selection bias because many samples were identified by GuardDog. (★ 372 · updated 2026-08-17)
- Related: GuardDog · pypi_malregistry
- GuardDog 🟢 — CLI for detecting malicious PyPI, npm, Go, RubyGems, GitHub Actions, and VSCode extension packages using Semgrep rules and package-metadata heuristics. (Datadog) (★ 1,184 · updated 2026-08-14)
- Related: malicious-software-packages-dataset · Packj
- package-analysis 🟢🔬 — Sandboxed static/dynamic analysis pipeline for open-source packages, capturing filesystem, process, and network behavior and publishing data for malicious-package research. (OpenSSF) (★ 903 · updated 2026-07-21)
- Related: malicious-packages · package-feeds
- malicious-code-ruleset 🟢 — Focused Semgrep ruleset for malicious-code patterns such as dynamic execution and obfuscation, used as an OMCBench baseline. (Apiiro) (★ 149 · updated 2025-02-24)
- Related: open-malicious-code-benchmark
- pypi_malregistry 🔬⚠️ — ASE'23 / USENIX Security'26 malicious-PyPI dataset with more than 10k malicious package versions. — note: no LICENSE file found and the repository contains malware samples; handle in an isolated environment. (★ 129 · updated 2026-07-21)
- Related: malicious-software-packages-dataset
- Activation-based Model Scanner (AMS) 🟢🔬 — PyPI scanner that uses safety-related activation fingerprints to detect degraded or removed safety training and compare an open-weight model with a known baseline. (Google Cloud Platform) — note: unofficial and unsupported Google research-derived project; GPU execution is recommended, calibration covers a limited set of model families, and the documented method can miss modifications that preserve measured safety directions. (★ 32 · updated 2026-08-27)
- Related: AASE research · model-provenance-kit
Pentest & Red-Team Agents
Autonomous and semi-autonomous AI agents for penetration testing, exploitation, and attack simulation.
- Cairn 🟢🔬⚠️ — Fact-intent graph coordinator that schedules Claude Code, Codex, or Pi workers for penetration-testing and CTF exploration through a shared persistent blackboard. Caveat: workers can execute arbitrary tools; Docker mode exposes the host Docker socket and local mode has no OS sandbox. Configured providers receive target context; use a dedicated isolated host. Root license is AGPL-3.0; README describes personal/educational use and offers a separate commercial license for use without AGPL obligations. (★ 3,198 · updated 2026-09-07)
- PentestGPT 🟢🔬 — The original USENIX'24 LLM pentest agent; re-released as an autonomous pipeline with strong benchmark results. (★ 14,906 · updated 2026-07-14)
- PentAGI 🟢 — Fully autonomous multi-agent pentest framework with Docker sandboxing. (VXControl) (★ 21,860 · updated 2026-08-06)
- CAI – Cybersecurity AI 🟢🟠 — Modular, bug-bounty-ready agent framework supporting 300+ LLM models. MIT for research; separate commercial license for production/on-prem. (Alias Robotics) (★ 9,745 · updated 2026-07-14)
- Strix 🟢 — Autonomous "AI hackers" that dynamically run code and validate vulnerabilities with PoCs (Apache-2.0). (★ 53,569 · updated 2026-08-17)
- hackingBuddyGPT 🟢🔬 — Minimal (~50 LOC) research framework for LLM-driven Linux priv-esc and web pentesting (FSE'23). (★ 1,209 · updated 2026-08-10)
- Nebula 🟢🟠 — AI pentesting CLI assistant with local-LLM support (Llama-3.1, Mistral, DeepSeek). (★ 1,087 · updated 2026-07-26)
- HexStrike-AI 🟢 — MCP server exposing 150+ security tools (nmap, gobuster, nuclei, …) to AI agents (MIT). (★ 11,116 · updated 2026-08-03)
- Deep Eye 🟢 — AI-assisted penetration-testing scanner that orchestrates multiple LLM providers for payload generation, 45+ vulnerability checks, CVE/RAG-assisted testing, AI triage, scan diffing, browser automation, proxying, and multi-format reports. — note: MIT-licensed; authorized use only. Heavy runtime surface: optional browser automation/proxying, plaintext API-key config, plugins with full OS access, and pickle model files called out in SECURITY.md. (★ 1,972 · updated 2026-08-14)
- Related: HexStrike-AI · pentest-ai
- Burp Suite MCP Server 🟢⚠️ — Official Burp Suite extension exposing Burp to AI clients through MCP. (PortSwigger) — note: GPL-3.0 licensed. (★ 1,071 · updated 2026-08-12)
- pentest-ai 🟢 — Offensive-security MCP server with 200+ wrapped tools, specialist agents, and OWASP-oriented probes for authorized testing. (★ 1,598 · updated 2026-08-17)
- Related: pentest-ai-agents
- Transilience Community Tools 🟢 — Claude Code skill and coordination-role bundle with supporting tools for authorized penetration testing, reconnaissance, source review, and AI threat testing. (Transilience AI) — note: agent-interpreted workflows, not a deterministic scanner or enforcement boundary; results depend on the configured coding agent/model, and published CTF results are maintainer-run after iterative benchmark tuning. The Docker helper uses host networking and
--dangerously-skip-permissions, whileenv-reader.pyprints requested secrets. Review and pin content, use disposable environments and scoped test credentials, and treat transcripts and context sent to model providers as sensitive. (★ 555 · updated 2026-07-29)- Sources: Container helper · Credential helper
- pentest-ai-agents 🟢 — Collection of Claude Code offensive-security subagents for authorized penetration-testing research. (★ 2,132 · updated 2026-08-16)
- Related: pentest-ai
- DarkMoon 🟢⚠️ — Autonomous AI penetration-testing platform that orchestrates specialized web, AD, Kubernetes, CMS, and framework agents through an MCP-controlled Docker toolbox with local privacy-tokenization for sensitive target data. — note: GPL-3.0 licensed; heavy Docker/LLM stack, use only for authorized testing. (★ 843 · updated 2026-08-06)
- Related: PentAGI · HexStrike-AI · pentest-ai
- T3MP3ST 🟢⚠️ — Autonomous offensive-security meta-harness that wraps local or API-backed coding agents into a multi-agent recon-to-exploit workflow with MCP/API, War Room UI, tool arsenal, and committed benchmark artifacts. — note: very new AGPL-3.0 project with bold benchmark claims; use only for authorized testing and verify independently before operational use. (★ 5,596 · updated 2026-08-12)
- Related: PentAGI · HexStrike-AI · pentest-ai
- Shannon 🟢🟠⚠️ — White-box autonomous AI pentester with strong XBOW-benchmark results. Shannon Lite is AGPL-3.0; Shannon Pro is commercial. (★ 46,882 · updated 2026-08-12)
- AIDA 🟢⚠️ — Model-agnostic autonomous pentest agent running inside an isolated Docker environment; AGPL-3.0 licensed. (★ 471 · updated 2026-07-19)
- HackSynth 🟢🔬⚠️ — Planner/summarizer LLM-agent framework for autonomous penetration testing and benchmark evaluation; AGPL-3.0 licensed. (★ 314 · updated 2025-06-24)
- VulnBot 🟢🔬 — Multi-agent collaborative penetration-testing framework with RAG support. (★ 185 · updated 2025-04-07)
- PentestAgent 🟢 — Black-box AI pentest framework with MCP, multi-agent spawning, and persistent sessions. (★ 2,958 · updated 2026-08-04)
- cyber-security-llm-agents 🟢⚠️ — AutoGen-based agents for cybersecurity tasks (shown at RSAC 2024). (NVISO) (★ 388 · updated 2024-05-07)
- Pentest-Swarm-AI 🟢 — Swarm-intelligence multi-agent pentest with stigmergic blackboard coordination (Go). (★ 2,205 · updated 2026-08-04)
- hackGPT 🟢⚠️ — LLM offensive-security toolkit. (★ 1,199 · updated 2026-08-12)
- ShiftGrid 🟢 — Prompt engine that turns Claude Code into a transparent, human-in-the-loop pentester, structuring engagements through checklists, observations, and notes exposed via an agent-facing API. — note: early-stage local Docker application with no built-in authentication; keep its API and UI ports bound to localhost. (★ 39 · updated 2026-08-02)
- BugTraceAI 🟢⚠️ — Self-hosted autonomous web-application security scanner that combines reconnaissance, specialist exploit agents, Go fuzzers, and Playwright validation to produce evidence-backed findings. (BugTraceAI) — note: AGPL-3.0 licensed and beta; use only for authorized testing. Requires an LLM provider or local Ollama endpoint and a substantial Docker/Playwright/Go runtime. (★ 174 · updated 2026-07-30)
- Related: Project overview · Web dashboard · Docker launcher
- HunterX 🟢 — AI-assisted offensive security engine that orchestrates reconnaissance, security-tool coordination, vulnerability detection and validation, evidence collection, and report-ready findings in one workflow. (NullC0d3) — note: early-stage and tool-orchestration-heavy; run only in an isolated, authorized assessment environment. (★ 12 · updated 2026-08-16)
- MCP Security Hub 🟢 — Collection of Dockerized MCP servers that expose offensive-security tools such as Nmap, Nuclei, SQLMap, Ghidra, Hashcat, and related assessment utilities to MCP-capable assistants. (FuzzingLabs) — note: orchestration and wrapper collection rather than a security boundary; its containers invoke offensive tools and must be used only in isolated, explicitly authorized environments. (★ 761 · updated 2026-04-08)
- Related: pentest-ai · Burp Suite MCP Server
- Forefy .context 🟢 — MIT-licensed collection of AI-agent Skills, Goals, and Dynamic Workflows for security auditing, authorized penetration testing, and research across web, cloud, blockchain, and defensive workflows. (Forefy) — note: agent-interpreted skill and workflow bundle rather than a deterministic scanner; includes active offensive procedures, so review and pin content before use and run it only in isolated, authorized assessments. The hosted AI Security Registry is a separate SaaS-backed catalog that publishes commit provenance and project-generated scan/audit metadata. (★ 133 · updated 2026-08-31)
- Related: AI Security Registry · Review methodology · OpenAPI schema
- RedAmon 🟢⚠️ — Self-hosted AI pentest framework combining reconnaissance, a Neo4j attack-surface graph, Kali-based tooling, configurable human approvals, and agent-assisted remediation pull requests, with MCP client and server interfaces. — note: authorized testing only, on a dedicated isolated host. The heavy Docker stack includes seccomp-unconfined components and services with Docker-socket access; approval gates are configurable, not an unavoidable boundary. Cloud LLMs receive target context; local inference avoids that provider transfer but does not disable external reconnaissance or other configured services. Bundled tools retain separate licenses, including copyleft and WPScan restrictions. (★ 2,906 · updated 2026-10-02)
AI-Powered Recon & Narrow ML Tools
Hyper-specific AI/ML tools for a single offensive-security, recon, or detection step — the subwiz/eyeballer pattern rather than broad autonomous agents. 🅐 = self-contained trained model or learned model/pattern engine; 🅑 = LLM wrapper that calls an external API.
Subdomain & DNS Prediction
- subwiz 🟢 — 🅐 Lightweight nanoGPT model that predicts resolvable subdomains via beam search; model weights are published on Hugging Face. (Hadrian Security) (★ 388 · updated 2025-12-18)
- Related: HadrianSecurity/subwiz model
- regulator 🟢⚠️ — 🅐 Learns and ranks regex-like naming patterns from known subdomains to generate likely new candidates. — note: no LICENSE file found; treat as source-available until clarified. (★ 392 · updated 2023-02-18)
- Related: subwiz
Recon Screenshot Triage
- eyeballer 🟢⚠️ — 🅐 Convolutional neural network that classifies pentest/recon screenshots (login pages, webapps, old-looking sites, parked domains, and custom 404s) for attack-surface triage. (Bishop Fox) — note: GPL-3.0 licensed. (★ 1,290 · updated 2024-02-19)
Software / Tech Fingerprinting
- GyoiThon 🟢🔬 — 🅐 Machine-learning-assisted web intelligence tool that fingerprints products, versions, CVEs, login pages, debug messages, and related web-server signals from HTTP responses. — note: historical research reference; Apache-2.0 licensed, but maintenance is low. (★ 826 · updated 2021-06-29)
AI-Assisted Fuzzing
- ffufai 🟢⚠️ — 🅑 AI wrapper around the ffuf web fuzzer that suggests file extensions and paths from the target URL and headers using OpenAI or Anthropic models. (Joseph Thacker) — note: requires an LLM API key; README states MIT but no LICENSE file was found. (★ 801 · updated 2025-12-04)
Password / Credential ML
- PassGPT 🔬⚠️ — 🅐 GPT-style password model trained on leaked passwords for research on password generation and strength estimation. (Rando et al.) license: CC BY-NC-4.0 · access: open 10-char model; 16-char variant gated · artifacts: PyTorch/Safetensors. Research-only / non-commercial use; related code: javirandor/passgpt.
- PassGAN 🔬 — 🅐 WGAN that learns password distributions from leaks to generate guesses; historical reference implementation of the PassGAN paper (MIT). — note: historical research reference; not an actively maintained password-auditing product. (★ 2,009 · updated 2018-09-30)
- neural_network_cracking 🔬 — 🅐 RNN password-guessing model from Fast, Lean, and Accurate: Modeling Password Guessability Using Neural Networks (USENIX Security 2016); Apache-2.0 licensed. (CMU CUPS Lab) — note: historical USENIX research implementation, not a maintained password-auditing product. (★ 243 · updated 2018-11-30)
Phishing Detection (Visual / URL)
- phishing-url-detection 🟢 — 🅐 Packaged URL phishing classifier with ONNX and pickle artifacts. license: MIT · access: open · artifacts: ONNX, pickle. Model card recommends ONNX over pickle for safer inference.
- Phishing Email Detection DistilBERT v2.4.1 🟢 — DistilBERT text-classification model for email and URL phishing detection, trained on a public Hugging Face phishing-email dataset. license: Apache-2.0 · access: open · artifacts: Safetensors. — note: strong download signal, but independently verify the very high published metrics before production use.
- PhishIntention 🔬 — 🅐 Deep-vision phishing detector that infers both brand intention and credential-taking intention from webpage appearance and dynamics (USENIX Security 2022). — note: CC0-1.0 licensed. (★ 263 · updated 2026-06-04)
- VisualPhishNet 🔬⚠️ — 🅐 Triplet CNN for zero-day phishing detection by visual similarity to trusted websites (ACM CCS 2020). (CISPA) — note: no LICENSE file found; dataset access is research-request based. (★ 30 · updated 2022-02-09)
AI/ML-Assisted Detection Rules & Engines
- SYARA 🟢🔬 — 🅐 Semantic YARA-like rule engine for text and multimodal signals, adding embedding similarity, classifier-backed rules, LLM evaluators, and pHash matching to familiar YARA-style syntax. — note: early-stage engine; useful for LLM-era intent signals such as phishing, prompt injection, jailbreaks, hallucination, and disinformation rather than classic binary-only YARA matching. (★ 18 · updated 2026-03-05)
- AutoYara 🟢🔬 — 🅐 Research implementation of automatic YARA rule generation via biclustering over byte n-grams for malware-family samples. — note: Apache-2.0 research code from the ACM AISec 2020 paper; README explicitly says it comes with no warranty or support. (★ 79 · updated 2025-10-08)
- yaraml_rules 🟢🔬 — Research code that trains scikit-learn classifiers on malware and benign corpora, then compiles the learned model into deployable YARA rules. (Sophos) — note: historical research reference; the maintained value is the ML-to-YARA technique, not a current detection product. (★ 215 · updated 2020-12-18)
- RuleLLM 🟢🔬 — 🅑 LLM-assisted malware-rule generator that clusters malicious code samples and produces/refines/validates YARA and Semgrep rules. — note: MIT-licensed research prototype; requires OpenAI-compatible API access plus YARA/Semgrep validators. (★ 12 · updated 2025-04-25)
Defensive Trained-Model Detectors
- open-appsec 🟢🟠⚠️ — ML-based web application and API firewall combining an offline-trained model with environment-specific traffic learning and block/log decisions. Caveat: the bundled Apache-2.0 basic model is recommended for monitor/test use; the production-oriented advanced model is a separate portal download with its own terms. Review standalone versus SaaS management, telemetry/data flows, and deployment privileges such as host IPC; detection claims are not a security guarantee. (★ 1,718 · updated 2026-09-01)
- DeepSQLi 🟢⚠️ — 🅐 Deep-learning SQL-injection detector with dataset, trained models, and a Flask Prediction API for GatewayD IDS/IPS integration. (GatewayD) — note: AGPL-3.0 licensed; defensive detector rather than offensive generator. (★ 7 · updated 2026-02-21)
- deepsecrets 🟢 — Semantic secrets scanner using lexing/parsing, entropy checks, and hashed-known-secret matching across 500+ languages. — note: useful narrow detector, but not a trained ML model. (★ 172 · updated 2026-06-04)
- VLAI Vulnerability Severity Classifier 🟢🔬 — RoBERTa-based vulnerability-severity classifier trained on CIRCL vulnerability scores to assist triage before manual CVSS scoring. (CIRCL) license: CC-BY-4.0 · access: open · artifacts: Safetensors.
AI-Powered SAST & Secure Code Review
Static analysis and secure code review enhanced with LLMs.
- Symfony Security Auditor 🟢 — Symfony-aware source-audit package with attacker/reviewer model passes, confidence filtering, baseline deduplication, and budget-bounded partial results. Caveat: reviewer-validated findings are model judgments, not confirmed exploitation. Offline-only mode defaults to false, so hosted providers receive selected source; local inference needs separately provisioned models. Requires PHP 8.3+ and a compatible Symfony version. (★ 95 · updated 2026-10-04)
- Clearwing 🟢 — Agent-driven source-security research pipeline that ranks files, coordinates vulnerability hunting and validation, optionally generates patches, and exports SARIF, JSON, and Markdown findings. Caveat: can compile untrusted targets and execute generated proofs or exploits; use isolated authorized environments and explicit spending limits. Hosted models receive source context; local providers are supported. Optional OTLP/Phoenix tracing can export run metadata. Findings need independent verification, and released packages can lag main. (★ 1,072 · updated 2026-10-02)
- Cloudflare Security Audit Skill 🟢 — Coding-agent skill for a multi-phase source-security audit with reconnaissance, coverage-led hunting, independent verification, structured findings, and a separate record-validation pass. (Cloudflare) — note: workflow skill, not a standalone scanner; it requires a capable coding agent with parallel subagents and an OS-enforced sandbox. Results are nondeterministic, and Cloudflare reports that one run finds only about half of the vulnerabilities found across repeated runs. (★ 18,424 · updated 2026-09-14)
- Mantis 🟢🔬 — Security-review skills and an ADK reference harness for vulnerability discovery, triage, reproduction, patching, and deterministic verification gates. (Google) — note: can generate and execute code while reproducing findings; use only on authorized source in an isolated restricted environment. Model output and proposed findings still require expert verification. (★ 1,678 · updated 2026-09-18)
- Sources: Google Cloud announcement
- VulnHunter (Capital One) 🟢 — Attacker-oriented source-review workflow with forward analysis from exposed entry points, a finding-falsification pass, evidence-backed remediation, and a separate fix-verification flow. (Capital One) — note: built and optimized for Claude Code with an Opus-class model, so source context is processed under the selected provider's terms and results remain model-dependent. Use only on code you are authorized to assess. (★ 1,016 · updated 2026-08-15)
- Sources: Capital One announcement
- Vulnhuntr 🟢 — Zero-shot vulnerability discovery in Python repos via LLM call-chain analysis; credited with a 0-day RCE in Ragflow. (Protect AI) (★ 2,741 · updated 2025-02-06)
- Related: IRIS
- deepsec 🟢 — Agent-powered security harness for scanning large codebases with coding agents, resumable parallel runs, custom matchers, and optional revalidation. (Vercel Labs) (★ 7,695 · updated 2026-08-13)
- Related: claude-code-security-review · sast-skills
- Codex Security 🟠 — CLI and TypeScript SDK that use Codex Security to find, validate, and help fix vulnerabilities in a codebase, with scan comparison and containerized bulk-scan support. (OpenAI) — note: the CLI/SDK are open source, but scans require Codex Security access and, for best results, OpenAI Trusted Access. (★ 9,911 · updated 2026-08-16)
- Related: deepsec · defending-code-reference-harness
- open·kritt 🟢⚠️ — Self-hosted platform that orchestrates Codex or Claude Code across focused vulnerability-research workflows, then validates, de-duplicates, ranks, and reports resulting findings. (Kritt AI) — note: jobs run as root in disposable Docker containers with writable target copies and direct internet access; the stack has no application auth by default and sends scanned code to the configured model provider. Deploy only on a dedicated, access-controlled host and scan authorized targets. (★ 1,877 · updated 2026-08-16)
- Related: deepsec · Visa Vulnerability Agentic Harness
- Visa Vulnerability Agentic Harness 🟢 — Agentic SAST pipeline for autonomous vulnerability discovery, exploitability verification, SARIF/Markdown reporting, remediation, and validation using frontier AI models. (Visa) — note: Apache-2.0; authorized use only. The default
scanprofile can continue into remediation and edit target source files; use--stop-after s9for detection-only runs. (★ 2,556 · updated 2026-08-04)- Related: deepsec · defending-code-reference-harness
- defending-code-reference-harness 🟢 — Reference Claude Code skills and autonomous vulnerability-discovery pipeline for threat modeling, static scanning, triage, execution-verified C/C++ memory-bug discovery, reporting, and patch generation. (Anthropic) — note: official reference implementation, not maintained as a product; the autonomous pipeline executes target code and should be run only inside the documented gVisor sandbox. (★ 7,270 · updated 2026-08-06)
- Related: deepsec · claude-code-security-review
- rust-in-peace 🟢 — Rust-security fork of Anthropic's defending-code reference harness, adding a Rust profile for agentic review of unsafe/FFI memory bugs, panic-DoS, deserialization-trust issues, and Miri/ASan/panic/hang-verified findings. (Sergey Gordeychik) — note: very new Apache-2.0 fork; autonomous runs execute target code and should use the documented sandbox. (★ 14 · updated 2026-08-11)
- Related: defending-code-reference-harness · deepsec
- Reproof 🟢 — Multi-language Kimi Code port of rust-in-peace, orchestrating agentic vulnerability discovery, PoC replay in fresh containers, comparison with known bugs, and patch re-testing. (Sergey Gordeychik) Caveat: newly published, maintainer-authored project; benchmark results are project-reported, not independently reproduced here. The autonomous pipeline executes target code and generated PoCs; use the documented Linux/Docker/gVisor setup on authorized targets. The configured model provider receives source and execution context and may incur costs; review generated patches and sensitive findings before use or disclosure. (★ 1 · updated 2026-10-05)
- Related: rust-in-peace · defending-code-reference-harness
- claude-code-security-review 🟠 — Official Claude-based semantic SAST GitHub Action that reviews PR diffs. (Anthropic) (★ 5,866 · updated 2026-02-11)
- IRIS 🟢🔬 — Neurosymbolic SAST combining LLMs with CodeQL for Java vulnerability detection (MIT). (★ 413 · updated 2026-07-02)
- sast-skills 🟢 — Agent skills that turn AI coding assistants into a multi-agent SAST scanner. (★ 1,276 · updated 2026-04-08)
- llm-sast-scanner 🟢 — SAST skill for AI coding agents with structured source-to-sink analysis across 34 vulnerability classes. License: MIT stated in README. (★ 274 · updated 2026-04-07)
- Related: sast-skills
- sast-ai-workflow 🟢 — LangGraph workflow for reviewing static-analysis findings, reducing false positives, and producing vulnerability review output. (Red Hat Ecosystem AppEng) (★ 20 · updated 2026-06-29)
- Related: seclab-taskflow-agent · Fraim
- llm-security-scanner 🟢⚠️ — LLM-powered code scanner that opens GitHub issues for findings. (★ 22 · updated 2025-04-02)
- Trail of Bits Skills 🟢⚠️ — Claude Code- and Codex-compatible security workflow skills for code review, differential review, false-positive analysis, supply-chain checks, GitHub Actions auditing, Semgrep rule generation, and vulnerability research. (Trail of Bits) — note: reusable agent workflow instructions rather than a standalone deterministic scanner; share-alike terms apply to adapted material. (★ 6,624 · updated 2026-08-14)
- OpenHack 🟢 — File-based source-guided white-box security-review workspace that orchestrates agents through reconnaissance, vulnerability hunting, validation, evidence capture, and reporting. (Hadrian Security) — note: requires an external coding harness/model and can consume substantial model tokens; OpenHack provides workflow state and review artifacts, not sandboxing or an execution-security boundary. Use only on authorized targets. (★ 730 · updated 2026-06-01)
- Related: deepsec · rust-in-peace
- Buttercup 🟢🔬⚠️ — Multi-component cyber reasoning system for finding, validating, and patching software vulnerabilities with coordinated agent workflows. (Trail of Bits) — note: AGPL-3.0 research/competition system rather than a lightweight scanner; deployment uses multiple services, Docker, and configured model providers. (★ 1,676 · updated 2026-08-10)
- Related: OpenHack · Visa Vulnerability Agentic Harness
AI-Powered Threat Modeling
Architecture-level threat model generation and design-phase risk analysis driven by LLM reasoning.
- ThreatForest 🟢🔬 — Strands-based threat-modeling pipeline that analyzes repositories, generates attack trees and mitigations, and maps attack steps to ATT&CK and other TTP frameworks through an embedding service. (AWS Samples) Caveat: sends project context to configured model providers and requires separately licensed embedding weights; optional Langfuse export defaults off. Keep its unauthenticated API on loopback. Generated paths and embedding-based TTP mappings require expert review, not automatic acceptance as demonstrated attack paths. (★ 80 · updated 2026-08-01)
- tachi 🟢 — Threat modeling and AI-reasoning vulnerability detection harness for Claude Code that dispatches 14 specialized threat agents (6 STRIDE, 5 LLM, 3 agentic) against an architecture description in Mermaid, C4, PlantUML, ASCII, or free text, producing SARIF 2.1.0 for code scanning, MAESTRO seven-layer classification, attack trees, CVSS-aligned composite risk scores, compensating-controls analysis of the target codebase, and a PDF report. (David Matousek) — note: runs inside Claude Code; architecture descriptions are processed by the configured Claude model. (★ 89 · updated 2026-08-13)
- Related: STRIDE GPT · defending-code-reference-harness
- STRIDE GPT 🟢 — LLM-powered threat modeling tool that generates STRIDE threat models, attack trees, data flow diagrams, DREAD risk scores, mitigations, and Gherkin test cases from application descriptions, architecture diagrams, or codebases (agentic analysis mode), with OWASP LLM Top 10 and Agentic (ASI) coverage, MITRE ATT&CK/ATLAS mapping, Markdown/JSON/SARIF/HTML output, and broad LLM provider support via LiteLLM including local hosting. (Matthew Adams) (★ 1,101 · updated 2026-08-12)
- Related: tachi
LLM-Driven Fuzzing
Two families: (a) LLMs generating harnesses/targets for traditional fuzzing, and (b) fuzzing the LLM itself.
Harness / target generation
- seclab-taskflows-fuzzing 🟢 — Taskflow package for LLM-assisted C/C++ fuzz-harness generation and repair, AFL++ campaigns, coverage feedback, and crash triage. (GitHub Security Lab) Caveat: requires a configured model provider and a Linux build/fuzzing toolchain; provider-backed analysis can disclose source code. Host-shell and file-write tools are not a sandbox, and setup can install system tools. Use a disposable isolated environment for authorized targets. (★ 20 · updated 2026-09-21)
- Related: seclab-taskflow-agent
- Ultrafuzz 🟢 — Agentic smart-contract fuzzing orchestrator that generates tests, coordinates analysis campaigns, and collects findings and reports. (Monad Foundation) Caveat: agents deliberately bypass permission and sandbox approvals and can access host files and the network. Run only on an ephemeral isolated VM without unrelated credentials; review a published release rather than the default unstable branch. Model/provider configuration determines cost and data disclosure; the harness is not an isolation boundary. (★ 98 · updated 2026-10-02)
- PromeFuzz 🟢🔬 — Research framework for C/C++ fuzz-harness generation using code metadata, documentation, API relationships, iterative compilation repair, and crash analysis (CCS 2025). Caveat: requires target build metadata, a native LLVM/Clang toolchain, and configured generation/embedding models. Hosted providers receive source and documentation context; generated harnesses and target code execute during campaigns. Use isolated authorized environments and do not generalize the paper's coverage or bug-finding results to arbitrary targets. (★ 61 · updated 2026-07-30)
- oss-fuzz-gen 🟢 — LLM-driven fuzz-harness generation for OSS-Fuzz; reported 26 real vulnerabilities (incl. CVE-2024-9143 in OpenSSL). (Google) (★ 1,430 · updated 2026-03-02)
- PromptFuzz 🟢🔬⚠️ — LLM-mutated prompts to generate fuzz drivers for C/C++ libraries (Rust). (★ 342 · updated 2026-05-15)
- Fuzz4All 🟢🔬 — "Universal" LLM-based fuzzer across compilers/languages (ICSE 2024). (★ 336 · updated 2025-08-11)
- ChatAFL 🟢🔬 — LLM-guided protocol fuzzing extending AFLNet (NDSS'24). (★ 392 · updated 2025-06-20)
- TitanFuzz 🟢🔬⚠️ — First LLM-based fuzzer for PyTorch/TensorFlow (ISSTA'23). (★ 94 · updated 2023-09-10)
Fuzzing the LLM
- LLMFuzzer 🟢🔬 — First open-source fuzzing framework for LLM API integrations. — note: historical research reference; maintenance appears low compared with current LLM security scanners. (★ 377 · updated 2024-02-12)
- ps-fuzz 🟠 — System-prompt hardening fuzzer; 16 attacks × 16 providers. (Prompt Security) (★ 703 · updated 2026-02-16)
- FuzzyAI 🟠 — Automated LLM fuzzer for jailbreaks/prompt injection. (CyberArk) (★ 1,563 · updated 2026-02-06)
- spikee 🟢 — Prompt-injection evaluation and exploitation kit with dataset generation, Burp integration, and pluggable judges. (ReversecLabs / WithSecure) (★ 263 · updated 2026-09-11)
- Related: promptmap
- promptmap 🟢⚠️ — Prompt-injection scanner for custom LLM applications in white-box and black-box modes; GPL-3.0 licensed. (★ 1,250 · updated 2025-12-01)
- Related: spikee
- ai-prompt-fuzzer 🟢 — Burp Suite extension fuzzing GenAI/LLM prompts. (PortSwigger) (★ 36 · updated 2025-09-04)
Threat Intelligence
AI/LLM tooling for CTI gathering, IOC/TTP extraction, and analysis.
- ZettelForge 🟢 — CTI investigation memory with entity and IOC extraction, graph and vector retrieval, local storage, and an MCP interface. Caveat: core storage and retrieval are local, but embedding models require an initial download and optional remote LLM providers receive analysis content. Rich extraction and synthesis depend on the configured model; synthesis without an LLM is a placeholder, not a completed analysis. Generated links and attribution need analyst validation. (★ 64 · updated 2026-07-10)
- VoidAccess 🟢 — OSINT and CTI pipeline combining source collection, IOC/entity extraction, enrichment, relationship analysis, and intelligence exports. Caveat: collection contacts external and potentially hostile sources; Tor is not an anonymity guarantee. Cloud model adapters disclose submitted content, while local models are configurable. Neural embeddings are optional and the SHA-256 fallback is not semantic ML; generated attribution and detection rules require analyst review. (★ 770 · updated 2026-08-04)
- cti-skills 🟢 — CTI workflow skills and integration clients for source assessment, IOC investigation, ATT&CK mapping, detection-rule drafting, and MISP/OpenCTI workflows. (Liberty91) Caveat: skills guide a host agent rather than provide an independent detection engine. Live integrations need provider credentials and can disclose intelligence; some clients can create, modify, or delete records. Confirmation instructions are not enforced authorization, so scope credentials and review writes separately. (★ 25 · updated 2026-09-30)
- trs 🟢 — LLM + ChromaDB tool to summarize threat reports and extract MITRE TTPs and IOCs. (★ 10 · updated 2023-11-15)
- TI-Mindmap-GPT 🟢 — Streamlit app: AI summaries, mindmaps, IOC/TTP extraction, and ATT&CK Navigator layers. (★ 111 · updated 2026-02-16)
- aiocrioc 🟢 — LLM + OCR IOC extraction (pulls IOCs from images/PDFs). (★ 38 · updated 2024-12-04)
- ThreatIngestor 🟢 — Extracts/aggregates IOCs from feeds; integrates with MISP/ThreatKB (pairs well with LLM post-processing). (★ 922 · updated 2026-05-26)
- IATelligence 🟢 — Explains imported Window