| agenta |
The LLMOps platform to build robust LLM apps. Easily experiment and evaluate different prompts, models, and workflows to build robust apps. |
 |
| AgentMark |
Type-Safe Markdown-based Agents |
 |
| AgentField |
Open-source control plane for building and operating AI agents like APIs at scale, with routing, memory, observability, identity, auth, and policy controls. |
 |
| Agnos |
Gateway-agnostic, self-hosted LLM control plane (MIT): an OpenAI-compatible governance proxy (auth, budgets, guardrails, cost, observability) that runs LiteLLM, Bifrost or Portkey as a swappable, stateless translation engine, so provider keys stay in your own control plane instead of the gateway. |
 |
| AI studio |
A Reliable Open Source AI studio to build core infrastructure stack for your LLM Applications. It allows you to gain visibility, make your application reliable, and prepare it for production with features such as caching, rate limiting, exponential retry, model fallback, and more. |
 |
| AISIX |
Open-source AI gateway written in Rust: one OpenAI-compatible API (plus a native Anthropic Messages API) in front of LLM providers, and a gateway for MCP servers and A2A agents, with shared API keys, rate limits, guardrails, caching, and Prometheus/OTLP observability. |
 |
| AIWG |
Deploys project-owned agents, skills, rules, and governed workflows across AI coding platforms. |
 |
| Arize-Phoenix |
ML observability for LLMs, vision, language, and tabular models. |
 |
| BitRouter |
Agent-native LLM router that optimizes your agent with every run — zero harness changes, with every model call reliable, traceable, secure, and cost-effective. Routes across OpenAI, Anthropic, Google, OpenRouter, Bedrock, GitHub Copilot, and more through one local endpoint, with cross-protocol translation, an MCP gateway, guardrails, observability, and multi-account failover. Written in Rust. |
 |
| boldrouter |
OpenAI-compatible AI gateway developed and hosted in Switzerland, with automatic routing, failover, and unified prepaid billing. |
|
| BudgetML |
Deploy a ML inference service on a budget in less than 10 lines of code. |
 |
| Caura |
Governed shared memory for AI agent fleets. Multi-agent, multi-tenant and MCP-native, with trust tiers, audit trails, a knowledge graph and self-improving retrieval. |
 |
| Cheshire Cat AI |
Web framework to create vertical AI agents. FastAPI based, plugin system inspired to WordPress, admin panel, vector DB included |
 |
| Chimeraforge |
LLM deployment planner and benchmarking CLI: given a model and a GPU it answers will it fit, will it hit your SLO, and what will it cost. Searches model x quantization x backend x GPU count against VRAM, quality, latency, cost, energy and an opt-in safety gate; models tensor/pipeline parallelism, MoE active parameters, FP8, KV-quant and self-host-vs-API break-even. Every number is labeled measured/estimated/unknown. Also ships an MCP server. |
 |
| claude-router |
Local prompt router that picks the right Claude model tier and prepends the right scaffold using local embeddings, before the API call. |
 |
| CogniCore |
Structured experience memory and learning infrastructure for AI agents, providing verified experience reuse, failure learning, environment-aware validity, and cross-session/cross-agent memory. |
 |
| Contexto |
Self-hosted context engine for AI agents with persistent conversation memory and recall. Works as a drop-in OpenAI-compatible proxy, OpenClaw plugin, or memory SDK — no code changes required. |
 |
| Cortex |
Generates typed MCP servers, interactive API documentation, and SDKs from OpenAPI, AsyncAPI, GraphQL, gRPC, and OpenRPC sources. |
 |
| CrewDock |
Self-hosted control plane for managing AI agents, with task approvals, activity logs, and cost tracking. |
 |
| Dataoorts |
Enjoy unlimited API calls with Serverless AI Workers/LLMs for just $25 per month. No rate or concurrency limits. |
|
| deeplake |
Stream large multimodal datasets to achieve near 100% GPU utilization. Query, visualize, & version control data. Access data w/o the need to recompute the embeddings for the model finetuning. |
 |
| depdesk |
Finds deprecated and retired LLM model identifiers in a codebase and fails CI before a provider retirement reaches production. Parses Python with the stdlib ast module, weights findings by real call volume from a usage export, and ships a hand-transcribed Anthropic and OpenAI catalog carrying the date each provider page was verified. Zero dependencies. |
 |
| Dify |
Open-source framework aims to enable developers (and even non-developers) to quickly build useful applications based on large language models, ensuring they are visual, operable, and improvable. |
 |
| Distil |
Reversible, certified context compression for LLM agents: gates each compression through a statistical non-inferiority test so the agent's decisions don't change. Stdlib-only, zero-dependency proxy for Anthropic/OpenAI/Gemini. |
 |
| Doubleword Control Layer |
The world's fastest open-source AI model gateway — ~450× less overhead than LiteLLM. Turns any model into a production-ready, OpenAI-compatible API with built-in auth, rate limits, and user controls. |
 |
| Dstack |
Cost-effective LLM development in any cloud (AWS, GCP, Azure, Lambda, etc). |
 |
| Ejentum |
Cognitive-harness MCP server with four tools (reasoning, code, anti-deception, memory) returning a structured scaffold (failure pattern, procedure, suppression vectors, falsification test) the agent absorbs before generating. Hosted at api.ejentum.com/mcp and on npm as ejentum-mcp. |
 |
| Embedchain |
Framework to create ChatGPT like bots over your dataset. |
 |
| Epsilla |
An all-in-one platform to create vertical AI agents powered by your private data and knowledge. |
|
| Evidently |
An open-source framework to evaluate, test and monitor ML and LLM-powered systems. |
 |
| FerryAPI |
Low-cost OpenAI-compatible AI API gateway for production workloads, with prepaid balance, usage billing, customer API keys, and provider account pools. |
|
| Fiddler AI |
Evaluate, monitor, analyze, and improve MLOps and LLMOps from pre-production to production. |
|
| FreeLLMAPI |
OpenAI-compatible proxy that stacks the free tiers of 28 LLM providers (~4B free tokens/month) behind one /v1 endpoint — smart routing, automatic failover, encrypted keys. |
 |
| Glide |
Cloud-Native LLM Routing Engine. Improve LLM app resilience and speed. |
 |
| GoModel |
AI gateway exposing a unified OpenAI-compatible API across OpenAI, Anthropic, Gemini, Groq, xAI, Ollama and other providers, with routing, usage tracking, rate limits, and guardrails. |
 |
| gotoHuman |
Bring a human into the loop in your LLM-based and agentic workflows. Prompt users to approve actions, select next steps, or review and validate generated results. |
|
| GPTCache |
Creating semantic cache to store responses from LLM queries. |
 |
| GPUStack |
An open-source GPU cluster manager for running and managing LLMs |
 |
| Haystack |
Quickly compose applications with LLM Agents, semantic search, question-answering and more. |
 |
| Hive |
Open-source AI agent framework for building goal-driven, self-improving autonomous agents with auto-generated graphs, evolution loops, and MCP integration. |
 |
| Helicone |
Open-source LLM observability platform for logging, monitoring, and debugging AI applications. Simple 1-line integration to get started. |
 |
| Humanloop |
The LLM evals platform for enterprises, providing tools to develop, evaluate, and observe AI systems. |
|
| Hypersigil |
Open-source prompt lifecycle management and gateway with a Web UI. |
 |
| Izlo |
Prompt management tools for teams. Store, improve, test, and deploy your prompts in one unified workspace. |
|
| Keywords AI |
A unified DevOps platform for AI software. Keywords AI makes it easy for developers to build LLM applications. |
|
| KubeIntellect |
LLM-orchestrated multi-agent framework for Kubernetes operations. Investigates a live cluster via kubectl, Prometheus (PromQL) and Loki (LogQL), correlates the evidence into a root cause, and gates every mutating action behind human approval with role-based access control. |
 |
| Kunavo |
Hosted AI API gateway with OpenAI-compatible and native Anthropic endpoints, one prepaid balance, and per-request cost accounting. |
|
| MLflow |
An open-source framework for the end-to-end machine learning lifecycle, helping developers track experiments, evaluate models/prompts, deploy models, and add observability with tracing. |
 |
| Laminar |
Open-source all-in-one platform for engineering AI products. Traces, Evals, Datasets, Labels. |
 |
| langchain |
Building applications with LLMs through composability |
 |
| LangFlow |
An effortless way to experiment and prototype LangChain flows with drag-and-drop components and a chat interface. |
 |
| Langfuse |
Open Source LLM Engineering Platform: Traces, evals, prompt management and metrics to debug and improve your LLM application. |
 |
| LangKit |
Out-of-the-box LLM telemetry collection library that extracts features and profiles prompts, responses and metadata about how your LLM is performing over time to find problems at scale. |
 |
| LangWatch |
LLM Ops platform with Analytics, Monitoring, Evaluations and an LLM Optimization Studio powered by DSPy |
 |
| lintlang |
Static analysis for AI agent configs, tool descriptions, and system prompts. Catches vague tool descriptions, missing stop conditions, and schema gaps before they reach runtime. Zero-LLM, deterministic checks, built for CI. |
 |
| LiteLLM 🚅 |
A simple & light 100 line package to standardize LLM API calls across OpenAI, Azure, Cohere, Anthropic, Replicate API Endpoints |
 |
| Literal AI |
Multi-modal LLM observability and evaluation platform. Create prompt templates, deploy prompts versions, debug LLM runs, create datasets, run evaluations, monitor LLM metrics and collect human feedback. |
|
| LlamaIndex |
Provides a central interface to connect your LLMs with external data. |
 |
| LLMApp |
LLM App is a Python library that helps you build real-time LLM-enabled data pipelines with few lines of code. |
 |
| LLMFlows |
LLMFlows is a framework for building simple, explicit, and transparent LLM applications such as chatbots, question-answering systems, and agents. |
 |
| LLMGraph |
No-code visual builder for LLM/AI workflows: turn your docs and models into RAG chatbots and AI agents, then deploy each to a REST API and an embeddable chat widget in one click. |
|
| LRM |
CLI/TUI tool for managing localization files (.resx, JSON, Android, iOS) with LLM-powered translation via Ollama, validation, and code scanning for unused/missing keys. |
 |
| Lunary |
Observability and prompt management for LLM chabots and agents. Debug agents with powerful tracing and logging. Usage analytics and dive deep into the history of your requests. Developer friendly modules with plug-and-play integration into LangChain. |
 |
| Mengram |
Open-source memory infrastructure for AI agents. Provides semantic (entities/facts), episodic (conversations), and procedural (learned behaviors) memory with auto-reflection. Python SDK, JS SDK, MCP server, and REST API. |
 |
| magentic |
Seamlessly integrate LLMs as Python functions. Use type annotations to specify structured output. Mix LLM queries and function calling with regular Python code to create complex LLM-powered functionality. |
 |
| Manag.ai |
Your all-in-one prompt management and observability platform. Craft, track, and perfect your LLM prompts with ease. |
|
| Mirascope |
Intuitive convenience tooling for lightning-fast, efficient development and ensuring quality in LLM-based applications |
 |
| NativePort |
Infrastructure for accessing web-search, scraping, browser-automation, voice, and model-inference providers through a single account, API key, and balance. Publishes dated, capability-specific leaderboards and a public methodology. |
|
| Neurolink |
Multi-provider AI agent framework that unifies 12+ LLM providers (OpenAI, Google, Anthropic, AWS, Azure, Groq, etc.) with workflow orchestration. Production-grade platform for building LLM applications with streaming, tool calling, caching, and enterprise features. Battle-tested at 15M+ requests/month. |
 |
| Nora |
Self-hosted control plane for deploying and operating OpenClaw and Hermes agent fleets on Docker and Kubernetes, with lifecycle management, observability, cost tracking, secrets, schedules, and REST, CLI, and MCP interfaces. |
 |
| OpenLIT |
OpenLIT is an OpenTelemetry-native GenAI and LLM Application Observability tool and provides OpenTelmetry Auto-instrumentation for monitoring LLMs, VectorDBs and Frameworks. It provides valuable insights into token & cost usage, user interaction, and performance related metrics. |
 |
| Opik |
Confidently evaluate, test, and ship LLM applications with a suite of observability tools to calibrate language model outputs across your dev and production lifecycle. |
 |
| OrcaRouter Lite |
Self-hosted, OpenAI-compatible LLM router. Bring your own provider keys and route across 100+ models from multiple providers with automatic failover and streaming; model="auto" selects a model by cost, latency, or quality. Includes a local analytics dashboard and does not send telemetry. |
 |
| Parea AI |
Platform and SDK for AI Engineers providing tools for LLM evaluation, observability, and a version-controlled enhanced prompt playground. |
 |
| Pezzo 🕹️ |
Pezzo is the open-source LLMOps platform built for developers and teams. In just two lines of code, you can seamlessly troubleshoot your AI operations, collaborate and manage your prompts in one place, and instantly deploy changes to any environment. |
 |
| Pilot Protocol |
Open-source overlay network giving AI agents a permanent virtual address, encrypted UDP tunnels with NAT traversal, and an explicit per-peer trust model, plus an app store of installable agent-native capabilities (discover → install → call). |
 |
| PraisonAI |
Production-ready Multi-AI Agents framework with self-reflection. Fastest agent instantiation (3.77μs), 100+ LLM support via LiteLLM, MCP integration, agentic workflows (route/parallel/loop/repeat), built-in memory, Python & JS SDKs. |
 |
| PromptDX |
A declarative, extensible, and composable approach for developing LLM prompts using Markdown and JSX. |
 |
| PromptHub |
Full stack prompt management tool designed to be usable by technical and non-technical team members. Test, version, collaborate, deploy, and monitor, all from one place. |
|
| promptfoo |
Open-source tool for testing & evaluating prompt quality. Create test cases, automatically check output quality and catch regressions, and reduce evaluation cost. |
 |
| PromptFoundry |
The simple prompt engineering and evaluation tool designed for developers building AI applications. |
 |
| PromptLayer 🍰 |
Prompt Engineering platform. Collaborate, test, evaluate, and monitor your LLM applications |
 |
| PromptMage |
Open-source tool to simplify the process of creating and managing LLM workflows and prompts as a self-hosted solution. |
 |
| PromptSite |
A lightweight Python library for prompt lifecycle management that helps you version control, track, experiment and debug with your LLM prompts with ease. Minimal setup, no servers, databases, or API keys required - works directly with your local filesystem, ideal for data scientists and engineers to easily integrate into existing LLM workflows |
|
| Prompteams |
Prompt management system. Version, test, collaborate, and retrieve prompts through real-time APIs. Have GitHub style with repos, branches, and commits (and commit history). |
|
| prompttools |
Open-source tools for testing and experimenting with prompts. The core idea is to enable developers to evaluate prompts using familiar interfaces like code and notebooks. In just a few lines of codes, you can test your prompts and parameters across different models (whether you are using OpenAI, Anthropic, or LLaMA models). You can even evaluate the retrieval accuracy of vector databases. |
 |
| Puzzlet AI |
The Git-Based LLM Engineering Platform. Achieve more from GenAI: Manage, evaluate, and improve your full-stack LLM application - with version control, type-safety, and local development built-in. |
|
| Quotaflow |
AI token and API resource utilization platform that helps teams reduce wasted subscribed quota and improve turnover across controlled internal pools. |
|
| systemprompt.io |
Systemprompt.io is a Rest API with quality tooling to enable the creation, use and observability of prompts in any AI system. Control every detail of your prompt for a SOTA prompt management experience. |
|
| TeamoRouter |
LLM routing gateway for OpenClaw. One API key to access Claude, GPT-4o, Gemini, DeepSeek, Kimi, MiniMax. Smart routing modes (teamo-best, teamo-balanced, teamo-eco) auto-pick the optimal model. Up to 50% off official prices. 2-second install via skill.md. |
|
| TokenMix |
AI gateway routing 171 LLMs from 14 providers (Claude, GPT, Gemini, DeepSeek, Qwen, and more) through one OpenAI-compatible endpoint. Pass-through pricing, automatic failover, no monthly subscription. |
|
| TreeScale |
All In One Dev Platform For LLM Apps. Deploy LLM-enhanced APIs seamlessly using tools for prompt optimization, semantic querying, version management, statistical evaluation, and performance tracking. As a part of the developer friendly API implementation TreeScale offers Elastic LLM product, which makes a unified API Endpoint for all major LLM providers and open source models. |
|
| TrueFoundry |
Deploy LLMOps tools like Vector DBs, Embedding server etc on your own Kubernetes (EKS,AKS,GKE,On-prem) Infra including deploying, Fine-tuning, tracking Prompts and serving Open Source LLM Models with full Data Security and Optimal GPU Management. Train and Launch your LLM Application at Production scale with best Software Engineering practices. |
|
| ReliableGPT 💪 |
Handle OpenAI Errors (overloaded OpenAI servers, rotated keys, or context window errors) for your production LLM Applications. |
 |
| Registry Broker |
Universal index and routing layer for AI agents. Aggregates agent metadata from multiple registries (NANDA, MCP, Virtuals, OpenRouter, A2A, X402 Bazaar) across web2 and web3, normalizes profiles, and provides protocol translation between agent ecosystems. |
 |
| Rhesis |
Open-source testing infrastructure for LLM and agentic applications. Collaborative platform enabling teams to define quality metrics, run evaluations, and ship confidently with version control and peer review workflows built for AI engineering. |
 |
| rote |
Open-source CLI that compiles a proven agent skill (a SKILL.md plus references) into a typed, deterministic pipeline. Fixed logic becomes reviewable Python or TypeScript with per-step tests, while judgment steps stay as typed LLM-judge signatures. Emits DBOS, Temporal, Cloudflare Workflows, Inngest, or plain Python/TS, and can serve compiled pipelines as MCP tools. |
 |
| Roundtable |
Zero-configuration unified AI assistant management built on the FastMCP framework. Provides seamless integration with Claude, ChatGPT, and other AI assistants through a single MCP interface with session management, logging, and production-ready operations. |
 |
| Portkey |
Control Panel with an observability suite & an AI gateway — to ship fast, reliable, and cost-efficient apps. |
|
| SAVI SDK |
Open-source observability SDK for LLM cost, PII masking, carbon, and compliance — drop-in wrapper for OpenAI/Anthropic/Bedrock/Cohere/Mistral/Vertex; works standalone with zero account via local_mode. |
 |
| Self-Hosted AI Stack |
Docker Compose stack for local AI with Ollama, a LiteLLM gateway, RAG, voice services, MCP tools, persistent data, and health checks. |
 |
| Semantic Cache Router |
Distributed semantic cache and stateful routing system that cuts LLM API costs by returning cached responses for semantically similar queries. Uses ANN vector search (cosine ≥ 0.8) and consistent hashing to pin requests to the same worker, achieving ~7× latency reduction on cache hits while scaling horizontally without cache thrash. |
 |
| Spendline |
Financial control layer for AI spend — per-customer cost attribution and hierarchical budgets enforced before the provider call. |
|
| Statewave |
Open-source memory runtime for AI agents. Compiles events into deterministic, provenance-tagged context bundles instead of query-time retrieval. Apache-2.0, self-hostable on Postgres + pgvector. |
 |
| TensorZero |
TensorZero is an open-source framework for building production-grade LLM applications. It unifies an LLM gateway, observability, optimization, evaluations, and experimentation. |
 |
| ThinkWatch Lite |
Desktop app for macOS, Windows and Linux that runs a local LLM gateway for Claude Code, Codex and other AI coding clients: routing rules with failover, conversion between the Anthropic, OpenAI and Gemini APIs, a record of each request's route and cost, and redaction of API keys before requests leave. The gateway engine, ThinkWatch Core, also runs on its own on a Linux server. |
 |
| Trinity |
Self-hosted platform that runs AI coding agents as persistent scheduled services. Each agent runs in a Docker container with its own workspace and git-backed memory, plus scheduling, an approvals queue, an MCP server and a web console. |
 |
| UnoRouter |
OpenAI-compatible LLM gateway with one API key for every major provider and smart routing across models. Drop-in for code, Claude Code, and chat clients like SillyTavern, Janitor.AI, RisuAI, and Chub. |
|
| Vellum |
An AI product development platform to experiment with, evaluate, and deploy advanced LLM apps. |
|
| Weights & Biases (Prompts) |
A suite of LLMOps tools within the developer-first W&B MLOps platform. Utilize W&B Prompts for visualizing and inspecting LLM execution flow, tracking inputs and outputs, viewing intermediate results, securely managing prompts and LLM chain configurations. |
|
| Wenlan |
Local-first AI knowledge base and LLM wiki that distills agent work into source-cited pages with graph context and hybrid retrieval across MCP clients. |
 |
| Wordware |
A web-hosted IDE where non-technical domain experts work with AI Engineers to build task-specific AI agents. It approaches prompting as a new programming language rather than low/no-code blocks. |
|
| XiuRouter |
Hosted multi-model API service with OpenAI Chat Completions and Responses, Anthropic Messages, and Gemini GenerateContent routes, scoped API keys, usage-based pricing, and request-level usage and cost records. |
|
| xTuring |
Build and control your personal LLMs with fast and efficient fine-tuning. |
 |
| ZenML |
Open-source framework for orchestrating, experimenting and deploying production-grade ML solutions, with built-in langchain & llama_index integrations. |
 |
| SwarmClaw |
Self-hosted multi-agent AI runtime with 23+ LLM providers, persistent memory, skills, schedules, sub-agent spawning, and MCP client + server support. Ships as desktop app, CLI, or Docker. |
 |
| ai-evaluation |
Evaluation framework for automated, reproducible scoring of LLM, agent, and workflow performance. |
 |
| future-agi |
Open-source self-hostable end-to-end agent engineering and optimization platform unifying tracing, evals, simulations, datasets, gateway, and guardrails for LLM and AI agent applications. |
 |
| Modelglass |
Sourced, versioned pricing and capability data for AI models (image, language, video, audio, plus coding/science/agentic benchmark verticals) to find the cheapest model that clears a capability bar. |
|