
全世界最好的大语言模型资源汇总 持续更新
挖掘那些真正有价值的项目,而不仅仅是噱头
[!TIP] 如果您对医疗数据集/大模型/多模态/评估相关资源感兴趣!请访问我们的 🤗 Awesome-AI4Med !
📢友情赞助
PackyCode 是一家稳定、高效的 API 中转服务商,一句话接入主流大模型。统一域名、统一密钥、智能容灾切换,97% 可用性。人民币1:1充值,无汇率无手续费坑,新用户首充立享折扣 + $1免费体验额度,多分组折扣低至 2 折起,提供专属Codex/Claude Code高速通道。
感谢 Infistar.cc 无限星河 赞助本项目!
⚡ 稳定高效的模型通道:价格低至官方渠道0.3折,模型倍率公开透明,多节点冗余保障,有效降低限流、429 与断连影响; 🧠 主流模型一站式接入:一个 API Key 即可调用 Claude、ChatGPT、Gemini、Kimi、GLM、DeepSeek 等主流模型,覆盖文本生成、图像生成、多模态推理等场景; 📚 契合本仓库资源实践:无论是多模态生成、辅助编程、模型推理还是 MCP / Agent 开发,都可直接把 Infistar 作为 OpenAI、Claude、Gemini 兼容上游,边学边练,零门槛跑通仓库中收录的各类项目与教程;
🎁 专属福利:通过 [专属推广链接] 注册即可领取 5 美元等值测试额度 / 首充专属优惠,快速接入并开始调用!
感谢 APIMart 赞助了本项目!APIMart 是专注 AI 图片/视频生成的低价 API 平台,GPT-Image-2 低至 $0.006/张,1 美元可出图 160+ 张。图片、视频一套异步 API 通吃,提交任务拿 ID、回调取结果,跑批万张不超时、换模型不改代码。按量付费、无月费,通过此 注册链接 注册即可开用。
目录
- 推荐 Suggestion 🌟
- 数据 Data
- 微调 Fine-Tuning
- Agentic RL 🌟
- 推理 Inference
- 评估 Evaluation
- 体验 Usage
- 知识库 RAG
- 智能体 Agents
- 研究 Research 🔥
- 代码 Coding
- 视频 Video
- 图片 Image 🔥
- 搜索 Search
- 语音 Speech 🔥
- 世界模型 World Models 🔥
- 龙虾 OpenClaw
- 统一模型 Unified Model 🌟
- 书籍 Book
- 课程 Course
- 教程 Tutorial
- 论文 Paper
- 社区 Community
- 模型上下文协议 MCP
- 技能 Skills
- 推理 Open o1
- 推理 Open o3
- 小语言模型 Small Language Model 🌟
- 小多模态模型 Small Vision Language Model 🌟
- 技巧 Tips
推荐 Suggestion
Podcast
- 谷歌AI的14年、Gemini翻身之战,与视觉理解模型:专访DeepMind前核心科学家Andrew Dai|Neolabs特辑
- 140. 对姚顺宇的4小时访谈:请允许我小疯一下!在Anthropic和Gemini训模型、技术预测、英雄主义已过去
- 张驰: A Year Inside ByteDance's AI Lab
- Luo Fuli: OpenClaw, Agent Frameworks — The AI Paradigm Has Already Changed Dramatically!
- A 7-hour marathon interview with Saining Xie: World Models, AMI Labs, Yann LeCun, Fei-Fei Li, and 42
- 翁家翌:OpenAI,GPT,强化学习,Infra,后训练,天授,tuixue,开源,CMU,清华|WhynotTV Podcast
- Lovart 创始人陈冕×罗永浩!且让我大闹一场,然后悄然离去
- MiniMax 创始人闫俊杰×罗永浩!大山并非无法翻越
- 影视飓风TIM×罗永浩!用影像打开世界的梦想家
- 129. 全球大模型第一股的上市访谈,和智谱CEO张鹏聊:敢问路在何方?
- 128. Manus决定出售前最后的访谈:啊,这奇幻的2025年漂流啊…
- 122. 朱啸虎现实主义故事的第三次连载:人工智能的盛筵与泡泡
- 119. Kimi Linear、Minimax M2?和杨松琳考古算法变种史,并预演未来架构改进方案
- 118. 对李想的第二次3小时访谈:CEO大模型、MoE、梁文锋、VLA、能量、记忆、对抗人性、亲密关系、人类的智慧
- 115. 对OpenAI姚顺雨3小时访谈:6年Agent研究、人与系统、吞噬的边界、既单极又多元的世界
- 113. 和杨植麟时隔1年的对话:K2、Agentic LLM、缸中之脑和“站在无限的开端”
数据 Data
[!NOTE]
此处命名为
数据,但这里并没有提供具体数据集,而是提供了处理获取大规模数据的方法
-
AotoLabel: Label, clean and enrich text datasets with LLMs.
-
LabelLLM: The Open-Source Data Annotation Platform.
-
data-juicer: A one-stop data processing system to make data higher-quality, juicier, and more digestible for LLMs!
-
OmniParser: a native Golang ETL streaming parser and transform library for CSV, JSON, XML, EDI, text, etc.
-
MinerU (
🔥): MinerU is a one-stop, open-source, high-quality data extraction tool, supports PDF/webpage/e-book extraction. -
PDF-Extract-Kit: A Comprehensive Toolkit for High-Quality PDF Content Extraction.
-
Parsera: Lightweight library for scraping web-sites with LLMs.
-
Sparrow: Sparrow is an innovative open-source solution for efficient data extraction and processing from various documents and images.
-
Docling: Get your documents ready for gen AI.
-
GOT-OCR2.0: OCR Model.
-
LLM Decontaminator: Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.
-
DataTrove: DataTrove is a library to process, filter and deduplicate text data at a very large scale.
-
llm-swarm: Generate large synthetic datasets like Cosmopedia.
-
Distilabel: Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
-
Common-Crawl-Pipeline-Creator: The Common Crawl Pipeline Creator.
-
Tabled: Detect and extract tables to markdown and csv.
-
Zerox: Zero shot pdf OCR with gpt-4o-mini.
-
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception.
-
TensorZero: make LLMs improve through experience.
-
Promptwright: Generate large synthetic data using a local LLM.
-
pdf-extract-api: Document (PDF) extraction and parse API using state of the art modern OCRs + Ollama supported models.
-
pdf2htmlEX: Convert PDF to HTML without losing text or format.
-
Extractous: Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
-
MegaParse: File Parser optimised for LLM Ingestion with no loss.
-
MarkItDown: Python tool for converting files and office documents to Markdown.
-
datasketch: datasketch gives you probabilistic data structures that can process and search very large amount of data super fast, with little loss of accuracy.
-
semhash: lightweight and flexible tool for deduplicating datasets using semantic similarity.
-
ReaderLM-v2: a 1.5B parameter language model that converts raw HTML into beautifully formatted markdown or JSON.
-
Bespoke Curator: Data Curation for Post-Training & Structured Data Extraction.
-
LangKit: An open-source toolkit for monitoring Large Language Models (LLMs). Extracts signals from prompts & responses, ensuring safety & security.
-
Curator: Synthetic Data curation for post-training and structured data extraction.
-
olmOCR: A toolkit for training language models to work with PDF documents in the wild.
-
Easy Dataset (
🔥): A powerful tool for creating fine-tuning datasets for LLM. -
BabelDOC: PDF scientific paper translation and bilingual comparison library.
-
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting.
-
EasyDistill: Easy Knowledge Distillation for Large Language Models.
-
ContextGem: a free, open-source LLM framework that makes it radically easier to extract structured data and insights from documents.
-
OCRFlux: a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.
-
DataFlow: Easy Data Preparation with latest LLMs-based Operators and Pipelines.
-
DatasetLoom (
multimodal): 一个面向多模态大模型训练的智能数据集构建与评估平台. -
Chandra: a highly accurate OCR model that converts images and PDFs into structured HTML/Markdown/JSON while preserving layout information.
-
HunyuanOCR: a leading end-to-end OCR expert VLM powered by Hunyuan's native multimodal architecture.
-
DeepSeek-OCR-2: Visual Causal Flow.
-
PaddleOCR-VL-1.5 (
🔥): Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing. -
GLM-OCR: a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture.
微调 Fine-Tuning
-
LLaMA-Factory (
🔥): Unify Efficient Fine-Tuning of 100+ LLMs. -
360-LLaMA-Factory: Unify Efficient Fine-Tuning of 100+ LLMs. (add Sequence Parallelism for supporting long context training)
-
unsloth (
🔥): 2-5X faster 80% less memory LLM finetuning. -
TRL: Transformer Reinforcement Learning.
-
Firefly: Firefly: 大模型训练工具,支持训练数十种大模型
-
Xtuner: An efficient, flexible and full-featured toolkit for fine-tuning large models.
-
torchtune: A Native-PyTorch Library for LLM Fine-tuning.
-
Swift: Use PEFT or Full-parameter to finetune 200+ LLMs or 15+ MLLMs.
-
AutoTrain: A new way to automatically train, evaluate and deploy state-of-the-art Machine Learning models.
-
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework (Support 70B+ full tuning & LoRA & Mixtral & KTO).
-
Ludwig: Low-code framework for building custom LLMs, neural networks, and other AI models.
-
mistral-finetune: A light-weight codebase that enables memory-efficient and performant finetuning of Mistral's models.
-
aikit: Fine-tune, build, and deploy open-source LLMs easily!
-
H2O-LLMStudio: H2O LLM Studio - a framework and no-code GUI for fine-tuning LLMs.
-
LitGPT: Pretrain, finetune, deploy 20+ LLMs on your own data. Uses state-of-the-art techniques: flash attention, FSDP, 4-bit, LoRA, and more.
-
LLMBox: A comprehensive library for implementing LLMs, including a unified training pipeline and comprehensive model evaluation.
-
PaddleNLP: Easy-to-use and powerful NLP and LLM library.
-
workbench-llamafactory: This is an NVIDIA AI Workbench example project that demonstrates an end-to-end model development workflow using Llamafactory.
-
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & Mixtral).
-
TinyLLaVA Factory: A Framework of Small-scale Large Multimodal Models.
-
LLM-Foundry: LLM training code for Databricks foundation models.
-
lmms-finetune: A unified codebase for finetuning (full, lora) large multimodal models, supporting llava-1.5, qwen-vl, llava-interleave, llava-next-video, phi3-v etc.
-
Simplifine: Simplifine lets you invoke LLM finetuning with just one line of code using any Hugging Face dataset or model.
-
Transformer Lab: Open Source Application for Advanced LLM Engineering: interact, train, fine-tune, and evaluate large language models on your own computer.
-
Liger-Kernel: Efficient Triton Kernels for LLM Training.
-
ChatLearn: A flexible and efficient training framework for large-scale alignment.
-
nanotron: Minimalistic large language model 3D-parallelism training.
-
Proxy Tuning: Tuning Language Models by Proxy.
-
Effective LLM Alignment: Effective LLM Alignment Toolkit.
-
Meta Lingua: a lean, efficient, and easy-to-hack codebase to research LLMs.
-
Vision-LLM Alignemnt: This repository contains the code for SFT, RLHF, and DPO, designed for vision-based LLMs, including the LLaVA models and the LLaMA-3.2-vision models.
-
finetune-Qwen2-VL: Quick Start for Fine-tuning or continue pre-train Qwen2-VL Model.
-
Online-RLHF: A recipe for online RLHF and online iterative DPO.
-
InternEvo: an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencies.
-
veRL (
🔥): Volcano Engine Reinforcement Learning for LLM. -
Axolotl: Axolotl is designed to work with YAML config files that contain everything you need to preprocess a dataset, train or fine-tune a model, run model inference or evaluation, and much more.
-
Oumi: Everything you need to build state-of-the-art foundation models, end-to-end.
-
Kiln: The easiest tool for fine-tuning LLM models, synthetic data generation, and collaborating on datasets.
-
DeepSeek-671B-SFT-Guide: An open-source solution for full parameter fine-tuning of DeepSeek-V3/R1 671B, including complete code and scripts from training to inference, as well as some practical experiences and conclusions.
-
MLX-VLM: MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
-
RL-Factory: Train your Agent model via our easy and efficient framework.
-
RM-Gallery: A One-Stop Reward Model Platform.
-
ART: rain multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training.
-
LMMs-Engine: A simple, any-to-any modality framework for pretraining and finetuning. Lean, flexible, and built for research.
-
dLLM: a library that unifies the training and evaluation of diffusion language models, bringing transparency and reproducibility to the entire development pipeline.
diffusion -
Miles: an enterprise-facing reinforcement learning framework for large-scale MoE post-training and production workloads.
-
Skills: a collection of pipelines to improve "skills" of large language models (LLMs).
-
Twinkle: a lightweight, client-server training framework engineered with modular, high-cohesion interfaces.
-
NeMo AutoModel: Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support.
-
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo.
-
Soup: One-config CLI for LLM post-training (SFT/DPO/GRPO/KTO/ORPO). Layer streaming trains an 8B model on a 4 GB laptop GPU by streaming the frozen base from host RAM one decoder layer at a time.
Agentic RL
-
veRL (
🔥): https://github.com/volcengine/verl -
slime (
🔥): https://github.com/THUDM/slime -
Agent Lightning: https://github.com/microsoft/agent-lightning
推理 Inference
-
ollama (
🔥): Get up and running with Llama 3, Mistral, Gemma, and other large language models. -
Open WebUI: User-friendly WebUI for LLMs (Formerly Ollama WebUI).
-
Text Generation WebUI: A Gradio web UI for Large Language Models. Supports transformers, GPTQ, AWQ, EXL2, llama.cpp (GGUF), Llama models.
-
Xinference: A powerful and versatile library designed to serve language, speech recognition, and multimodal models.
-
LangChain: Build context-aware reasoning applications.
-
LlamaIndex: A data framework for your LLM applications.
-
lobe-chat: an open-source, modern-design LLMs/AI chat framework. Supports Multi AI Providers, Multi-Modals (Vision/TTS) and plugin system.
-
TensorRT-LLM: TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs.
-
vllm (
🔥): A high-throughput and memory-efficient inference and serving engine for LLMs. -
LlamaChat: Chat with your favourite LLaMA models in a native macOS app.
-
NVIDIA ChatRTX: ChatRTX is a demo app that lets you personalize a GPT large language model (LLM) connected to your own content—docs, notes, or other data.
-
LM Studio (
🔥): Discover, download, and run local LLMs. -
chat-with-mlx: Chat with your data natively on Apple Silicon using MLX Framework.
-
LLM Pricing: Quickly Find the Perfect Large Language Models (LLM) API for Your Budget! Use Our Free Tool for Instant Access to the Latest Prices from Top Providers.
-
Open Interpreter: A natural language interface for computers.
-
Chat-ollama: An open source chatbot based on LLMs. It supports a wide range of language models, and knowledge base management.
-
chat-ui: Open source codebase powering the HuggingChat app.
-
MemGPT: Create LLM agents with long-term memory and custom tools.
-
koboldcpp: A simple one-file way to run various GGML and GGUF models with KoboldAI's UI.
-
LLMFarm: llama and other large language models on iOS and MacOS offline using GGML library.
-
enchanted: Enchanted is iOS and macOS app for chatting with private self hosted language models such as Llama2, Mistral or Vicuna using Ollama.
-
Flowise: Drag & drop UI to build your customized LLM flow.
-
Jan: Jan is an open source alternative to ChatGPT that runs 100% offline on your computer. Multiple engine support (llama.cpp, TensorRT-LLM).
-
LMDeploy: LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
-
RouteLLM: A framework for serving and evaluating LLM routers - save LLM costs without compromising quality!
-
MInference: About To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
-
Mem0: The memory layer for Personalized AI.
-
SGLang (
🔥): SGLang is yet another fast serving framework for large language models and vision language models. -
AirLLM: AirLLM optimizes inference memory usage, allowing 70B large language models to run inference on a single 4GB GPU card without quantization, distillation and pruning. And you can run 405B Llama3.1 on 8GB vram now.
-
LLMHub: LLMHub is a lightweight management platform designed to streamline the operation and interaction with various language models (LLMs).
-
LiteLLM (
🔥): Call all LLM APIs using the OpenAI format [Bedrock, Huggingface, VertexAI, TogetherAI, Azure, OpenAI, Groq etc.] -
GuideLLM: GuideLLM is a powerful tool for evaluating and optimizing the deployment of large language models (LLMs).
-
LLM-Engines: A unified inference engine for large language models (LLMs) including open-source models (VLLM, SGLang, Together) and commercial models (OpenAI, Mistral, Claude).
-
OARC: ollama_agent_roll_cage (OARC) is a local python agent fusing ollama llm's with Coqui-TTS speech models, Keras classifiers, Llava vision, Whisper recognition, and more to create a unified chatbot agent for local, custom automation.
-
g1: Using Llama-3.1 70b on Groq to create o1-like reasoning chains.
-
MemoryScope: MemoryScope provides LLM chatbots with powerful and flexible long-term memory capabilities, offering a framework for building such abilities.
-
OpenLLM: Run any open-source LLMs, such as Llama 3.1, Gemma, as OpenAI compatible API endpoint in the cloud.
-
Infinity: The AI-native database built for LLM applications, providing incredibly fast hybrid search of dense embedding, sparse embedding, tensor and full-text.
-
optillm: an OpenAI API compatible optimizing inference proxy which implements several state-of-the-art techniques that can improve the accuracy and performance of LLMs.
-
LLaMA Box: LLM inference server implementation based on llama.cpp.
-
ZhiLight: A highly optimized inference acceleration engine for Llama and its variants.
-
DashInfer: DashInfer is a native LLM inference engine aiming to deliver industry-leading performance atop various hardware architectures.
-
LocalAI: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required.
-
ktransformers: A Flexible Framework for Experiencing Cutting-edge LLM Inference Optimizations.
-
SkyPilot: Run AI and batch jobs on any infra (Kubernetes or 14+ clouds). Get unified execution, cost savings, and high GPU availability via a simple interface.
-
Chitu: High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
-
TokenSwift: From Hours to Minutes: Lossless Acceleration of Ultra Long Sequence Generation.
-
Cherry Studio (
🔥): a desktop client that supports for multiple LLM providers, available on Windows, Mac and Linux. -
Shimmy: Python-free Rust inference server — OpenAI-API compatible. GGUF + SafeTensors, hot model swap, auto-discovery, single binary.
-
LlamaBarn: Run local LLMs on your Mac with a simple menu bar app.
-
Parallax: a distributed model serving framework that lets you build your own AI cluster anywhere.
-
xLLM: A high-performance inference engine for LLMs, optimized for diverse AI accelerators.
-
Rapid-MLX: OpenAI-compatible local LLM inference server for Apple Silicon, 2-4x faster than Ollama.
-
TokenSpeed: a speed-of-light LLM inference engine designed for agentic workloads, with TensorRT-LLM-level performance and vLLM-level usability. Our goal is to be the most performant inference engine for production agentic workloads.
-
FreeToken: Unlock datacenter-class intelligence on the hardware you already own.
-
NInfer: High-performance single-GPU inference for selected model checkpoints and GPUs.
评估 Evaluation
-
lm-evaluation-harness: A framework for few-shot evaluation of language models.
-
opencompass (
🔥): OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets. -
llm-comparator: LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed.
-
Weave: A lightweight toolkit for tracking and evaluating LLM applications.
-
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures.
-
Evaluation guidebook: If you've ever wondered how to make sure an LLM performs well on your specific task, this guide is for you!
-
Ollama Benchmark: LLM Benchmark for Throughput via Ollama (Local LLMs).
-
VLMEvalKit: Open-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 40+ benchmarks.
-
EvalScope: A streamlined and customizable framework for efficient large model evaluation and performance benchmarking.
-
DeepEval: a simple-to-use, open-source LLM evaluation framework, for evaluating and testing large-language model systems.
-
Lighteval: Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends.
-
QwQ/eval: QwQ is the reasoning model series developed by Qwen team, Alibaba Cloud.
-
Evalchemy: A unified and easy-to-use toolkit for evaluating post-trained language models.
-
MathArena: Evaluation of LLMs on latest math competitions.
-
YourBench: A Dynamic Benchmark Generation Framework.
-
MedEvalKit: A Unified Medical Evaluation Framework.
-
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards.
体验 Usage
-
Vals AI
evaluation
知识库 RAG
-
AnythingLLM: The all-in-one AI app for any LLM with full RAG and AI Agent capabilites.
-
MaxKB: 基于 LLM 大语言模型的知识库问答系统。开箱即用,支持快速嵌入到第三方业务系统
-
RAGFlow: An open-source RAG (Retrieval-Augmented Generation) engine based on deep document understanding.
-
Dify: An open-source LLM app development platform. Dify's intuitive interface combines AI workflow, RAG pipeline, agent capabilities, model management, observability features and more, letting you quickly go from prototype to production.
-
FastGPT: A knowledge-based platform built on the LLM, offers out-of-the-box data processing and model invocation capabilities, allows for workflow orchestration through Flow visualization.
-
Langchain-Chatchat: 基于 Langchain 与 ChatGLM 等不同大语言模型的本地知识库问答
-
QAnything: Question and Answer based on Anything.
-
Quivr: A personal productivity assistant (RAG) ⚡️🤖 Chat with your docs (PDF, CSV, ...) & apps using Langchain, GPT 3.5 / 4 turbo, Private, Anthropic, VertexAI, Ollama, LLMs, Groq that you can share with users ! Local & Private alternative to OpenAI GPTs & ChatGPT powered by retrieval-augmented generation.
-
RAG-GPT: RAG-GPT, leveraging LLM and RAG technology, learns from user-customized knowledge bases to provide contextually relevant answers for a wide range of queries, ensuring rapid and accurate information retrieval.
-
Verba: Retrieval Augmented Generation (RAG) chatbot powered by Weaviate.
-
FlashRAG: A Python Toolkit for Efficient RAG Research.
-
GraphRAG: A modular graph-based Retrieval-Augmented Generation (RAG) system.
-
LightRAG: LightRAG helps developers with both building and optimizing Retriever-Agent-Generator pipelines.
-
GraphRAG-Ollama-UI: GraphRAG using Ollama with Gradio UI and Extra Features.
-
nano-GraphRAG: A simple, easy-to-hack GraphRAG implementation.
-
RAG Techniques: This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. RAG systems combine information retrieval with generative models to provide accurate and contextually rich responses.
-
ragas: Evaluation framework for your Retrieval Augmented Generation (RAG) pipelines.
-
kotaemon: An open-source clean & customizable RAG UI for chatting with your documents. Built with both end users and developers in mind.
-
RAGapp: The easiest way to use Agentic RAG in any enterprise.
-
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text.
-
LightRAG: Simple and Fast Retrieval-Augmented Generation.
-
TEN: the Next-Gen AI-Agent Framework, the world's first truly real-time multimodal AI agent framework.
-
AutoRAG: RAG AutoML tool for automatically finding an optimal RAG pipeline for your data.
-
KAG: KAG is a knowledge-enhanced generation framework based on OpenSPG engine, which is used to build knowledge-enhanced rigorous decision-making and information retrieval knowledge services.
-
Fast-GraphRAG: RAG that intelligently adapts to your use case, data, and queries.
-
DB-GPT GraphRAG: DB-GPT GraphRAG integrates both triplet-based knowledge graphs and document structure graphs while leveraging community and document retrieval mechanisms to enhance RAG capabilities, achieving comparable performance while consuming only 50% of the tokens required by Microsoft's GraphRAG. Refer to the DB-GPT Graph RAG User Manual for details.
-
Chonkie: The no-nonsense RAG chunking library that's lightweight, lightning-fast, and ready to CHONK your texts.
-
RAGLite: RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with PostgreSQL or SQLite.
-
KAG: KAG is a logical form-guided reasoning and retrieval framework based on OpenSPG engine and LLMs.
-
CAG: CAG leverages the extended context windows of modern large language models (LLMs) by preloading all relevant resources into the model’s context and caching its runtime parameters.
-
MiniRAG: an extremely simple retrieval-augmented generation framework that enables small models to achieve good RAG performance through heterogeneous graph indexing and lightweight topology-enhanced retrieval.
-
XRAG: a benchmarking framework designed to evaluate the foundational components of advanced Retrieval-Augmented Generation (RAG) systems.
-
Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation.
-
RAG-Anything: All-in-One RAG System.
智能体 Agents
- AutoGen: AutoGen is a framework that enables the development of LLM applications using multiple agents that can converse with each other to solve tasks. AutoGen AIStudio
- CrewAI: Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewAI empowers agents to work together seamlessly, tackling complex tasks.
- Coze
- AgentGPT: Assemble, configure, and deploy autonomous AI Agents in your browser.
- XAgent: An Autonomous LLM Agent for Complex Task Solving.
- MobileAgent: The Powerful Mobile Device Operation Assistant Family.
- Lagent: A lightweight framework for building LLM-based agents.
- Qwen-Agent: Agent framework and applications built upon Qwen2, featuring Function Calling, Code Interpreter, RAG, and Chrome extension.
- LinkAI: 一站式 AI 智能体搭建平台
- Baidu APPBuilder
- agentUniverse: agentUniverse is a LLM multi-agent framework that allows developers to easily build multi-agent applications. Furthermore, through the community, they can exchange and share practices of patterns across different domains.
- LazyLLM: 低代码构建多Agent大模型应用的开发工具
- AgentScope: Start building LLM-empowered multi-agent applications in an easier way.
- AgentField: Open-source control plane for building and operating AI agents like APIs at scale, with routing, memory, observability, identity, auth, and policy controls.
- MoA: Mixture of Agents (MoA) is a novel approach that leverages the collective strengths of multiple LLMs to enhance performance, achieving state-of-the-art results.
- Agently: AI Agent Application Development Framework.
- OmAgent: A multimodal agent framework for solving complex tasks.
- Tribe: No code tool to rapidly build and coordinate multi-agent teams.
- CAMEL: First LLM multi-agent framework and an open-source community dedicated to finding the scaling law of agents.
- PraisonAI: PraisonAI application combines AutoGen and CrewAI or similar frameworks into a low-code solution for building and managing multi-agent LLM systems, focusing on simplicity, customisation, and efficient human-agent collaboration.
- IoA: An open-source framework for collaborative AI agents, enabling diverse, distributed agents to team up and tackle complex tasks through internet-like connectivity.
- llama-agentic-system : Agentic components of the Llama Stack APIs.
- Agent Zero: Agent Zero is not a predefined agentic framework. It is designed to be dynamic, organically growing, and learning as you use it.
- Agents: An Open-source Framework for Data-centric, Self-evolving Autonomous Language Agents.
- AgentScope: Start building LLM-empowered multi-agent applications in an easier way.
- FastAgency: The fastest way to bring multi-agent workflows to production.
- Swarm: Framework for building, orchestrating and deploying multi-agent systems. Managed by OpenAI Solutions team. Experimental framework.
- Agent-S: an open agentic framework that uses computers like a human.
- PydanticAI: Agent Framework / shim to use Pydantic with LLMs.
- Agentarium: open-source framework for creating and managing simulations populated with AI-powered agents.
- smolagents: a barebones library for agents. Agents write python code to call tools and orchestrate other agents.
- Cooragent: Cooragent is an AI agent collaboration community.
- Agno: Agno is a lightweight library for building Agents with memory, knowledge, tools and reasoning.
- Suna: Open Source Generalist AI Agent.
- rowboat: Let AI build multi-agent workflows for you in minutes.
- EvoAgentX: Building a Self-Evolving Ecosystem of AI Agents.
- ii-agent: a new open-source framework to build and deploy intelligent agents.
- OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation.
- OpenManus: No fortress, purely open ground. OpenManus is Coming.
- JoyAgent-JDGenie: 业界首个开源高完成度轻量化通用多智能体产品.
- coze-studio: An AI agent development platform with all-in-one visual tools, simplifying agent creation, debugging, and deployment like never before.
- OxyGent: An advanced Python framework that empowers developers to quickly build production-ready intelligent systems.
- LazyCraft: LazyCraft 是一个基于 LazyLLM 构建的 AI Agent 应用开发与管理平台,旨在协助开发者以 低门槛、低成本 快速构建和发布大模型应用。
- OpenAgents: AI Agent Networks for Open Collaboration.
- SandBox: All-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.
- DeepAnalyze: First agentic LLM for autonomous data science, supporting specific data tasks (data preparation, analysis, modeling, visualization, and insight) and data-oriented deep research (produce analyst-grade research reports).
- Astron Agent: Enterprise-grade, commercial-friendly agentic workflow platform for building next-generation SuperAgents.
- Youtu-Agent: A simple yet powerful agent framework that delivers with open-source models.
- MiroThinker: an open-source search agent model, built for tool-augmented reasoning and real-world information seeking, aiming to match the deep research experience of OpenAI Deep Research and Gemini Deep Research.
- Nexent: A zero-code platform for auto-generating agents — no orchestration, no complex drag-and-drop required, using pure language to develop any agent you want.
- Yunjue-Agent: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks.
- Hindsight: State-of-the-art long-term memory for AI agents by Vectorize. Open-source, self-hostable, with integrations for LangChain, CrewAI, LlamaIndex, MCP, and more.
- AgentsMesh: The AI Agent Workforce Platform. Self-hostable multi-agent orchestration with remote AI workstations (AgentPods), PTY sandbox + git worktree isolation, channels-based agent collaboration, built-in Kanban, and per-pod MCP server. Supports Claude Code, Codex CLI, Gemini CLI, Aider, OpenCode.
- BitFun: Open-source agentic development environment with a Rust/Tauri desktop app and CLI for coding, research, office work, browser and desktop automation, extensible through MCP, Skills, and custom agents.
Harness
-
pi: AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI.
-
DeepSeek Harness: Everything is a Plugin.
-
OpenSquilla: a token-efficient, microkernel AI agent.
-
PenguinHarness: Your Automated Agent Builder, Right on Your Desktop / Server.
-
FrontierAgent: an open-source agent runtime, terminal product, and evaluation suite for long-horizon research and file-based work.
研究 Research
写作
- PaperDebugger: https://github.com/PaperDebugger/PaperDebugger
- XtraGPT served as refiner: https://github.com/Xtra-Computing/XtraGPT
- Chat Overleaf: https://github.com/anuin-cat/chat-overleaf
- 文智云助手: https://overleaf.top/
- LiteWrite: https://litewrite.ai/
- Prism: https://openai.com/zh-Hans-CN/prism/
- claude-prism: https://github.com/delibae/claude-prism
审稿
- PaperReview: https://paperreview.ai/
- aiXiv: https://aixiv.science/
- OpenJudge Review: https://openjudge.me/paper_review
PPT
- PPTAgent: https://github.com/icip-cas/PPTAgent
- Paper PPT Agent: https://github.com/CRui5in/paper-ppt-agent
- PPT Master: https://github.com/hugohe3/ppt-master
- LandPPT: https://github.com/sligter/LandPPT
- Kami: https://github.com/tw93/kami
- beautiful-html-templates: https://github.com/zarazhangrui/beautiful-html-templates
- guizang-ppt-skill: https://github.com/op7418/guizang-ppt-skill
- GordenSuperPPTSkills: https://github.com/GordenSun/GordenSuperPPTSkills
- dashiAI-ppt-skill: https://github.com/chuspeeism/dashiAI-ppt-skill
其他
- Paper2Video: https://github.com/showlab/Paper2Video
- Paper2Poster: https://github.com/Paper2Poster/Paper2Poster
- AutoPR: https://github.com/irgolic/AutoPR
- Auto-Slides: https://github.com/Westlake-AGI-Lab/Auto-Slides
- EvoPresent: https://github.com/eric-ai-lab/EvoPresent
- Paper2All: https://github.com/YuhangChen1/Paper2All
- AutoPage: https://github.com/AutoLab-SAI-SJTU/AutoPage
- pdf2video: https://github.com/DangJin/pdf2video
- Idea2Paper: https://github.com/AgentAlphaAGI/Idea2Paper
- PaperX: https://github.com/yutao1024/PaperX
- figures4papers: https://github.com/ChenLiu-1996/figures4papers
- PaperBanana: https://github.com/dwzhu-pku/PaperBanana
- PaperBanana-Pro: https://github.com/elpsykongloo/PaperBanana-Pro
- AutoFigure: https://github.com/ResearAI/AutoFigure
- FigureWeave: https://github.com/Krisocer/FigureWeave
- EditDeck: https://github.com/Morgensonne/EditDeck
- AutoFigure-Edit: https://github.com/ResearAI/AutoFigure-Edit
- Kahneman4Review: https://huggingface.co/spaces/nuojohnchen/Kahneman4Review
- Academic Figure Generator: https://github.com/LigphiDonk/academic-figure-generator
- PaperFit: https://github.com/OpenRaiser/PaperFit
全自动科研
-
EvoScientist: https://github.com/EvoScientist/EvoScientist
-
Auto-claude-code-research-in-sleep: https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep
-
ArgusBot: https://github.com/waltstephen/ArgusBot
-
Dr.Claw: https://github.com/OpenLAIR/dr-claw
-
Redigg: https://github.com/redigg/redigg
-
AutoResearchClaw: https://github.com/aiming-lab/AutoResearchClaw
-
NanoResearch: https://github.com/OpenRaiser/NanoResearch
-
ScienceClaw: https://github.com/AgentTeam-TaichuAI/ScienceClaw
-
EurekaClaw: https://github.com/EurekaClaw/EurekaClaw
-
Claude-scholar: https://github.com/Galaxy-Dawn/claude-scholar/
-
claude-scientific-skills: https://github.com/K-Dense-AI/claude-scientific-skills
-
K-Dense BYOK: https://github.com/K-Dense-AI/k-dense-byok
-
latex-paper-skills: https://github.com/yunshenwuchuxun/latex-paper-skills
-
AutoResearch : https://github.com/karpathy/autoresearch
-
RD-Agent : https://github.com/microsoft/RD-Agent
-
DeepScientist : https://github.com/ResearAI/DeepScientist
-
Deep Researcher Agent: https://github.com/Xiangyue-Zhang/auto-deep-researcher-24x7
-
academic-research-skills: https://github.com/Imbad0202/academic-research-skills
-
Supervisor-Skills: https://github.com/HKUSTDial/Supervisor-Skills
代码 Coding
-
Cloi CLI: Local debugging agent that runs in your terminal.
视频 Video
模型
[!NOTE] 🤝Awesome-Video-Diffusion
- HunyuanVideo
- CogVideo
- Wan2.1
- Open-Sora
- Open-Sora-Plan
- LTX-Video
- Step-Video-T2V
- Step1X-Edit
Editing - Wan2.1-VACE
Editing - ICEdit
Editing - mochi-1-preview
- Wan2.1-Fun
- Wan2.1-FLF2V
首尾帧 - MAGI-1
自回归模型 - SkyReels-V2
- FramePack
- Pusa-VidGen
- Wan2.2
- MoGA
长视频 - LongCat-Video
- HunyuanVideo-1.5
- LTX-2
- daVinci-MagiHuman
- LongLive
- JoyAI-Echo
- NAVA
- LTX-2.3 (
🔥) - LingBot-Video
- MiniMax-H3 (
🔥) - MAGI-2-preview
- LTX-2.5 (
🔥)
编辑
- Wan2.1-VACE-14B: https://huggingface.co/Wan-AI/Wan2.1-VACE-14B
- Ditto: https://github.com/EzioBy/Ditto
- Bernini (
🔥): https://github.com/bytedance/Bernini - JoyAI-Video-Edit: https://huggingface.co/jdopensource/JoyAI-Video-Edit
训练
- https://github.com/hao-ai-lab/FastVideo
- https://github.com/tdrussell/diffusion-pipe
- https://github.com/VideoVerses/VideoTuna
- (
🔥) https://github.com/modelscope/DiffSynth-Studio - https://github.com/huggingface/diffusers
- https://github.com/kohya-ss/musubi-tuner
- https://github.com/spacepxl/HunyuanVideo-Training
- https://github.com/Tele-AI/TeleTron
- https://github.com/Yaofang-Liu/Mochi-Full-Finetuner
- https://github.com/bghira/SimpleTuner
- https://github.com/X-GenGroup/Flow-Factory
- https://github.com/shengshu-ai/minWM
world model
推理
实用工具
-
PySceneDetect: Python and OpenCV-based scene cut/transition detection program & library.
-
DOVER: Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives.
-
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding.
图片 Image
生成
- awesome-nano-banana
- Awesome-Nano-Banana-images
- HunyuanImage-3.0:https://github.com/Tencent-Hunyuan/HunyuanImage-3.0
- Seedream 4.0:https://arxiv.org/abs/2509.20427
- LongCat-Image:https://huggingface.co/meituan-longcat/LongCat-Image
- Z-Image-Turbo:https://huggingface.co/Tongyi-MAI/Z-Image-Turbo
- Qwen-Image:https://huggingface.co/Qwen/Qwen-Image
- Qwen-Image-2512:https://huggingface.co/Qwen/Qwen-Image-2512
- Z-Image:https://huggingface.co/Tongyi-MAI/Z-Image
- ERNIE-Image: https://huggingface.co/baidu/ERNIE-Image
- ERNIE-Image-Turbo: https://huggingface.co/Baidu/ERNIE-Image-Turbo
- Nucleus-Image: https://huggingface.co/NucleusAI/Nucleus-Image
- HiDream-O1-Image: https://huggingface.co/HiDream-ai/HiDream-O1-Image
- Ideogram 4: https://github.com/ideogram-oss/ideogram4
- Boogu-Image: https://github.com/boogu-project/Boogu-Image
- Krea-2-Raw: https://huggingface.co/krea/Krea-2-Raw
- SenseNova-U1: https://github.com/OpenSenseNova/SenseNova-U1
- Mage-Flow: https://huggingface.co/collections/microsoft/mage
- Qwen-Image-2.1: https://huggingface.co/Qwen/Qwen-Image-2.1
- Ming-Image-0.1: https://huggingface.co/inclusionAI/Ming-Image-0.1-Design
编辑
- ChronoEdit-14B: https://huggingface.co/nvidia/ChronoEdit-14B-Diffusers
- Eigen-Banana-Qwen-Image-Edit: https://huggingface.co/eigen-ai-labs/eigen-banana-qwen-image-edit
- Qwen-Image-Edit-2509: https://huggingface.co/Qwen/Qwen-Image-Edit-2509
- Upscale: https://huggingface.co/vafipas663/Qwen-Edit-2509-Upscale-LoRA
- Multiple-angles: https://huggingface.co/dx8152/Qwen-Edit-2509-Multiple-angles
- Multi-Angle-Lighting: https://huggingface.co/dx8152/Qwen-Edit-2509-Multi-Angle-Lighting
- LongCat-Image-Edit: https://huggingface.co/meituan-longcat/LongCat-Image-Edit
- Qwen-Image-Edit-2511: https://huggingface.co/Qwen/Qwen-Image-Edit-2511
- Qwen-Image-Edit-2511-Upscale2K: https://huggingface.co/valiantcat/Qwen-Image-Edit-2511-Upscale2K
- Qwen-Image-Edit-2511-Multiple-Angles-LoRA: https://huggingface.co/fal/Qwen-Image-Edit-2511-Multiple-Angles-LoRA
- FireRed-Image-Edit: https://huggingface.co/FireRedTeam/FireRed-Image-Edit-1.0
- JoyAI-Image-Edit: https://huggingface.co/jdopensource/JoyAI-Image-Edit
统一
- GLM-Image: https://huggingface.co/zai-org/GLM-Image
- https://huggingface.co/black-forest-labs/FLUX.2-klein-4B
- DreamLite: https://github.com/ByteVisionLab/DreamLite
- SenseNova-U1: https://huggingface.co/sensenova/SenseNova-U1-8B-MoT-SFT
训练
- Ostris:https://github.com/ostris/ai-toolkit
- FlymyAI:https://github.com/FlyMyAI/flymyai-lora-trainer
- Nitro-T:https://github.com/AMD-AGI/Nitro-T
- (
🔥) DiffSynth-Studio:https://github.com/modelscope/DiffSynth-Studio - Musubi Tuner: https://github.com/kohya-ss/musubi-tuner
- SimpleTuner: https://github.com/bghira/SimpleTuner
- MS Training: https://www.modelscope.cn/aigc/modelTraining
- Finetune HunyuanImage-3.0: https://github.com/PhotonAISG/hunyuan-image3-finetune
- OneTrainer: https://github.com/Nerogar/OneTrainer
- Finetune LongCat-Image and Edit: https://github.com/meituan-longcat/LongCat-Image/tree/main/train_examples
- Flow-Factory: https://github.com/X-GenGroup/Flow-Factory
- UniRL: https://github.com/Tencent-Hunyuan/UniRL
评估
- ULMEvalKit:https://github.com/ULMEvalKit/ULMEvalKit
推理
-
TypemovieInfer: https://github.com/typemovie/TypemovieInfer
搜索 Search
-
OpenSearch GPT: SearchGPT / Perplexity clone, but personalised for you.
-
MindSearch: An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT).
-
nanoPerplexityAI: The simplest open-source implementation of perplexity.ai.
-
curiosity: Try to build a Perplexity-like user experience.
-
MiniPerplx: A minimalistic AI-powered search engine that helps you find information on the internet.
语音 Speech
TTS
- SpeechGPT-2.0-preview: https://github.com/OpenMOSS/SpeechGPT-2.0-preview
- Moss-TTSD:https://github.com/OpenMOSS/MOSS-TTSD
- Index-TTS:https://github.com/index-tts/index-tts
- MegaTTS3:https://github.com/bytedance/MegaTTS3
- F5-TTS:https://github.com/SWivid/F5-TTS
- GPT-SoVITS:https://github.com/RVC-Boss/GPT-SoVITS
- CosyVoice:https://github.com/FunAudioLLM/CosyVoice
- Spark-TTS:https://github.com/SparkAudio/Spark-TTS
- OpenVoice:https://github.com/myshell-ai/OpenVoice
- Dia:https://github.com/nari-labs/dia
- ChatTTS:https://github.com/2noise/ChatTTS
- Fish Speech:https://github.com/fishaudio/fish-speech
- Edge-TTS:https://github.com/rany2/edge-tts
- Bark:https://github.com/suno-ai/bark
- kokoro: https://github.com/hexgrad/kokoro
- Higgs Audio V2: https://github.com/boson-ai/higgs-audio 【Training】
- KittenTTS: https://github.com/KittenML/KittenTTS
- ZipVoice: https://github.com/k2-fsa/ZipVoice
- VyvoTTS: https://github.com/Vyvo-Labs/VyvoTTS
- VibeVoice: https://github.com/microsoft/VibeVoice
- Index-TTS-2: https://huggingface.co/IndexTeam/IndexTTS-2
- FireRedTTS2: https://github.com/FireRedTeam/FireRedTTS2
- VoxCPM: https://github.com/OpenBMB/VoxCPM/
- Neutts-Air: https://github.com/neuphonic/neutts-air
- Maya1: https://huggingface.co/maya-research/maya1
- VibeVoice: https://huggingface.co/collections/microsoft/vibevoice
- GLM-TTS: https://github.com/zai-org/GLM-TTS
- Fun-CosyVoice3: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512
- Qwen3-TTS:https://huggingface.co/collections/Qwen/qwen3-tts
- Ming-Omni-TTS: https://github.com/inclusionAI/Ming-omni-tts
- VoxCPM2: https://huggingface.co/openbmb/VoxCPM2
- OmniVoice: https://github.com/k2-fsa/OmniVoice
- MOSS-TTS-v1.5: https://huggingface.co/OpenMOSS-Team/MOSS-TTS-v1.5
- Higgs Audio v3 TTS: https://huggingface.co/bosonai/higgs-audio-v3-tts-4b
- Confucius4-TTS: https://github.com/netease-youdao/Confucius4-TTS
- Dia-1.6B: https://huggingface.co/nari-labs/Dia-1.6B-0626
- FireRedTTS3: https://github.com/FireRedTeam/FireRedTTS3
STT/ASR
- Kyutai: https://github.com/kyutai-labs/delayed-streams-modeling
- Whisper: https://github.com/openai/whisper
- Audio Flamingo 3: https://huggingface.co/nvidia/audio-flamingo-3
- Voxtral: https://huggingface.co/mistralai/Voxtral-Mini-3B-2507
- Step-Audio2: https://github.com/stepfun-ai/Step-Audio2
- SoulX-Podcast: https://huggingface.co/collections/Soul-AILab/soulx-podcast
- Omnilingual ASR: https://github.com/facebookresearch/omnilingual-asr
- Fun-ASR: https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512
- FunASR: https://github.com/modelscope/FunASR
- SenseVoice: https://github.com/FunAudioLLM/SenseVoice
- VibeVoice-ASR: https://huggingface.co/microsoft/VibeVoice-ASR
- Qwen3-ASR: https://github.com/QwenLM/Qwen3-ASR
- Mega-ASR: https://github.com/xzf-thu/Mega-ASR
- MOSS-Transcribe-Diarize: https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize
Voice Interaction
-
Fun-Audio-Chat: https://huggingface.co/FunAudioLLM/Fun-Audio-Chat-8B
-
Chroma 1.0: https://huggingface.co/FlashLabs/Chroma-4B
-
AudioInteraction: https://huggingface.co/zhifeixie/AudioInteraction
世界模型 World Models
模型
- ShadowDancer: https://github.com/AlayaLab/ShadowDancer
- PhiZero: https://github.com/yaoyao-jpg/PhiZero
- Wonder: https://arxiv.org/pdf/2607.26037
- open-dreamer: https://github.com/next-state/open-dreamer
- Cosmos3-Edge: https://huggingface.co/nvidia/Cosmos3-Edge
- WorldWander: https://github.com/showlab/WorldWander
- Matrix-Game-3.5: https://github.com/Riemann-Dynamics/Matrix-Game-3.5
- ABot-World: https://github.com/amap-cvlab/ABot-World
- LingBot-World-V2: https://huggingface.co/collections/robbyant/lingbot-world-v2
- AlayaWorld: https://github.com/AlayaLab/AlayaWorld
- MoWorld: https://moxin-tech.github.io/moworld/
- Warp-as-History: https://github.com/yyfz/Warp-as-History
- MIRA: https://github.com/mira-wm/mira
Multiplayer - ActWorld: https://interactwm.github.io/ActWorld/
- Sana-WM: https://nvlabs.github.io/Sana/WM/
- Cosmos-3: https://github.com/nvidia/cosmos
- DreamX-World-5B: https://huggingface.co/GD-ML/DreamX-World-5B
- DreamX-World-5B-Cam: https://modelscope.cn/models/GD-ML/DreamX-World-5B-Cam
- Astra: https://eternalevan.github.io/Astra-project/
- Yume-5B: https://huggingface.co/stdstu123/Yume-5B-720P
- LingBot-World: https://github.com/robbyant/lingbot-world
- HY-WorldPlay: https://huggingface.co/tencent/HY-WorldPlay
- Matrix-Game-3.0: https://huggingface.co/Skywork/Matrix-Game-3.0
- Matrix-Game-2.0: https://huggingface.co/Skywork/Matrix-Game-2.0
- Waypoint-1.5-1B: https://huggingface.co/Overworld/Waypoint-1.5-1B
- Hunyuan-GameCraft: https://hunyuan-gamecraft.github.io/
- AlayaWorld: https://github.com/AlayaLab/AlayaWorld
- WorldWander: https://github.com/showlab/WorldWander
框架
-
nano-world-model: https://github.com/simchowitzlabpublic/nano-world-model
-
stable-worldmodel: https://github.com/galilai-group/stable-worldmodel
-
OpenWorldLib: https://github.com/OpenDCAI/OpenWorldLib
龙虾 OpenClaw
-
MultiUserClaw: https://github.com/johnson7788/MultiUserClaw
-
ClawManager: https://github.com/Yuan-lab-LLM/ClawManager
-
OpenHanako: https://github.com/liliMozi/openhanako
统一模型 Unified Model
现在统一模型已经从
理解+生成变成理解+生成+编辑
-
Janus-Pro:http://arxiv.org/abs/2508.05954
-
Any-GPT:https://arxiv.org/abs/2402.12226
-
Next-GPT:https://arxiv.org/pdf/2309.05519.pdf
-
Dream-LLM:https://arxiv.org/abs/2309.11499
-
Chameleon:https://arxiv.org/abs/2405.09818
-
MedViLaM:https://arxiv.org/abs/2409.19684
-
TokenFlow:https://github.com/ByteFlow-AI/TokenFlow
-
OneDiffusion:https://github.com/lehduong/OneDiffusion
-
MetaMorph: https://arxiv.org/abs/2412.14164
-
LlamaFusion:https://arxiv.org/abs/2412.15188
-
InstructSeg:https://arxiv.org/abs/2412.14006
-
ILLUME: https://arxiv.org/abs/2412.06673
-
SynerGen-VL:https://arxiv.org/abs/2412.09604
-
Align Anything:https://arxiv.org/abs/2412.15838
-
Transfusion: https://arxiv.org/abs/2408.11039
-
JanusFlow: https://arxiv.org/abs/2411.07975
-
HealthGPT:https://arxiv.org/abs/2502.09838
Medical -
Qwen2.5-Omni:https://arxiv.org/abs/2503.20215
-
Bifrost-1:https://arxiv.org/abs/2508.05954
-
OmniGen2:https://arxiv.org/abs/2506.18871
-
VeOmni:https://github.com/ByteDance-Seed/VeOmni
Training -
NextStep-1:https://arxiv.org/abs/2508.10711
-
UniUGG: https://arxiv.org/abs/2508.11952
3D -
Omni-Video:https://arxiv.org/abs/2507.06119
-
Lumina-DiMOO:https://github.com/Alpha-VLLM/Lumina-DiMOO
-
Hyper-Bagel:https://arxiv.org/abs/2509.18824
-
Ming-UniVision:https://arxiv.org/abs/2510.06590
-
EditVerse:https://arxiv.org/abs/2509.20360
-
LightBagel: https://arxiv.org/abs/2510.22946
-
DreamLLM: https://arxiv.org/abs/2309.11499
-
X-Omni: https://arxiv.org/abs/2507.22058
-
Ming-flash-omni-Preview: https://huggingface.co/inclusionAI/Ming-flash-omni-Preview
-
Omni-View: https://arxiv.org/abs/2511.07222
-
NExT-OMNI: https://arxiv.org/abs/2510.13721
-
Uni-MoE-2.0-Omni: https://arxiv.org/abs/2511.12609
-
LongCat-Flash-Omni: https://huggingface.co/meituan-longcat/LongCat-Flash-Omni
-
ShapeLLM-Omni: https://arxiv.org/abs/2506.01853
-
UniGen-1.5: https://arxiv.org/abs/2511.14760
-
UniModel: https://arxiv.org/abs/2511.16917
-
HBridge: https://arxiv.org/abs/2511.20520
-
OpenOmni: https://github.com/RainBowLuoCS/OpenOmni
-
Ming-Flash-Omni: https://arxiv.org/abs/2510.24821
-
InternVL-U: https://github.com/OpenGVLab/InternVL-U
-
LongCat-Next: https://github.com/meituan-longcat/LongCat-Next
-
SenseNova-U1: https://huggingface.co/sensenova/SenseNova-U1-8B-MoT-SFT
-
JoyAI-VL-Interaction: https://arxiv.org/abs/2606.14777
Interaction -
MOSS-VL-Realtime: https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime
Interaction -
SenseNova-Vision: https://huggingface.co/collections/sensenova/sensenova-vision
-
Mage-VL: https://huggingface.co/microsoft/Mage-VL
Interaction
书籍 Book
-
Taming LLMs: A Practical Guide to LLM Pitfalls with Open Source Software
-
《The Smol Training Playbook: The Secrets to Building World-Class LLMs》
课程 Course
-
ACL 2023 Tutorial: Retrieval-based Language Models and Applications
-
llm-course: Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
-
普林斯顿 COS 597G (Fall 2022): Understanding Large Language Models
-
openai-cookbook: Examples and guides for using the OpenAI API.
-
Hands on llms: Learn about LLM, LLMOps, and vector DBS for free by designing, training, and deploying a real-time financial advisor LLM system.
-
LangGPT: Empowering everyone to become a prompt expert!
-
build nanoGPT: Video+code lecture on building nanoGPT from scratch.
-
LLM101n: Let's build a Storyteller.
-
Smol Vision: Recipes for shrinking, optimizing, customizing cutting edge vision models.
-
RAG++ : From POC to production: Advanced RAG course.
-
Weights & Biases AI Academy: Finetuning, building with LLMs, Structured outputs and more LLM courses.
-
Learn RAG From Scratch – Python AI Tutorial from a LangChain Engineer
-
RAG_Techniques: This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. RAG systems combine information retrieval with generative models to provide accurate and contextually rich responses.
-
NanoChat: The best ChatGPT that $100 can buy.
教程 Tutorial
论文 Paper
[!NOTE] 🤝Huggingface Daily Papers、Cool Papers、ML Papers Explained
-
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
-
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
-
DataComp-LM: In search of the next generation of training sets for language models
-
MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series
-
Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model
-
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
-
Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent
-
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models
-
TÜLU 3: Pushing Frontiers in Open Language Model Post-Training
-
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
-
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
-
Baichuan-M1: Pushing the Medical Capability of Large Language Models
-
Predictable Scale: Part I -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
-
SkyLadder: Better and Faster Pretraining via Context Window Scheduling
-
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
-
Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
-
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
-
MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining
-
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training 85M-Midtraining Data 22M Instruct Data
-
Olmo3: Charting a path through the model flow to lead open-source AI. Website
社区 Community
模型上下文协议 MCP
- MCP是啥?技术原理是什么?一个视频搞懂MCP的一切。Windows系统配置MCP,Cursor,Cline 使用MCP
- MCP是什么?为啥是下一代AI标准?MCP原理+开发实战!在Cursor、Claude、Cline中使用MCP,让AI真正自动化!
- 从零编写MCP并发布上线,超简单!手把手教程
MCP工具聚合:
技能 Skills
推理 Open o1
[!NOTE]
开放的技术是我们永恒的追求
- https://github.com/atfortes/Awesome-LLM-Reasoning
- https://github.com/hijkzzz/Awesome-LLM-Strawberry
- https://github.com/wjn1996/Awesome-LLM-Reasoning-Openai-o1-Survey
- https://github.com/srush/awesome-o1
- https://github.com/open-thought/system-2-research
- https://github.com/ninehills/blog/issues/121
- https://github.com/OpenSource-O1/Open-O1
- https://github.com/GAIR-NLP/O1-Journey
- https://github.com/marlaman/show-me
- https://github.com/bklieger-groq/g1
- https://github.com/Jaimboh/Llamaberry-Chain-of-Thought-Reasoning-in-AI
- https://github.com/pseudotensor/open-strawberry
- https://huggingface.co/collections/peakji/steiner-preview-6712c6987110ce932a44e9a6
- https://github.com/SimpleBerry/LLaMA-O1
- https://huggingface.co/collections/Skywork/skywork-o1-open-67453df58e12f6c3934738d0
- https://huggingface.co/collections/Qwen/qwq-674762b79b75eac01735070a
- https://github.com/SkyworkAI/skywork-o1-prm-inference
- https://github.com/RifleZhang/LLaVA-Reasoner-DPO
- https://github.com/ADaM-BJTU
- https://github.com/ADaM-BJTU/OpenRFT
- https://github.com/RUCAIBox/Slow_Thinking_with_LLMs
- https://github.com/richards199999/Thinking-Claude
- https://huggingface.co/AGI-0/Art-v0-3B
- https://huggingface.co/deepseek-ai/DeepSeek-R1
- https://huggingface.co/deepseek-ai/DeepSeek-R1-Zero
- https://github.com/huggingface/open-r1
- https://github.com/hkust-nlp/simpleRL-reason
- https://github.com/Jiayi-Pan/TinyZero
- https://github.com/baichuan-inc/Baichuan-M1-14B
- https://github.com/EvolvingLMMs-Lab/open-r1-multimodal
- https://github.com/open-thoughts/open-thoughts
- Mini-R1: https://www.philschmid.de/mini-deepseek-r1
- LLaMA-Berry: https://arxiv.org/abs/2410.02884
- MCTS-DPO: https://arxiv.org/abs/2405.00451
- OpenR: https://github.com/openreasoner/openr
- https://arxiv.org/abs/2410.02725
- LLaVA-o1: https://arxiv.org/abs/2411.10440
- Marco-o1: https://arxiv.org/abs/2411.14405
- OpenAI o1 report: https://openai.com/index/deliberative-alignment
- DRT-o1: https://github.com/krystalan/DRT-o1
- Virgo:https://arxiv.org/abs/2501.01904
- HuatuoGPT-o1:https://arxiv.org/abs/2412.18925
- o1 roadmap:https://arxiv.org/abs/2412.14135
- Mulberry:https://arxiv.org/abs/2412.18319
- https://arxiv.org/abs/2412.09413
- https://arxiv.org/abs/2501.02497
- Search-o1:https://arxiv.org/abs/2501.05366v1
- https://arxiv.org/abs/2501.18585
- https://github.com/simplescaling/s1
- https://github.com/Deep-Agent/R1-V
- https://github.com/StarRing2022/R1-Nature
- https://github.com/Unakar/Logic-RL
- https://github.com/datawhalechina/unlock-deepseek
- https://github.com/GAIR-NLP/LIMO
- https://github.com/Zeyi-Lin/easy-r1
- https://github.com/jackfsuia/nanoRLHF/tree/main/examples/r1-v0
- https://github.com/FanqingM/R1-Multimodal-Journey
- https://github.com/dhcode-cpp/X-R1
- https://github.com/agentica-project/deepscaler
- https://github.com/ZihanWang314/RAGEN
- https://github.com/sail-sg/oat-zero
- https://github.com/TideDra/lmm-r1
- https://github.com/FlagAI-Open/OpenSeek
- https://github.com/SwanHubX/ascend_r1_turtorial
- https://github.com/om-ai-lab/VLM-R1
- https://github.com/wizardlancet/diagnosis_zero
- https://github.com/lsdefine/simple_GRPO
- https://github.com/brendanhogan/DeepSeekRL-Extended
- https://github.com/Wang-Xiaodong1899/Open-R1-Video
- https://github.com/lsdefine/simple_GRPO
- https://github.com/Open-Reasoner-Zero/Open-Reasoner-Zero
- https://github.com/lucasjinreal/Namo-R1
- https://github.com/hiyouga/EasyR1
- https://github.com/Fancy-MLLM/R1-Onevision
- https://github.com/tulerfeng/Video-R1
- https://huggingface.co/qihoo360/TinyR1-32B-Preview
- https://github.com/facebookresearch/swe-rl
- https://github.com/turningpoint-ai/VisualThinker-R1-Zero
- https://github.com/yuyq96/R1-Vision
- https://github.com/sungatetop/deepseek-r1-vision
- https://huggingface.co/qihoo360/Light-R1-32B
- https://github.com/Liuziyu77/Visual-RFT
- https://github.com/Mohammadjafari80/GSM8K-RLVR
- https://github.com/ModalMinds/MM-EUREKA
- https://github.com/joey00072/nanoGRPO
- https://github.com/PeterGriffinJin/Search-R1
- https://openi.pcl.ac.cn/PCL-Reasoner/GRPO-Training-Suite
- https://github.com/dvlab-research/Seg-Zero
- https://github.com/HumanMLLM/R1-Omni
- https://github.com/OpenManus/OpenManus-RL
- https://arxiv.org/pdf/2503.07536
- https://github.com/Osilly/Vision-R1
- https://github.com/LengSicong/MMR1
- https://github.com/phonism/CP-Zero
- https://github.com/SkyworkAI/Skywork-R1V
- https://arxiv.org/abs/2503.13939v1
- https://github.com/0russwest0/Agent-R1
- https://github.com/MetabrainAGI/Awaker2.5-R1
- https://github.com/LG-AI-EXAONE/EXAONE-Deep
- https://github.com/qiufengqijun/open-r1-reprod
- https://github.com/SUFE-AIFLM-Lab/Fin-R1
- https://github.com/sail-sg/understand-r1-zero
- https://github.com/baibizhe/Efficient-R1-VLLM
- https://github.com/hkust-nlp/simpleRL-reason
- https://arxiv.org/abs/2502.19655
- https://arxiv.org/abs/2503.21620v1
- https://arxiv.org/abs/2503.16081
- https://github.com/ShadeCloak/ADORA
- https://github.com/appletea233/Temporal-R1
- https://github.com/inclusionAI/AReaL
- https://github.com/lzhxmu/CPPO
- https://arxiv.org/abs/2503.23829
- https://github.com/TencentARC/SEED-Bench-R1
- https://github.com/McGill-NLP/nano-aha-moment
- https://github.com/VLM-RL/Ocean-R1
- https://github.com/OpenGVLab/VideoChat-R1
- https://github.com/ByteDance-Seed/Seed-Thinking-v1.5
- https://github.com/SkyworkAI/Skywork-OR1
- https://github.com/MoonshotAI/Kimi-VL
- https://arxiv.org/abs/2504.08600
- https://github.com/ZhangXJ199/TinyLLaVA-Video-R1
- https://arxiv.org/abs/2504.11914
- https://github.com/policy-gradient/GRPO-Zero
- https://github.com/linkangheng/PR1
- https://github.com/jiangxinke/Agentic-RAG-R1
- https://github.com/shangshang-wang/Tina
- https://github.com/aliyun/qwen-dianjin
- https://github.com/RAGEN-AI/RAGEN
- https://github.com/XiaomiMiMo/MiMo
- https://github.com/yuanzhoulvpi2017/nano_rl
- https://huggingface.co/a-m-team/AM-Thinking-v1
- https://huggingface.co/Intelligent-Internet/II-Medical-8B
- https://github.com/CSfufu/Revisual-R1
[↥ back to top](#Contents)
推理 Open o3
-
Mini-o3: https://arxiv.org/abs/2509.07969
-
Simple-o3: https://arxiv.org/abs/2508.12109
-
Open o3 Video: https://arxiv.org/abs/2510.20579
小语言模型 Small Language Model
-
https://github.com/loubnabnl/nanotron-smol-cluster (使用Cosmopedia训练cosmo-1b)
-
https://huggingface.co/Nanbeige/Nanbeige4-3B-Thinking-2511
23T tokens预训练模型
小多模态模型 Small Vision Language Model
-
https://github.com/yuanzhoulvpi2017/zero_nlp/tree/main/train_llava
技巧 Tips
-
What We Learned from a Year of Building with LLMs (Part III): Strategy
-
LLMs for Text Classification: A Guide to Supervised Learning
-
Unsupervised Text Classification: Categorize Natural Language With LLMs
-
Text Classification With LLMs: A Roundup of the Best Methods
-
MiniMind: 3小时完全从0训练一个仅有26M的小参数GPT,最低仅需2G显卡即可推理训练.
-
LLM-Travel: 致力于深入理解、探讨以及实现与大模型相关的各种技术、原理和应用
-
Reader-LM: Small Language Models for Cleaning and Converting HTML to Markdown
-
pytorch-llama: LLaMA 2 implemented from scratch in PyTorch.
-
Preference Optimization for Vision Language Models with TRL 【support model】
-
Distributed Training Guide: Best practices & guides on how to write distributed pytorch training code.
贡献者:
如果你觉得本项目对你有帮助,欢迎引用:
@misc{wang2024llm,
title={awesome-LLM-resourses},
author={Rongsheng Wang},
year={2024},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/WangRongsheng/awesome-LLM-resourses}},
}


