
AI Game DevTools (AI-GDT) 🎮
Your AI Game Dev Hub. 🚀
The ultimate resource hub for AI-powered game development tools. Discover cutting-edge LLMs, World Model, Agent, Code, Image, Texture, Shader, 3D Model, Animation, Video, Audio, Music, Singing Voice and Analytics. 🔥
Website | 官方网站
Table of Contents
- LLM (LLM & Tool)
- VLM (Visual)
- Game (World Model & Agent)
- Code
- Image
- Texture
- Shader
- 3D Model
- Avatar
- Animation
- Video
- Audio
- Music
- Singing Voice
- Speech
- Analytics
Project List
LLM (LLM & Tool)
| Source | Description | Paper | Game Engine | Type |
|---|---|---|---|---|
| AgentGPT | 🤖 Assemble, configure, and deploy autonomous AI Agents in your browser. | Tool | ||
| AICommand | ChatGPT integration with Unity Editor. | Unity | Tool | |
| AIOS | LLM Agent Operating System. | Tool | ||
| AI Scientist | The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. | arXiv | Tool | |
| Assistant CLI | A comfortable CLI tool to use ChatGPT service🔥 | Tool | ||
| Auferet | An AI game master for solo text adventures and tabletop-style RPGs, with persistent memory of your story and your own uploaded lore. | Writer | ||
| Auto-GPT | An experimental open-source attempt to make GPT-4 fully autonomous. | Tool | ||
| BabyAGI | This Python script is an example of an AI-powered task management system. | Tool | ||
| 👶🤖🖥️ BabyAGI UI | BabyAGI UI is designed to make it easier to run and develop with babyagi in a web app, like a ChatGPT. | Tool | ||
| baichuan-7B | A large-scale 7B pretraining language model developed by Baichuan. | Tool | ||
| Baichuan-13B | A 13B large language model developed by Baichuan Intelligent Technology. | Tool | ||
| Baichuan 2 | A series of large language models developed by Baichuan Intelligent Technology. | Tool | ||
| Bisheng | Bisheng is an open LLM devops platform for next generation AI applications. | Tool | ||
| Character-LLM | A Trainable Agent for Role-Playing. | arXiv | Tool | |
| ChatDev | Communicative Agents for Software Development. | arXiv | Tool | |
| ChatGPT-API-unity | Binds ChatGPT chat completion API to pure C# on Unity. | Unity | Tool | |
| ChatGPTForUnity | ChatGPT for unity. | Unity | Tool | |
| ChatRWKV | ChatRWKV is like ChatGPT but powered by RWKV (100% RNN) language model, and open source. | Tool | ||
| ChatYuan | Large Language Model for Dialogue in Chinese and English. | Tool | ||
| Chinese-LLaMA-Alpaca-3 | (Chinese Llama-3 LLMs) developed from Meta Llama 3. | Tool | ||
| Chrome-GPT | An AutoGPT agent that controls Chrome on your desktop. | Tool | ||
| CogVLM | CogVLM, a powerful open-source visual language foundation model. | arXiv | Tool | |
| CoreNet | A library for training deep neural networks. | Tool | ||
| Cosmos | Cosmos is a world model development platform that consists of world foundation models, tokenizers and video processing pipeline to accelerate the development of Physical AI at Robotics & AV labs. | LLM | ||
| DBRX | DBRX is a large language model trained by Databricks. | Tool | ||
| DCLM | DataComp for Language Models. | arXiv | Tool | |
| DeepSeek-R1 | DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. | LLM | ||
| DeepSeek-V3 | DeepSeek-V3 is a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. | arXiv | LLM | |
| DemoGPT | Auto Gen-AI App Generator with the Power of Llama 2 | Tool | ||
| Design2Code | Automating Front-End Engineering | Tool | ||
| Devika | Devika is an Agentic AI Software Engineer. | Tool | ||
| Devon | An open-source pair programmer. | Tool | ||
| Dora | Generating powerful websites, one prompt at a time. | Tool | ||
| Flowise | Drag & drop UI to build your customized LLM flow using LangchainJS. | Tool | ||
| Gemini | Gemini is built from the ground up for multimodality — reasoning seamlessly across text, images, video, audio, and code. | Tool | ||
| Gemma | Gemma is a family of lightweight, state-of-the art open models built from research and technology used to create Google Gemini models. | Tool | ||
| gemma.cpp | lightweight, standalone C++ inference engine for Google's Gemma models. | Tool | ||
| GLM-4 | GLM-4-9B is the open-source version of the latest generation of pre-trained models in the GLM-4 series launched by Zhipu AI. | Tool | ||
| GLM-4.5 | GLM-4.5: An open-source large language model designed for intelligent agents by Z.ai. | LLM | ||
| GPT4All | A chatbot trained on a massive collection of clean assistant data including code, stories and dialogue. | Tool | ||
| GPT-4o | GPT-4o (“o” for “omni”) is a step towards much more natural human-computer interaction—it accepts as input any combination of text, audio, image, and video and generates any combination of text, audio, and image outputs. | Tool | ||
| gpt-oss | gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI. | LLM | ||
| GPTScript | Develop LLM Apps in Natural Language. | Tool | ||
| Grok-1 | The weights and architecture of our 314 billion parameter Mixture-of-Experts model, Grok-1. | Tool | ||
| HuggingChat | Making the community's best AI chat models available to everyone. | Tool | ||
| Hugging Face API Unity Integration | This Unity package provides an easy-to-use integration for the Hugging Face Inference API, allowing developers to access and use Hugging Face AI models within their Unity projects. | Unity | Tool | |
| Hunyuan-MT | The Hunyuan-MT comprises a translation model, Hunyuan-MT-7B, and an ensemble model, Hunyuan-MT-Chimera. The translation model is used to translate source text into the target language, while the ensemble model integrates multiple translation outputs to produce a higher-quality result. | LLM | ||
| ImageBind | ImageBind One Embedding Space to Bind Them All. | arXiv | Tool | |
| Index-1.9B | A SOTA lightweight multilingual LLM. | Tool | ||
| InteractML-Unity | InteractML, an Interactive Machine Learning Visual Scripting framework for Unity3D. | Unity | Tool | |
| InteractML-Unreal Engine | Bringing Machine Learning to Unreal Engine. | Unreal Engine | Tool | |
| InternLM | InternLM has open-sourced a 7 billion parameter base model, a chat model tailored for practical scenarios and the training system. | arXiv | Tool | |
| InternLM-XComposer | InternLM-XComposer2 is a groundbreaking vision-language large model (VLLM) excelling in free-form text-image composition and comprehension. | arXiv | Tool | |
| Jan | Bring AI to your Desktop. | Tool | ||
| Janus | Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation. | arXiv | LLM | |
| Kimi K2 | Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. | LLM | ||
| Lamini | Lamini allows any engineering team to outperform general purpose LLMs through RLHF and fine- tuning on their own data. | Tool | ||
| LaMini-LM | LaMini-LM is a collection of small-sized, efficient language models distilled from ChatGPT and trained on a large-scale dataset of 2.58M instructions. | Tool | ||
| LangChain | LangChain is a framework for developing applications powered by language models. | Tool | ||
| LangFlow | ⛓️ LangFlow is a UI for LangChain, designed with react-flow to provide an effortless way to experiment and prototype flows. | Tool | ||
| LaVague | Automate automation with Large Action Model framework. | Tool | ||
| Lemur | Open Foundation Models for Language Agents. | Tool | ||
| Lepton AI | A Pythonic framework to simplify AI service building. | Tool | ||
| Lit-LLaMA | Implementation of the LLaMA language model based on nanoGPT. Supports flash attention, Int8 and GPTQ 4bit quantization, LoRA and LLaMA-Adapter fine-tuning, pre-training. | Tool | ||
| llama2-webui | Run Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). | Tool | ||
| Llama 3 | The official Meta Llama 3 GitHub site. | Tool | ||
| Llama 3.1 | Llama is an accessible, open large language model (LLM) designed for developers, researchers, and businesses to build, experiment, and responsibly scale their generative AI ideas. | Tool | ||
| LLaSM | Large Language and Speech Model. | Tool | ||
| LLM Answer Engine | Build a Perplexity-Inspired Answer Engine Using Next.js, Groq, Mixtral, Langchain, OpenAI, Brave & Serper. | Tool | ||
| llm.c | LLM training in simple, raw C/CUDA. | Tool | ||
| LLMUnity | Create characters in Unity with LLMs! | Unity | Tool | |
| LLocalSearch | LLocalSearch is a completely locally running search engine using LLM Agents. | Tool | ||
| LogicGamesSolver | A Python tool to solve logic games with AI, Deep Learning and Computer Vision. | Tool | ||
| LongCat-Flash | LongCat-Flash is a powerful and efficient language model with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates 18.6B∼31.3B parameters (averaging∼27B) based on contextual demands, optimizing both computational efficiency and performance. | LLM | ||
| LongWriter | LongWriter: Unleashing 10,000+ Word Generation From Long Context LLMs. | arXiv | Tool | |
| Large World Model (LWM) | Large World Model (LWM) is a general-purpose large-context multimodal autoregressive model. | arXiv | Tool | |
| Lumina-T2X | Lumina-T2X is a unified framework for Text to Any Modality Generation. | arXiv | Tool | |
| MetaGPT | The Multi-Agent Framework | Tool | ||
| MiniCPM-2B | An end-side LLM outperforms Llama2-13B. | Tool | ||
| MiniGPT-4 | Enhancing Vision-language Understanding with Advanced Large Language Models. | arXiv | Tool | |
| MiniGPT-5 | Interleaved Vision-and-Language Generation via Generative Vokens. | arXiv | Tool | |
| MiniMax-01 | MiniMax-01: Scaling Foundation Models with Lightning Attention. | arXiv | LLM | |
| Mixtral 8x7B | A high quality Sparse Mixture-of-Experts. | arXiv | Tool | |
| Mistral 7B | The best 7B model to date, Apache 2.0. | Tool | ||
| Mistral Large | Mistral Large is a new cutting-edge text generation model. It reaches top-tier reasoning capabilities. | Tool | ||
| MLC LLM | Enable everyone to develop, optimize and deploy AI models natively on everyone's devices. | Tool | ||
| MobiLlama | Towards Accurate and Lightweight Fully Transparent GPT. | arXiv | Tool | |
| MoE-LLaVA | Mixture of Experts for Large Vision-Language Models. | arXiv | Tool | |
| Moshi | Moshi is an experimental conversational AI. | Tool | ||
| Moshi | Moshi: a speech-text foundation model for real time dialogue. | Tool | ||
| MOSS | An open-source tool-augmented conversational language model from Fudan University. | Tool | ||
| mPLUG-Owl🦉 | Modularization Empowers Large Language Models with Multimodality. | arXiv | Tool | |
| Nemotron-4 | A 15-billion-parameter large multilingual language model trained on 8 trillion text tokens. | arXiv | Tool | |
| NExT-GPT | Any-to-Any Multimodal Large Language Model. | Tool | ||
| OLMo | Open Language Model | arXiv | Tool | |
| OmniLMM | Large multi-modal models for strong performance and efficient deployment. | Tool | ||
| OneLLM | One Framework to Align All Modalities with Language. | arXiv | Tool | |
| Open-Assistant | OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so. | Tool | ||
| Open Deep Research | An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. | LLM | ||
| OpenDevin | An autonomous AI software engineer. | Tool | ||
| Orion-14B | Orion-14B is a family of models includes a 14B foundation LLM, and a series of models. | arXiv | Tool | |
| Panda | Overseas Chinese open source large language model, based on Llama-7B, -13B, -33B, -65B for continuous pre-training in the Chinese field. | Tool | ||
| Perplexica | An AI-powered search engine. | Tool | ||
| Pi | AI chatbot designed for personal assistance and emotional support. | Tool | ||
| Qwen1.5 | Qwen1.5 is the improved version of Qwen. | Tool | ||
| Qwen2 | Qwen2 is the large language model series developed by Qwen team, Alibaba Cloud. | LLM | ||
| Qwen2.5-Coder | Qwen2.5-Coder is the code version of Qwen2.5, the large language model series developed by Qwen team, Alibaba Cloud. | arXiv | LLM | |
| Qwen-7B | The official repo of Qwen-7B (通义千问-7B) chat & pretrained large language model proposed by Alibaba Cloud. | LLM | ||
| Qwen3 | Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud. | arXiv | LLM | |
| RepoAgent | RepoAgent is an Open-Source project driven by Large Language Models(LLMs) that aims to provide an intelligent way to document projects. | arXiv | Tool | |
| s1 | s1: Simple test-time scaling. | arXiv | LLM | |
| Sanity AI Engine | Sanity AI Engine for the Unity Game Development Tool. | Unity | Tool | |
| SearchGPT | 🌳 Connecting ChatGPT with the Internet | Tool | ||
| Seed-OSS | Seed-OSS is a series of open-source large language models developed by ByteDance's Seed Team, designed for powerful long-context, reasoning, agent and general capabilities, and versatile developer-friendly features. | LLM | ||
| ShareGPT4V | Improving Large Multi-Modal Models with Better Captions. | Tool | ||
| SkyThought | Sky-T1: Train your own O1 preview model within $450. | LLM | ||
| Skywork | Skywork series models are pre-trained on 3.2TB of high-quality multilingual (mainly Chinese and English) and code data. | Tool | ||
| StableLM | Stability AI Language Models. | arXiv | Tool | |
| Stanford Alpaca | An Instruction-following LLaMA Model. | LLM | ||
| Text generation web UI | A gradio web UI for running Large Language Models like LLaMA, llama.cpp, GPT-J, OPT, and GALACTICA. | Tool | ||
| TinyChatEngine | On-Device LLM Inference Library. | Tool | ||
| ToolBench | An open platform for training, serving, and evaluating large language model for tool learning. | Tool | ||
| Unity ChatGPT | Unity ChatGPT Experiments. | Unity | Tool | |
| Unity OpenAI-API Integration | Integrate openai GPT-3 language model and ChatGPT API into a Unity project. | Unity | Tool | |
| Unreal Engine 5 Llama LoRA | A proof-of-concept project that showcases the potential for using small, locally trainable LLMs to create next-generation documentation tools. | Unreal Engine | Tool | |
| UnrealGPT | A collection of Unreal Engine 5 Editor Utility widgets powered by GPT3/4. | Unreal Engine | Tool | |
| Video-LLaVA | Learning United Visual Representation by Alignment Before Projection. | arXiv | Tool | |
| WebGPT | Run GPT model on the browser with WebGPU. | Tool | ||
| Web3-GPT | Deploy smart contracts with AI | Tool | ||
| WordGPT | 🤖 Bring the power of ChatGPT to Microsoft Word | Tool | ||
| XAgent | An Autonomous LLM Agent for Complex Task Solving. | Tool | ||
| Yi | A series of large language models trained from scratch by developers. | Tool | ||
| 01 Project | The open-source language model computer. | Tool | ||
| SimpleOllamaUnity | Ollama integration for Unity Engine (works in runtime and editor) | Unity | Tool | |
| AI-Writer | AI writes novels, generates fantasy and romance web articles, etc. Chinese pre-trained generative model. | Writer | ||
| Notebook.ai | Notebook.ai is a set of tools for writers, game designers, and roleplayers to create magnificent universes – and everything within them. | Writer | ||
| Novel | Notion-style WYSIWYG editor with AI-powered autocompletions. | Writer | ||
| NovelAI | Driven by AI, painlessly construct unique stories, thrilling tales, seductive romances, or just fool around. | Writer | ||
| Unity-MCP | Open-source MCP server connecting AI agents to the Unity Editor and runtime, with 100+ built-in tools. | Unity | Tool | |
| Godot-MCP | Open-source MCP server connecting AI agents to the Godot Editor and runtime (Godot 4.x, C#). | Godot | Tool | |
| Unreal-MCP | Open-source MCP server connecting AI agents to Unreal Engine 5.7, editor and runtime (C++ plugin + .NET sidecar). | Unreal Engine | Tool | |
| GameDev-MCP-Server | Open-source, engine-agnostic MCP server shared by Unity-MCP, Godot-MCP, and Unreal-MCP. | Unity/Godot/Unreal Engine | Tool | |
| MCP-Plugin-dotnet | Open-source .NET library/SDK that turns any .NET application into an MCP server. | Tool | ||
| ReflectorNet | Open-source .NET reflection toolkit for AI-driven scenarios. | Tool |
VLM (Visual)
| Source | Description | Paper | Game Engine | Type |
|---|---|---|---|---|
| Cambrian-1 | Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs. | arXiv | Multimodal LLMs | |
| CogVLM2 | GPT4V-level open-source multi-modal model based on Llama3-8B. | Visual | ||
| CoTracker | It is Better to Track Together. | arXiv | Visual | |
| dots.vlm1 | dots.vlm1 is the first vision-language model in the dots model family. Built upon a 1.2 billion-parameter vision encoder and the DeepSeek V3 large language model (LLM), dots.vlm1 demonstrates strong multimodal understanding and reasoning capabilities. | VLM | ||
| EVF-SAM | EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model. | arXiv | Visual | |
| FaceHi | It is Better to Track Together. | Visual | ||
| GLM-V | GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning. | arXiv | VLM | |
| InternLM-XComposer2 | InternLM-XComposer2 is a groundbreaking vision-language large model (VLLM) excelling in free-form text-image composition and comprehension. | arXiv | Visual | |
| Kangaroo | Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input. | Visual | ||
| Kwai Keye-VL | Kwai Keye-VL is a cutting-edge multimodal large language model meticulously crafted by the Kwai Keye Team at Kuaishou. | arXiv | VLM | |
| LGVI | Towards Language-Driven Video Inpainting via Multimodal Large Language Models. | Visual | ||
| LLaVA++ | Extending Visual Capabilities with LLaMA-3 and Phi-3. | Visual | ||
| LLaVA-OneVision | LLaVA-OneVision: Easy Visual Task Transfer. | arXiv | Visual | |
| LongVA | Long Context Transfer from Language to Vision. | arXiv | Visual | |
| Lumina-DiMOO | Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding. | VLM | ||
| MaskViT | Masked Visual Pre-Training for Video Prediction. | arXiv | Visual | |
| MiniCPM-Llama3-V 2.5 | A GPT-4V Level MLLM on Your Phone. | Visual | ||
| MiniCPM-V 4.0 | MiniCPM-V 4.0: A GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone. | Visual | ||
| MoE-LLaVA | Mixture of Experts for Large Vision-Language Models. | arXiv | Visual | |
| MotionLLM | Understanding Human Behaviors from Human Motions and Videos. | arXiv | Visual | |
| PLLaVA | Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning. | arXiv | Visual | |
| POINTS-Reader | POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion. | arXiv | Visual | |
| Qwen-VL | A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. | arXiv | Visual | |
| Sapiens | Sapiens: Foundation for Human Vision Models. | arXiv | Visual | |
| ShareGPT4V | Improving Large Multi-modal Models with Better Captions. | arXiv | Visual | |
| SOLO | SOLO: A Single Transformer for Scalable Vision-Language Modeling. | arXiv | Visual | |
| VideoAgent | VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding. | arXiv | Agent | |
| Video-CCAM | Video-CCAM: Advancing Video-Language Understanding with Causal Cross-Attention Masks. | Visual | ||
| Video-LLaVA | Learning United Visual Representation by Alignment Before Projection. | arXiv | Visual | |
| VideoLLaMA 2 | Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs. | arXiv | Visual | |
| VideoLLaMA 3 | VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding. | arXiv | Visual | |
| Video-MME | The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. | arXiv | Visual | |
| Vitron | A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing. | Visual | ||
| VILA | VILA: On Pre-training for Visual Language Models. | arXiv | Visual |
Game (World Model & Agent)
| Source | Description | Paper | Game Engine | Type |
|---|---|---|---|---|
| AgentBench | A Comprehensive Benchmark to Evaluate LLMs as Agents. | arXiv | Agent | |
| Agent Group Chat | An Interactive Group Chat Simulacra For Better Eliciting Collective Emergent Behavior. | arXiv | Agent | |
| Agent K | An autoagentic AGI that is self-evolving and modular. | Agent | ||
| Agent Laboratory | Agent Laboratory: Using LLM Agents as Research Assistants. | arXiv | Agent | |
| AgentScope | Start building LLM-empowered multi-agent applications in an easier way. | arXiv | Agent | |
| AgentSims | An Open-Source Sandbox for Large Language Model Evaluation. | Agent | ||
| AI Town | AI Town is a virtual town where AI characters live, chat and socialize. | Agent | ||
| anime.gf | Local & Open Source Alternative to CharacterAI. | Game | ||
| Astrocade | Create games with AI | Game | ||
| Atomic Agents | The Atomic Agents framework is designed to be modular, extensible, and easy to use. | Agent | ||
| AutoAgents | A Framework for Automatic Agent Generation. | Agent | ||
| AutoGen | Enable Next-Gen Large Language Model Applications. | arXiv | Agent | |
| AWorld | AWorld: The Agent Runtime for Self-Improvement. | Agent | ||
| behaviac | Behaviac is a framework of the game AI development. | Framework | ||
| Biomes | Biomes is an open source sandbox MMORPG built for the web using web technologies such as Next.js, Typescript, React and WebAssembly. | Game | ||
| Buffer of Thoughts | Thought-Augmented Reasoning with Large Language Models. | arXiv | Agent | |
| Byzer-Agent | Easy, fast, and distributed agent framework for everyone. | Agent | ||
| Cat Town | A C(h)atGPT-powered simulation with cats. | Agent | ||
| Cat Town | A C(h)atGPT-powered simulation with cats. | Agent | ||
| CharacterGLM | Customizing Chinese Conversational AI Characters with Large Language Models. | arXiv | Agent | |
| ChatDev | Communicative Agents for Software Development. | arXiv | Agent | |
| CogAgent | CogAgent is an open-source visual language model improved based on CogVLM. | arXiv | Agent | |
| ComoRAG | ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative Reasoning. | arXiv | Agent | |
| Cradle | Towards General Computer Control. | Agent | ||
| crewAI | Framework for orchestrating role-playing, autonomous AI agents. | Agent | ||
| Datarus Jupyter Agent | The Datarus Jupyter Agent is a powerful multi-step reasoning system that executes complex analytical workflows with step-by-step reasoning, automatic error recovery, and comprehensive result synthesis. | Agent | ||
| Dify | Dify is an open-source LLM app building platform. | Agent | ||
| Digital Life Project | Autonomous 3D Characters with Social Intelligence. | arXiv | Agent | |
| everything-ai | Your fully proficient, AI-powered and local chatbot assistant🤖. | Agent | ||
| fabric | fabric is an open-source framework for augmenting humans using AI. | Agent | ||
| FastGPT | FastGPT is a knowledge-based platform built on the LLM. | Agent | ||
| fastRAG | Efficient Retrieval Augmentation and Generation Framework. | Agent | ||
| GameAISDK | Image-based game AI automation framework. | Framework | ||
| GameNGen | Diffusion Models Are Real-Time Game Engines. | arXiv | Game | |
| GameGen-O | GameGen-O: Open-world Video Game Generation. | Game | ||
| GenAgent | GenAgent: Build Collaborative AI Systems with Automated Workflow Generation - Case Studies on ComfyUI. | arXiv | Agent | |
| Generative Agents | Interactive Simulacra of Human Behavior. | arXiv | Agent | |
| Genesis | Genesis: A Generative and Universal Physics Engine for Robotics and Beyond. | Game | ||
| Genie | Generative Interactive Environments. | Game | ||
| Genie 3 | Genie 3: A new frontier for world models. Genie 3 is a general purpose world model that can generate an unprecedented diversity of interactive environments. | Game | ||
| gigax | Runtime, LLM-powered NPCs. | Game | ||
| HippoRAG | Neurobiologically Inspired Long-Term Memory for Large Language Models. | arXiv | Agent | |
| Hunyuan-GameCraft | Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition. | arXiv | Game | |
| HunyuanWorld 1.0 | HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels. | arXiv | Game | |
| HunyuanWorld-Voyager | HunyuanWorld-Voyager is a novel video diffusion framework that generates world-consistent 3D point-cloud sequences from a single image with user-defined camera path. Voyager can generate 3D-consistent scene videos for world exploration following custom camera trajectories. | Game | ||
| HY-World 1.5 | HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric Consistency. | Game | ||
| Interactive LLM Powered NPCs | Interactive LLM Powered NPCs, is an open-source project that completely transforms your interaction with non-player characters (NPCs) in any game! | Game | ||
| IoA | An open-source framework for collaborative AI agents, enabling diverse, distributed agents to team up and tackle complex tasks through internet-like connectivity. | Agent | ||
| Jaaz | Jaaz - The world's first open-source multimodal creative assistant. AI design agent, local alternative for Lovart. Canva + Cursor. AI agent with ability to design, edit and generate images, posters, storyboards, etc. | Agent | ||
| KwaiAgents | A generalized information-seeking agent system with Large Language Models (LLMs). | arXiv | Agent | |
| LangChain | Get your LLM application from prototype to production. | Agent | ||
| Langflow | Langflow is a UI for LangChain, designed with react-flow to provide an effortless way to experiment and prototype flows. | Agent | ||
| LangGraph Studio | LangGraph Studio offers a new way to develop LLM applications by providing a specialized agent IDE that enables visualization, interaction, and debugging of complex agentic applications. | Agent | ||
| LARP | Language-Agent Role Play for open-world games. | arXiv | Agent | |
| LLama Agentic System | Agentic components of the Llama Stack APIs. | Agent | ||
| LlamaIndex | LlamaIndex is a data framework for your LLM application. | Agent | ||
| Matrix-Game | Matrix-Game: Interactive World Foundation Model. Matrix-Game is a 17B-parameter interactive world foundation model for controllable game world generation. | Game | ||
| Matrix-Game 2.0 | Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World Model. | Game | ||
| MindSearch | 🔍 An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT). | Agent | ||
| Mixture of Agents (MoA) | Mixture-of-Agents Enhances Large Language Model Capabilities. | arXiv | Agent | |
| MMRole | MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents. | arXiv | Agent | |
| Moonlander.ai | Start building 3D games without any coding using generative AI. | Framework | ||
| MuG Diffusion | MuG Diffusion is a charting AI for rhythm games based on Stable Diffusion (one of the most powerful AIGC models) with a large modification to incorporate audio waves. | Game | ||
| NVIDIA NeMo Agent Toolkit | NVIDIA NeMo Agent toolkit is a flexible, lightweight, and unifying library that allows you to easily connect existing enterprise agents to data sources and tools across any framework. | Agent | ||
| Oasis | Oasis is an interactive world model developed by Decart and Etched. Based on diffusion transformers, Oasis takes in user keyboard input and generates gameplay in an autoregressive manner. | Game | ||
| OmAgent | A multimodal agent framework for solving complex tasks. | Agent | ||
| OpenAgents | An Open Platform for Language Agents in the Wild. | Agent | ||
| OpenGame | OpenGame: Open Agentic Coding for Games. | arXiv | Game | |
| Opus | An AI app that turns text into a video game. | Game | ||
| Pipecat | Open Source framework for voice and multimodal conversational AI. | Agent | ||
| Qwen-Agent | Qwen-Agent is a framework for developing LLM applications based on the instruction following, tool usage, planning, and memory capabilities of Qwen. | Agent | ||
| Ragas | Ragas is a framework that helps you evaluate your Retrieval Augmented Generation (RAG) pipelines. | Agent | ||
| RPBench-Auto | An automated pipeline for evaluating LLMs for role-playing. | Game | ||
| Rosebud AI | Vibe coding platform for creating 3D games and interactive web apps with AI. | Game | ||
| SIMA | A generalist AI agent for 3D virtual environments. | Agent | ||
| StoryGames.ai | AI for Dreamers Make Games. | Game | ||
| SWE-agent | Agent Computer Interfaces Enable Software Engineering Language Models. | arXiv | Agent | |
| TaskGen | A Task-based agentic framework building on StrictJSON outputs by LLM agents. | Agent | ||
| TEN Agent | TEN Agent is the world’s first real-time multimodal agent integrated with the OpenAI Realtime API, RTC, and features weather checks, web search, vision, and RAG capabilities. | Agent | ||
| Translation Agent | Agentic translation using reflection workflow. | Agent | ||
| Twitter Personality is a web application that analyzes your Twitter handle to create a personalized personality profile using Wordware AI Agent. | Agent | |||
| Unbounded | Unbounded: A Generative Infinite Game of Character Life Simulation. | arXiv | Game | |
| Video2Game | Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video. | arXiv | Game | |
| V-IRL | Grounding Virtual Intelligence in Real Life. | arXiv | Agent | |
| WebDesignAgent | An agent used for webdesign. | Agent | ||
| XAgent | An Autonomous LLM Agent for Complex Task Solving. | Agent |
Code
| Source | Description | Paper | Game Engine | Type |
|---|---|---|---|---|
| AI Code Translator | Use AI to translate code from one language to another. | Code | ||
| aiXcoder-7B | aiXcoder-7B Code Large Language Model. | Code | ||
| bloop | bloop is a fast code search engine written in Rust. | Code | ||
| Chapyter | ChatGPT Code Interpreter in Jupyter Notebooks. | Code | ||
| CodeGeeX | An Open Multilingual Code Generation Model. | arXiv | Code | |
| CodeGeeX2 | A More Powerful Multilingual Code Generation Model. | Code | ||
| CodeGeeX4 | CodeGeeX4: Open Multilingual Code Generation Model. | Code | ||
| CodeGen | CodeGen is an open-source model for program synthesis. Trained on TPU-v4. Competitive with OpenAI Codex. | arXiv | Code | |
| CodeGen2 | CodeGen2 models for program synthesis. | arXiv | Code | |
| Code Llama | Code Llama is a large language models for code based on Llama 2. | Code | ||
| CodeTF | One-stop Transformer Library for State-of-the-art Code LLM. | Code | ||
| CodeT5 | Open Code LLMs for Code Understanding and Generation. | Code | ||
| Code World Model (CWM) | Code World Model (CWM) is a 32-billion-parameter open-weights LLM, to advance research on code generation with world models. | Code | ||
| Cursor | Write, edit, and chat about your code with GPT-4 in a new type of editor. | Code | ||
| DeepSeek Coder | DeepSeek Coder: Let the Code Write Itself. | arXiv | Code | |
| OpenAI Codex | OpenAI Codex is a descendant of GPT-3. | Code | ||
| PandasAI | Pandas AI is a Python library that integrates generative artificial intelligence capabilities into Pandas, making dataframes conversational. | Code | ||
| RobloxScripterAI | RobloxScripterAI is an AI-powered code generation tool for Roblox. | Roblox | Code | |
| Roblox GUI Maker | Roblox GUI Maker generates Roblox Studio GUI layouts and Lua starter code from prompts for faster game UI prototyping. | Roblox | Code | |
| Scikit-LLM | Seamlessly integrate powerful language models like ChatGPT into scikit-learn for enhanced text analysis tasks. | Code | ||
| SoTaNa | The Open-Source Software Development Assistant. | arXiv | Code | |
| Stable Code 3B | Coding on the Edge. | Code | ||
| StarCoder | 💫 StarCoder is a language model (LM) trained on source code and natural language text. | arXiv | Code | |
| StarCoder 2 | StarCoder2 is a family of code generation models (3B, 7B, and 15B), trained on 600+ programming languages from The Stack v2 and some natural language text such as Wikipedia, Arxiv, and GitHub issues. | arXiv | Code | |
| Tura | A terminal-native coding agent that turns intent into verified code changes with repo-aware controls and auditable execution. | Code | ||
| UnityGen AI | UnityGen AI is an AI-powered code generation plugin for Unity. | Unity | Code | |
| Void | Void is an open source Cursor alternative. Write code with the best AI tools, retain full control over your data, and access powerful AI features. | Code |
Image
| Source | Description | Paper | Game Engine | Type |
|---|---|---|---|---|
| AnyDoor | Zero-shot Object-level Image Customization. | arXiv | Image | |
| AnyText | Multilingual Visual Text Generation And Editing. | arXiv | Image | |
| AutoStudio | Crafting Consistent Subjects in Multi-turn Interactive Image Generation. | arXiv | Image | |
| BAGEL | BAGEL - Unified Model for Multimodal Understanding and Generation. BAGEL is an open‑source multimodal foundation model with 7B active parameters (14B total) trained on large‑scale interleaved multimodal data. | arXiv | Image | |
| Blender-ControlNet | Using ControlNet right in Blender. | Blender | Image | |
| BriVL | Bridging Vision and Language Model. | arXiv | Image | |
| CatVTON | CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models. | arXiv | Image | |
| CLIPasso | A method for converting an image of an object to a sketch, allowing for varying levels of abstraction. | arXiv | Image | |
| ClipDrop | Create stunning visuals in seconds. | Image | ||
| ComfyUI | A powerful and modular stable diffusion GUI with a graph/nodes interface. | Image | ||
| ConceptLab | Creative Generation using Diffusion Prior Constraints. | arXiv | Image | |
| ControlNet | ControlNet is a neural network structure to control diffusion models by adding extra conditions. | arXiv | Image | |
| CSGO | CSGO: Content-Style Composition in Text-to-Image Generation. | arXiv | Image | |
| DALL·E 2 | DALL·E 2 is an AI system that can create realistic images and art from a description in natural language. | Image | ||
| Dashtoon Studio | Dashtoon Studio is an AI powered comic creation platform. | Comic | ||
| DeepAI | DeepAI offers a suite of tools that use AI to enhance your creativity. | Image | ||
| DeepFloyd IF | IF by DeepFloyd Lab at StabilityAI. | Image | ||
| Depth Anything V2 | Depth Anything V2 | arXiv | Image | |
| Depth map library and poser | Depth map library for use with the Control Net extension for Automatic1111/stable-diffusion-webui. | Image | ||
| Diffuse to Choose | Enriching Image Conditioned Inpainting in Latent Diffusion Models for Virtual Try-All. | arXiv | Image | |
| Disco Diffusion | A frankensteinian amalgamation of notebooks, models and techniques for the generation of AI Art and Animations. | Image | ||
| DragGAN | Interactive Point-based Manipulation on the Generative Image Manifold. | arXiv | Image | |
| Draw Things | AI- assisted image generation in Your Pocket. | Image | ||
| DWPose | Effective Whole-body Pose Estimation with Two-stages Distillation. | arXiv | Image | |
| EasyPhoto | Your Smart AI Photo Generator. | Image | ||
| Flux | This repo contains minimal inference code to run text-to-image and image-to-image with our Flux latent rectified flow transformers. | Image | ||
| Follow-Your-Click | Open-domain Regional Image Animation via Short Prompts. | arXiv | Image | |
| Fooocus | Focus on prompting and generating. | Image | ||
| GIFfusion | Create GIFs and Videos using Stable Diffusion. | Image | ||
| Grounded-Segment-Anything | Automatically Detect , Segment and Generate Anything with Image, Text, and Audio Inputs. | arXiv | Image | |
| HivisionIDPhotos | HivisionIDPhotos: a lightweight and efficient AI ID photos tools. | Image | ||
| Hua | Hua is an AI image editor with Stable Diffusion (and more). | Image | ||
| Hunyuan-DiT | A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding. | arXiv | Image | |
| HunyuanImage-2.1 | HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation. | Image | ||
| HunyuanImage-3.0 | HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation. | Image | ||
| IC-Light | IC-Light is a project to manipulate the illumination of images. | Image | ||
| Ideogram | Helping people become more creative. | Image | ||
| Imagen | Imagen is an AI system that creates photorealistic images from input text. | Image | ||
| img2img-turbo | One-Step Image-to-Image with SD-Turbo. | Image | ||
| Img2Prompt | Get prompts from stable diffusion generated images. | Image | ||
| Infinity | Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis. | arXiv | Image | |
| InstantID | Zero-shot Identity-Preserving Generation in Seconds. | arXiv | Image | |
| InternLM-XComposer2 | InternLM-XComposer2 is a groundbreaking vision-language large model (VLLM) excelling in free-form text-image composition and comprehension. | arXiv | Image | |
| IRG | IRG - Interleaving Reasoning for Better Text-to-Image Generation. | arXiv | Image | |
| KOALA | Self-Attention Matters in Knowledge Distillation of Latent Diffusion Models for Memory-Efficient and Fast Image Synthesis. | Image | ||
| Kolors | Kolors: Effective Training of Diffusion Model for Photorealistic Text-to-Image Synthesis. | Image | ||
| Komiko | Komiko is an AI-powered storytelling platform that lets you create original characters, comics, and animations with ease. | Comic | ||
| KREA | Generate images and videos with a delightful AI-powered design tool. | Image | ||
| LaVi-Bridge | Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation. | arXiv | Image | |
| LayerDiffusion | Transparent Image Layer Diffusion using Latent Transparency. | arXiv | Image | |
| Lexica | A Stable Diffusion prompts search engine. | Image | ||
| LlamaGen | Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation. | arXiv | Image | |
| Lumina-Image 2.0 | Lumina-Image 2.0 : A Unified and Efficient Image Generative Model. | Image | ||
| Lumina-mGPT | Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining. | arXiv | Image | |
| MakeAnything | MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation. | arXiv | Image | |
| MetaShoot | MetaShoot is a digital twin of a photo studio, developed as a plugin for Unreal Engine that gives any creator the ability to produce highly realistic renders in the easiest and quickest way. | Unreal Engine | Image | |
| Midjourney | Midjourney is an independent research lab exploring new mediums of thought and expanding the imaginative powers of the human species. | Image | ||
| MIGC | MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis. | arXiv | Image | |
| MimicBrush | Zero-shot Image Editing with Reference Imitation. | arXiv | Image | |
| NextStep-1 | NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale. | arXiv | Image | |
| OmniGen | OmniGen: Unified Image Generation. | arXiv | Image | |
| OmniGen2 | OmniGen2: Exploration to Advanced Multimodal Generation. | arXiv | Image | |
| Oniichan | AI sprite generator and game character creator. Generate game-ready character sprites and original characters from text prompts using a custom finetuned model, with editing, inpainting, and a reusable character library. | Comic | ||
| Omost | Omost is a project to convert LLM's coding capability to image generation (or more accurately, image composing) capability. | Image | ||
| Openpose Editor | Openpose Editor for AUTOMATIC1111's stable-diffusion-webui. | Image | ||
| Outfit Anyone | Ultra-high quality virtual try-on for Any Clothing and Any Person. | Image | ||
| PaintsUndo | PaintsUndo: A Base Model of Drawing Behaviors in Digital Paintings. | Image | ||
| PhotoMaker | Customizing Realistic Human Photos via Stacked ID Embedding. | arXiv | Image | |
| Photoroom | AI Background Generator. | Image | ||
| Plask | AI image generation in the cloud. | Image | ||
| PosterCraft |