← Open Source
BradyFU

Awesome-Multimodal-Large-Language-Models

:sparkles::sparkles:Latest Advances on Multimodal Large Language Models

ListsPaper collections
Open on GitHub
Momentum
+-2stars in 24 hours-0.0%
18.0k
Stars
1.14k
Forks
+2
This week
12
Contributors
Created 2023-05-19 · Updated 2026-10-05 · #18233 today
Top developers
README

Awesome-Multimodal-Large-Language-Models

 ![](./images/mig_logo.png) 

✨ Highlights of NJU-MiG

🔥🔥 Surveys of MLLMs | 💬 WeChat (MLLM微信交流群)

  • 🌟 MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
    arXiv 2025, Paper, Project

  • 🌟 A Survey of Unified Multimodal Understanding and Generation: Advances and Challenges
    arXiv 2025, Paper, Project

  • A Survey on Multimodal Large Language Models
    NSR 2024, Paper, Project


🔥🔥 VITA Series Omni MLLMs | 💬 WeChat (VITA微信交流群)

  • VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
    NeurIPS 2025 Highlight, Paper, Project

  • VITA: Towards Open-Source Interactive Omni Multimodal LLM
    arXiv 2024, Paper, Project

  • VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
    NeurIPS 2025, Paper, Project


🔥🔥 MME Series MLLM Benchmarks

  • 🔥 Video-MME-v2: Towards the Next Stage in Video Understanding Evaluation

[🍎 Project Page] [📖 Paper] [🤗 Dataset] [🏆 Leaderboard]

  • 🌟 MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
    arXiv 2025, Paper, Project

  • MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
    NeurIPS 2025 DB Highlight, Paper, Dataset, Eval Tool, ✒️ Citation

  • Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
    CVPR 2025, Paper, Project, Dataset


Table of Contents


Awesome Papers

Multimodal Instruction Tuning (& Latest Works)

Title Venue Date Code Demo
Gemini 4 Argon: our next era of frontier intelligence Blog 2026-09-30 - -
Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery. Blog 2026-09-18 - -
GPT-6 Astra: A new generation of intelligence Blog 2026-09-04 - -
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber Blog 2026-09-02 - -
GLM-5.3-Flash: Frontier Intelligence, Flash Cost Zhipu 2026-08-26 Huggingface -
DeepSeek-V4-Flash-Vision-Exp DeepSeek 2026-08-21 - -
Qwen3.8-Max: A New Bar for Coding and Cowork Blog 2026-08-15 Huggingface Demo
Star
DeepSeek Harness
DeepSeek 2026-08-13 Github -
Star
Qwen-MM-Plugins
Qwen 2026-08-11 Github -
Star
Kimi K3: Open Frontier Intelligence
Kimi 2026-07-27 Github Demo
Star
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
arXiv 2026-07-16 Github Demo
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity arXiv 2026-06-30 - -
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models Blog 2026-06-15 - -
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation arXiv 2026-06-15 - Demo
Star
OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains
arXiv 2026-06-12 Github Demo
GitHub Repo stars
Cosmos 3: Omnimodal World Models for Physical AI
arXiv 2026-06-05 Github Demo
π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities arXiv 2026-04-24 - Demo
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence DeepSeek 2026-04-24 Huggingface -
Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model Blog 2026-04-22 Huggingface Demo
Xiaomi MiMo-V2.5 Blog 2026-04-22 - Demo
Star
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
arXiv 2026-04-06 Github Demo
Introducing Muse Spark: Scaling Towards Personal Superintelligence Blog 2026-04-08 - Demo
Star
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing
arXiv 2026-04-03 Github Local Demo
Gemma 4: Byte for byte, the most capable open models Blog 2026-04-02 - Demo
Qwen3.6-Plus: Towards Real World Agents Blog 2026-04-02 - -
Qwen3.5-Omni: Scaling Up, Toward Native Omni-Modal AGI Blog 2026-03-30 - Demo
Xiaomi MiMo-V2-Omni Blog 2026-03-18 - -
Star
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
arXiv 2026-03-10 Github Local Demo
Star
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
arXiv 2026-03-06 Github -
Beyond Language Modeling: An Exploration of Multimodal Pretraining arXiv 2026-03-03 - -
Gemini 3.1 Pro: A smarter model for your most complex tasks Blog 2026-02-19 - -
Star
Qwen3.5: Towards Native Multimodal Agents
Blog 2026-02-16 Github Demo
Star
MiniCPM-o 4.5
Blog 2026-02-06 Github Demo
Star
Kimi K2.5: Visual Agentic Intelligence
arXiv 2026-02-02 Github -
Star
DeepSeek-OCR 2: Visual Causal Flow
DeepSeek 2026-01-27 Github -
Seed1.8 Model Card: Towards Generalized Real-World Agency Bytedance Seed 2025-12-18 - -
Introducing GPT-5.2 OpenAI 2025-12-11 - -
Introducing Mistral 3 Blog 2025-12-02 Huggingface -
Star
Qwen3-VL Technical Report
arXiv 2025-11-26 Github Demo
Star
Emu3.5: Native Multimodal Models are World Learners
arXiv 2025-10-30 Github -
Star
VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
arXiv 2025-10-21 Github Local Demo
Star
DeepSeek-OCR: Contexts Optical Compression
arXiv 2025-10-21 Github -
Star
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
arXiv 2025-10-17 Github -
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching arXiv 2025-10-16 - -
Star
InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue arXiv 2025-10-15 Github -
Star
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
arXiv 2025-10-10 Github -
Star
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
arXiv 2025-10-09 Github Demo
Star
Qwen3-Omni Technical Report
arXiv 2025-09-22 Github Demo
Star
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
arXiv 2025-08-27 Github Demo
MiniCPM-V 4.5: A GPT-4o Level MLLM for Single Image, Multi Image and Video Understanding on Your Phone - 2025-08-26 Github Demo
Star
Thyme: Think Beyond Images
arXiv 2025-08-18 Github Demo
Introducing GPT-5 OpenAI 2025-08-07 - -
Star
dots.vlm1
rednote-hilab 2025-08-06 Github Demo
Star
Step3: Cost-Effective Multimodal Intelligence
StepFun 2025-07-31 Github Demo
Star
GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
arXiv 2025-07-02 Github Demo
Star
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
arXiv 2025-06-30 Github -
Qwen VLo: From "Understanding" the World to "Depicting" It Qwen 2025-06-26 - Demo
Star
MMSearch-R1: Incentivizing LMMs to Search
arXiv 2025-06-25 Github -
Star
Show-o2: Improved Native Unified Multimodal Models
arXiv 2025-06-18 Github -
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Google 2025-06-17 - -
Star
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
arXiv 2025-06-16 Github -
Star
MiMo-VL Technical Report
arXiv 2025-06-04 Github -
Star
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
arXiv 2025-05-29 Github -
Star
Emerging Properties in Unified Multimodal Pretraining
arXiv 2025-05-23 Github Demo
Star
MMaDA: Multimodal Large Diffusion Language Models
arXiv 2025-05-21 Github Demo
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation arXiv 2025-05-20 - -
Star
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
arXiv 2025-05-14 Github Local Demo
Seed1.5-VL Technical Report arXiv 2025-05-11 - -
Star
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
arXiv 2025-05-08 Github -
Star
VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
arXiv 2025-05-06 Github Local Demo
Star
Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
arXiv 2025-04-23 Github -
Star
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
arXiv 2025-04-21 Github -
Star
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
arXiv 2025-04-21 Github -
Star
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
arXiv 2025-04-14 Github Demo
Introducing GPT-4.1 in the API OpenAI 2025-04-14 - -
Star
Kimi-VL Technical Report
arXiv 2025-04-10 Github Demo
The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation Meta 2025-04-05 Hugging Face -
Star
Qwen2.5-Omni Technical Report
Qwen 2025-03-26 Github Demo
Addendum to GPT-4o System Card: Native image generation OpenAI 2025-03-25 - -
Star
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
arXiv 2025-03-17 Github -
Nexus-O: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision arXiv 2025-03-07 - -
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs arXiv 2025-03-03 Hugging Face Demo
Star
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuray
arXiv 2025-02-19 Github -
Star
Qwen2.5-VL Technical Report
arXiv 2025-02-19 Github Demo
Star
Baichuan-Omni-1.5 Technical Report
Tech Report 2025-01-26 Github Local Demo
Star
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
arXiv 2025-01-10 Github -
Star
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
arXiv 2025-01-03 Github -
Star
QVQ: To See the World with Wisdom
Qwen 2024-12-25 Github Demo
Star
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
arXiv 2024-12-13 Github -
Apollo: An Exploration of Video Understanding in Large Multimodal Models arXiv 2024-12-13 - -
Star
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
arXiv 2024-12-12 Github Local Demo
StreamChat: Chatting with Streaming Video arXiv 2024-12-11 Coming soon -
CompCap: Improving Multimodal Large Language Models with Composite Captions arXiv 2024-12-06 - -
Star
LinVT: Empower Your Image-level Large Language Model to Understand Videos
arXiv 2024-12-06 Github -
Star
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
arXiv 2024-12-06 Github Demo
Star
NVILA: Efficient Frontier Visual Language Models
arXiv 2024-12-05 Github Demo
Star
Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning
arXiv 2024-12-04 Github -
Star
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
arXiv 2024-11-27 Github -
Star
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
arXiv 2024-11-27 Github Local Demo
Star
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
arXiv 2024-10-22 Github Demo
Star
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
arXiv 2024-10-09 Github -
Star
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
arXiv 2024-10-04 Github Local Demo
Star
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
CVPR 2024-09-26 Github Demo
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models arXiv 2024-09-25 Huggingface Demo
Star
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
arXiv 2024-09-18 Github Demo
Star
ChartMoE: Mixture of Expert Connector for Advanced Chart Understanding
ICLR 2024-09-05 Github Local Demo
Star
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via Hybrid Architecture
arXiv 2024-09-04 Github -
Star
EAGLE: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
arXiv 2024-08-28 Github Demo
Star
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
arXiv 2024-08-28 Github -
Star
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
arXiv 2024-08-09 Github -
Star
VITA: Towards Open-Source Interactive Omni Multimodal LLM
arXiv 2024-08-09 Github -
Star
LLaVA-OneVision: Easy Visual Task Transfer
arXiv 2024-08-06 Github Demo
Star
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
arXiv 2024-08-03 Github Demo
VILA^2: VILA Augmented VILA arXiv 2024-07-24 - -
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models arXiv 2024-07-22 - -
EVLM: An Efficient Vision-Language Model for Visual Understanding arXiv 2024-07-19 - -
Star
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
arXiv 2024-07-10 Github -
Star
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
arXiv 2024-07-03 Github Demo
Star
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
arXiv 2024-06-27 Github Local Demo
Star
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
AAAI 2024-06-27 Github -
Star
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
arXiv 2024-06-24 Github Local Demo
Star
Long Context Transfer from Language to Vision
arXiv 2024-06-24 Github Local Demo
Star
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
ICML 2024-06-22 Github -
Star
TroL: Traversal of Layers for Large Language and Vision Models
EMNLP 2024-06-18 Github Local Demo
Star
Unveiling Encoder-Free Vision-Language Models
arXiv 2024-06-17 Github Local Demo
Star
VideoLLM-online: Online Video Large Language Model for Streaming Video
CVPR 2024-06-17 Github Local Demo
Star
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
CoRL 2024-06-15 Github Demo
Star
Comparison Visual Instruction Tuning
arXiv 2024-06-13 Github Local Demo
Star
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
arXiv 2024-06-12 Github -
Star
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
arXiv 2024-06-11 Github Local Demo
Star
Parrot: Multilingual Visual Instruction Tuning
arXiv 2024-06-04 Github -
Star
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
arXiv 2024-05-31 Github -
Star
Matryoshka Query Transformer for Large Vision-Language Models
arXiv 2024-05-29 Github Demo
Star
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models
arXiv 2024-05-24 Github -
Star
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
arXiv 2024-05-24 Github Demo
Star
Libra: Building Decoupled Vision System on Large Language Models
ICML 2024-05-16 Github Local Demo
Star
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
arXiv 2024-05-09 Github Local Demo
Star
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
arXiv 2024-04-25 Github Demo
Star
Graphic Design with Large Multimodal Model
arXiv 2024-04-22 Github -
BRAVE: Broadening the visual encoding of vision-language models ECCV 2024-04-10 - -
Star
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
arXiv 2024-04-09 Github Demo
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs arXiv 2024-04-08 - -
Star
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
CVPR 2024-04-08 Github -
Star
VITRON: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
NeurIPS 2024-04-04 Github Local Demo
TOMGPT: Reliable Text-Only Training Approach for Cost-Effective Multi-modal Large Language Model ACM TKDD 2024-03-28 - -
Star
LITA: Language Instructed Temporal-Localization Assistant arXiv 2024-03-27 Github Local Demo
Star
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
arXiv 2024-03-27 Github Demo
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training arXiv 2024-03-14 - -
Star
MoAI: Mixture of All Intelligence for Large Language and Vision Models
arXiv 2024-03-12 Github Local Demo
Star
DeepSeek-VL: Towards Real-World Vision-Language Understanding
arXiv 2024-03-08 Github Demo
Star
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
arXiv 2024-03-07 Github Demo
Star
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World arXiv 2024-02-29 Github -
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation CVPR 2024-02-26 Coming soon Coming soon
Star
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
arXiv 2024-02-19 Github -
Star
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
arXiv 2024-02-18 Github -
Star
ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language Model
arXiv 2024-02-18 Github Demo
Star
CoLLaVO: Crayon Large Language and Vision mOdel
arXiv 2024-02-17 Github -
Star
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
ICML 2024-02-12 Github -
Star
CogCoM: Train Large Vision-Language Models Diving into Details through Chain of Manipulations
arXiv 2024-02-06 Github -
Star
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
arXiv 2024-02-06 Github -
Star
GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning
NeurIPS 2024-02-03 Github -
Enhancing Multimodal Large Language Models with Vision Detection Models: An Empirical Study arXiv 2024-01-31 Coming soon -
Star
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge Blog 2024-01-30 Github Demo
Star
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
arXiv 2024-01-29 Github Demo
Star
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
arXiv 2024-01-29 Github Demo
Star
Yi-VL
- 2024-01-23 Github Local Demo
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities arXiv 2024-01-22 - -
Star
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
ACL 2024-01-04 Github Local Demo
Star
MobileVLM : A Fast, Reproducible and Strong Vision Language Assistant for Mobile Devices
arXiv 2023-12-28 Github -
Star
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
CVPR 2023-12-21 Github Demo
Star
Osprey: Pixel Understanding with Visual Instruction Tuning
CVPR 2023-12-15 Github Demo
Star
CogAgent: A Visual Language Model for GUI Agents
arXiv 2023-12-14 Github Coming soon
Pixel Aligned Language Models arXiv 2023-12-14 Coming soon -
Star
VILA: On Pre-training for Visual Language Models
CVPR 2023-12-13 Github Local Demo
See, Say, and Segment: Teaching LMMs to Overcome False Premises arXiv 2023-12-13 Coming soon -
Star
Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
ECCV 2023-12-11 Github Demo
Star
Honeybee: Locality-enhanced Projector for Multimodal LLM
CVPR 2023-12-11 Github -
Gemini: A Family of Highly Capable Multimodal Models Google 2023-12-06 - -
Star
OneLLM: One Framework to Align All Modalities with Language
arXiv 2023-12-06 Github Demo
Star
Lenna: Language Enhanced Reasoning Detection Assistant
arXiv 2023-12-05 Github -
VaQuitA: Enhancing Alignment in LLM-Assisted Video Understanding arXiv 2023-12-04 - -
Star
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
arXiv 2023-12-04 Github Local Demo
Star
Making Large Multimodal Models Understand Arbitrary Visual Prompts
CVPR 2023-12-01 Github Demo
Star
Dolphins: Multimodal Language Model for Driving
arXiv 2023-12-01 Github -
Star
LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning
arXiv 2023-11-30 Github Coming soon
Star
VTimeLLM: Empower LLM to Grasp Video Moments
arXiv 2023-11-30 Github Local Demo
Star
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
arXiv 2023-11-30 Github -
Star
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
arXiv 2023-11-28 Github Coming soon
Star
LLMGA: Multimodal Large Language Model based Generation Assistant
arXiv 2023-11-27 Github Demo
Star
ChartLlama: A Multimodal LLM for Chart Understanding and Generation
arXiv 2023-11-27 Github -
Star
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
arXiv 2023-11-21 Github Demo
Star
LION : Empowering Multimodal Large Language Model with Dual-Level Visual Knowledge
arXiv 2023-11-20 Github -
Star
An Embodied Generalist Agent in 3D World
arXiv 2023-11-18 Github Demo
Star
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
arXiv 2023-11-16 Github Demo
Star
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
CVPR 2023-11-14 Github -
Star
To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
arXiv 2023-11-13 Github -
Star
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
arXiv 2023-11-13 Github Demo
Star
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
CVPR 2023-11-11 Github Demo
Star
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
arXiv 2023-11-09 Github Demo
Star
NExT-Chat: An LMM for Chat, Detection and Segmentation
arXiv 2023-11-08 Github Local Demo
Star
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
arXiv 2023-11-07 Github Demo
Star
OtterHD: A High-Resolution Multi-modality Model
arXiv 2023-11-07 Github -
CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding arXiv 2023-11-06 Coming soon -
Star
GLaMM: Pixel Grounding Large Multimodal Model
CVPR 2023-11-06 Github Demo
Star
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
arXiv 2023-11-02 Github -
Star
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
arXiv 2023-10-14 Github Local Demo
Star
SALMONN: Towards Generic Hearing Abilities for Large Language Models
ICLR 2023-10-20 Github -
Star
Ferret: Refer and Ground Anything Anywhere at Any Granularity
arXiv 2023-10-11 Github -
Star
CogVLM: Visual Expert For Large Language Models
arXiv 2023-10-09 Github Demo
Star
Improved Baselines with Visual Instruction Tuning
arXiv 2023-10-05 Github Demo
Star
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
ICLR 2023-10-03 Github Demo
Star
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs arXiv 2023-10-01 Github -
Star
Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants
arXiv 2023-10-01 Github Local Demo
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model arXiv 2023-09-27 - -
Star
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
arXiv 2023-09-26 Github Local Demo
Star
DreamLLM: Synergistic Multimodal Comprehension and Creation
ICLR 2023-09-20 Github Coming soon
An Empirical Study of Scaling Instruction-Tuned Large Multimodal Models arXiv 2023-09-18 Coming soon -
Star
TextBind: Multi-turn Interleaved Multimodal Instruction-following
arXiv 2023-09-14 Github Demo
Star
NExT-GPT: Any-to-Any Multimodal LLM
arXiv 2023-09-11 Github Demo
Star
Sight Beyond Text: Multi-Modal Training Enhances LLMs in Truthfulness and Ethics
arXiv 2023-09-13 Github -
Star
ImageBind-LLM: Multi-modality Instruction Tuning
arXiv 2023-09-07 Github Demo
Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning arXiv 2023-09-05 - -
Star
PointLLM: Empowering Large Language Models to Understand Point Clouds
arXiv 2023-08-31 Github Demo
Star
✨Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
arXiv 2023-08-31 Github Local Demo
Star
MLLM-DataEngine: An Iterative Refinement Approach for MLLM
arXiv 2023-08-25 Github -
Star
Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models
arXiv 2023-08-25 Github Demo
Star
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
arXiv 2023-08-24 Github Demo
Star
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
ICLR 2023-08-23 Github Demo
Star
StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data
arXiv 2023-08-20 Github -
Star
BLIVA: A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions
arXiv 2023-08-19 Github Demo
Star
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
arXiv 2023-08-08 Github -
Star
The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
ICLR 2023-08-03 Github Demo
Star
LISA: Reasoning Segmentation via Large Language Model
arXiv 2023-08-01 Github Demo
Star
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
arXiv 2023-07-31 Github Local Demo
Star
3D-LLM: Injecting the 3D World into Large Language Models
arXiv 2023-07-24 Github -
ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning
arXiv 2023-07-18 - Demo
Star
BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
arXiv 2023-07-17 Github Demo
Star
SVIT: Scaling up Visual Instruction Tuning
arXiv 2023-07-09 Github -
Star
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
arXiv 2023-07-07 Github Demo
Star
What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
arXiv 2023-07-05 Github -
Star
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
arXiv 2023-07-04 Github Demo
Star
Visual Instruction Tuning with Polite Flamingo
arXiv 2023-07-03 Github Demo
Star
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
arXiv 2023-06-29 Github Demo
Star
Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
arXiv 2023-06-27 Github Demo
Star
MotionGPT: Human Motion as a Foreign Language
arXiv 2023-06-26 Github -
Star
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
arXiv 2023-06-15 Github Coming soon
Star
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
arXiv 2023-06-11 Github Demo
Star
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
arXiv 2023-06-08 Github Demo
Star
MIMIC-IT: Multi-Modal In-Context Instruction Tuning
arXiv 2023-06-08 Github Demo
M3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning arXiv 2023-06-07 - -
Star
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
arXiv 2023-06-05 Github Demo
Star
LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
arXiv 2023-06-01 Github -
Star
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
arXiv 2023-05-30 Github Demo
Star
PandaGPT: One Model To Instruction-Follow Them All
arXiv 2023-05-25 Github Demo
Star
ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst
arXiv 2023-05-25 Github -
Star
Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models
arXiv 2023-05-24 Github Local Demo
Star
DetGPT: Detect What You Need via Reasoning
arXiv 2023-05-23 Github Demo
Star
Pengi: An Audio Language Model for Audio Tasks
NeurIPS 2023-05-19 Github -
Star
VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
arXiv 2023-05-18 Github -
Star
Listen, Think, and Understand
arXiv 2023-05-18 Github Demo
Star
VisualGLM-6B
- 2023-05-17 Github Local Demo
Star
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
arXiv 2023-05-17 Github -
Star
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
arXiv 2023-05-11 Github Local Demo
Star
VideoChat: Chat-Centric Video Understanding
arXiv 2023-05-10 Github Demo
Star
MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
arXiv 2023-05-08 Github Demo
Star
X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
arXiv 2023-05-07 Github -
Star
LMEye: An Interactive Perception Network for Large Language Models
arXiv 2023-05-05 Github Local Demo
Star
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
arXiv 2023-04-28 Github Demo
Star
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
arXiv 2023-04-27 Github Demo
Star
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
arXiv 2023-04-20 Github -
Star
Visual Instruction Tuning
NeurIPS 2023-04-17 GitHub Demo
Star
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
ICLR 2023-03-28 Github Demo
Star
MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
ACL 2022-12-21 Github -

Multimodal Hallucination

Title Venue Date Code Demo
Star
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
arXiv 2024-10-04 Github -
Star
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
arXiv 2024-10-03 Github -
FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs arXiv 2024-09-20 Link -
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation arXiv 2024-08-01 - -
Star
Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs
ECCV 2024-07-31 Github -
Star
Evaluating and Analyzing Relationship Hallucinations in LVLMs
ICML 2024-06-24 Github -
Star
AGLA: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
arXiv 2024-06-18 Github -
CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models arXiv 2024-06-04 Coming soon -
Mitigating Object Hallucination via Data Augmented Contrastive Tuning arXiv 2024-05-28 Coming soon -
VDGD: Mitigating LVLM Hallucinations in Cognitive Prompts by Bridging the Visual Perception Gap arXiv 2024-05-24 Coming soon -
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback arXiv 2024-04-22 - -
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding arXiv 2024-03-27 - -
Star
What if...?: Counterfactual Inception to Mitigate Hallucination Effects in Large Multimodal Models
arXiv 2024-03-20 Github -
Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization arXiv 2024-03-13 - -
Star
Debiasing Multimodal Large Language Models
arXiv 2024-03-08 Github -
Star
HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
arXiv 2024-03-01 Github -
IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding arXiv 2024-02-28 - -
Star
Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
arXiv 2024-02-22 Github -
Star
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
arXiv 2024-02-18 Github -
Star
The Instinctive Bias: Spurious Images lead to Hallucination in MLLMs
arXiv 2024-02-06 Github -
Star
Unified Hallucination Detection for Multimodal Large Language Models
arXiv 2024-02-05 Github -
A Survey on Hallucination in Large Vision-Language Models arXiv 2024-02-01 - -
Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models arXiv 2024-01-18 - -
Star
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
arXiv 2023-12-12 Github -
Star
MOCHa: Multi-Objective Reinforcement Mitigating Caption Hallucinations
arXiv 2023-12-06 Github -
Star
Mitigating Fine-Grained Hallucination by Fine-Tuning Large Vision-Language Models with Caption Rewrites
arXiv 2023-12-04 Github -
Star
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
arXiv 2023-12-01 Github Demo
Star
OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
CVPR 2023-11-29 Github -
Star
Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
CVPR 2023-11-28 Github -
Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization arXiv 2023-11-28 Github Comins Soon
Mitigating Hallucination in Visual Language Models with Visual Supervision arXiv 2023-11-27 - -
Star
HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
arXiv 2023-11-22 Github -
Star
An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
arXiv 2023-11-13 Github -
Star
FAITHSCORE: Evaluating Hallucinations in Large Vision-Language Models
arXiv 2023-11-02 Github -
Star
Woodpecker: Hallucination Correction for Multimodal Large Language Models
arXiv 2023-10-24 Github Demo
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models arXiv 2023-10-09 - -
Star
HallE-Switch: Rethinking and Controlling Object Existence Hallucinations in Large Vision Language Models for Detailed Caption
arXiv 2023-10-03 Github -
Star
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
ICLR 2023-10-01 Github -
Star
Aligning Large Multimodal Models with Factually Augmented RLHF
arXiv 2023-09-25 Github Demo
Evaluation and Mitigation of Agnosia in Multimodal Large Language Models arXiv 2023-09-07 - -
CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning arXiv 2023-09-05 - -
Star
Evaluation and Analysis of Hallucination in Large Vision-Language Models
arXiv 2023-08-29 Github -
Star
VIGC: Visual Instruction Generation and Correction
arXiv 2023-08-24 Github Demo
Detecting and Preventing Hallucinations in Large Vision Language Models arXiv 2023-08-11 - -
Star
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
ICLR 2023-06-26 Github Demo
Star
Evaluating Object Hallucination in Large Vision-Language Models
EMNLP 2023-05-17 Github -

Multimodal In-Context Learning

Title Venue Date Code Demo
Visual In-Context Learning for Large Vision-Language Models arXiv 2024-02-18 - -
Star
RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning in Multi-Modal Large Language Model
RSS 2024-02-16 Github -
Star
Can MLLMs Perform Text-to-Image In-Context Learning?
arXiv 2024-02-02 Github -
Star
Generative Multimodal Models are In-Context Learners
CVPR 2023-12-20 Github Demo
Hijacking Context in Large Multi-modal Models arXiv 2023-12-07 - -
Towards More Unified In-context Visual Understanding arXiv 2023-12-05 - -
Star
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
arXiv 2023-09-14 Github Demo
Star
Link-Context Learning for Multimodal LLMs
arXiv 2023-08-15 Github Demo
Star
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
arXiv 2023-08-02 Github Demo
Star
Med-Flamingo: a Multimodal Medical Few-shot Learner
arXiv 2023-07-27 Github Local Demo
Star
Generative Pretraining in Multimodality
ICLR 2023-07-11 Github Demo
AVIS: Autonomous Visual Information Seeking with Large Language Models arXiv 2023-06-13 - -
Star
MIMIC-IT: Multi-Modal In-Context Instruction Tuning
arXiv 2023-06-08 Github Demo
Star
Exploring Diverse In-Context Configurations for Image Captioning
NeurIPS 2023-05-24 Github -
Star
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
arXiv 2023-04-19 Github Demo
Star
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace
arXiv 2023-03-30 Github Demo
Star
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
arXiv 2023-03-20 Github Demo
Star
ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information Extraction
ICCV 2023-03-09 Github -
Star
Prompting Large Language Models with Answer Heuristics for Knowledge-based Visual Question Answering
CVPR 2023-03-03 Github -
Star
Visual Programming: Compositional visual reasoning without training
CVPR 2022-11-18 Github Local Demo
Star
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
AAAI 2022-06-28 Github -
Star
Flamingo: a Visual Language Model for Few-Shot Learning
NeurIPS 2022-04-29 Github Demo
Multimodal Few-Shot Learning with Frozen Language Models NeurIPS 2021-06-25 - -

Multimodal Chain-of-Thought

Title Venue Date Code Demo
Star
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
arXiv 2024-11-21 Github -
Star
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
arXiv 2024-04-24 Github Local Demo
Star
Visual CoT: Unleashing Chain-of-Thought Reasoning in Multi-Modal Language Models
arXiv 2024-03-25 Github Local Demo
Star
Compositional Chain-of-Thought Prompting for Large Multimodal Models
CVPR 2023-11-27 Github -
Star
DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models
NeurIPS 2023-10-25 Github -
Star
Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
arXiv 2023-06-27 Github Demo
Star
Explainable Multimodal Emotion Reasoning
arXiv 2023-06-27 Github -
Star
EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought
arXiv 2023-05-24 Github -
Let’s Think Frame by Frame: Evaluating Video Chain of Thought with Video Infilling and Prediction arXiv 2023-05-23 - -
T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question Answering arXiv 2023-05-05 - -
Star
Caption Anything: Interactive Image Description with Diverse Multimodal Controls
arXiv 2023-05-04 Github Demo
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings arXiv 2023-05-03 Coming soon -
Star
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
arXiv 2023-04-19 Github Demo
Chain of Thought Prompt Tuning in Vision Language Models arXiv 2023-04-16 Coming soon -
Star
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
arXiv 2023-03-20 Github Demo
Star
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
arXiv 2023-03-08 Github Demo
Star
Multimodal Chain-of-Thought Reasoning in Language Models
arXiv 2023-02-02 Github -
Star
Visual Programming: Compositional visual reasoning without training
CVPR 2022-11-18 Github Local Demo
Star
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
NeurIPS 2022-09-20 Github -

LLM-Aided Visual Reasoning

Title Venue Date Code Demo
Star
VideoDeepResearch: Long Video Understanding With Agentic Tool Using
arXiv 2025-06-12 Github Local Demo
Star
Beyond Embeddings: The Promise of Visual Table in Multi-Modal Models
arXiv 2024-03-27 Github -
Star
V∗: Guided Visual Search as a Core Mechanism in Multimodal LLMs
arXiv 2023-12-21 Github Local Demo
Star
LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing
arXiv 2023-11-01 Github Demo
MM-VID: Advancing Video Understanding with GPT-4V(vision) arXiv 2023-10-30 - -
Star
ControlLLM: Augment Language Models with Tools by Searching on Graphs
arXiv 2023-10-26 Github -
Star
Woodpecker: Hallucination Correction for Multimodal Large Language Models
arXiv 2023-10-24 Github Demo
Star
MindAgent: Emergent Gaming Interaction
arXiv 2023-09-18 Github -
Star
Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language
arXiv 2023-06-28 Github Demo
Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models arXiv 2023-06-15 - -
Star
AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
arXiv 2023-06-14 Github -
AVIS: Autonomous Visual Information Seeking with Large Language Models arXiv 2023-06-13 - -
Star
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
arXiv 2023-05-30 Github Demo
Mindstorms in Natural Language-Based Societies of Mind arXiv 2023-05-26 - -
Star
LayoutGPT: Compositional Visual Planning and Generation with Large Language Models
arXiv 2023-05-24 Github -
Star
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
arXiv 2023-05-24 Github Local Demo
Star
Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation
arXiv 2023-05-10 Github -
Star
Caption Anything: Interactive Image Description with Diverse Multimodal Controls
arXiv 2023-05-04 Github Demo
Star
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
arXiv 2023-04-19 Github Demo
Star
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace
arXiv 2023-03-30 Github Demo
Star
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
arXiv 2023-03-20 Github Demo
Star
ViperGPT: Visual Inference via Python Execution for Reasoning
arXiv 2023-03-14 Github Local Demo
Star
ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions
arXiv 2023-03-12 Github Local Demo
ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information Extraction ICCV 2023-03-09 - -
Star
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
arXiv 2023-03-08 Github Demo
Star
Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners
CVPR 2023-03-03 Github -
Star
From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models
CVPR 2022-12-21 Github Demo
Star
SuS-X: Training-Free Name-Only Transfer of Vision-Language Models
arXiv 2022-11-28 Github -
Star
PointCLIP V2: Adapting CLIP for Powerful 3D Open-world Learning
CVPR 2022-11-21 Github -
Star
Visual Programming: Compositional visual reasoning without training
CVPR 2022-11-18 Github Local Demo
Star
Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
arXiv 2022-04-01 Github -

Foundation Models

Title Venue Date Code Demo
Introducing GPT-5 OpenAI 2025-08-07 - -
Star
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
arXiv 2025-01-22 Github Demo
Star
Emu3: Next-Token Prediction is All You Need
arXiv 2024-09-27 Github Local Demo
Llama 3.2: Revolutionizing edge AI and vision with open, customizable models Meta 2024-09-25 - Demo
Pixtral-12B Mistral 2024-09-17 - -
Star
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
arXiv 2024-08-16 Github -
The Llama 3 Herd of Models arXiv 2024-07-31 - -
Chameleon: Mixed-Modal Early-Fusion Foundation Models arXiv 2024-05-16 - -
Hello GPT-4o OpenAI 2024-05-13 - -
The Claude 3 Model Family: Opus, Sonnet, Haiku Anthropic 2024-03-04 - -
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context Google 2024-02-15 - -
Gemini: A Family of Highly Capable Multimodal Models Google 2023-12-06 - -
Fuyu-8B: A Multimodal Architecture for AI Agents Blog 2023-10-17 [Huggingfa