👁️🗨️ Awesome VLM Architectures 
Awesome VLM Architectures is a citation-first visual catalog of 155+ Vision-Language Model (VLM/MLLM) architectures, spanning contrastive encoders, multimodal LLMs, native multimodal models, unified understanding and generation, video, OCR, GUI agents, and embodied AI. Each entry links to primary sources and summarizes the model architecture, modality alignment or fusion, training stages, datasets, and distinctive design choices, with an architectural figure when available.
Use this repository to compare multimodal model families, trace architectural ideas over time, or retrieve grounded references for research and AI-agent workflows. The catalog includes a verified release timeline through July 2026 and covers foundational systems such as CLIP and Flamingo alongside current multimodal reasoning and agentic models. Expand any model panel for its detailed architecture summary.
Last reviewed: August 1, 2026.
Contents
Citation
If this repository is useful in your work, you may cite it below. Please also cite the original paper for claims about any individual model; this catalog is a guide to the literature, not a substitute for it.
Created and maintained by Gökay Aydoğan at fal.ai (ORCID; [email protected]).
Architecture images are credited individually in the figure credits and remain subject to the rights described in the figure notice.
📚 BibTeX
@misc{aydogan2024awesomevlmarchitectures,
author = {Gökay Aydoğan},
title = {Awesome VLM Architectures},
year = {2024},
howpublished = {\url{https://github.com/gokayfem/awesome-vlm-architectures}},
note = {GitHub repository, fal.ai},
url = {https://github.com/gokayfem/awesome-vlm-architectures}
}
Models
All architecture panels are ordered by release date, newest first. Models released on the same day retain editorial catalog order.
🧭 Chronological Model Index (155 architectures, newest first)
2026: MODUS | Argus-Unified | Kimi K3 | Mage-VL | Inkling | Hy-Embodied-VLM | MonkeyOCRv2 | MiniMax M3 | InternVideo3 | Keye-VL 2.0 | Zamba2-VL | Cosmos 3 | Lance | ZAYA1-VL | Falcon Perception | GLM-5V-Turbo | PLaMo 2.1-VL | EXAONE 4.5 | BidirLM and BidirLM-Omni | Gemma 4 | Penguin-VL | Phi-4-Reasoning-Vision | V-SONAR and V-LCM | Qwen3.5 | Youtu-VL | Kimi K2.5 and K2.6 | Step3-VL-10B
2025: ERNIE 5.0 | DeepSeek-OCR | PaddleOCR-VL | Qwen3-VL | Step3 | GLM-4.1V-Thinking | ERNIE 4.5-VL | MiMo-VL | BAGEL | Seed1.5-VL | InternVL3 and InternVL3.5 | Kimi-VL | Llama 4 Scout and Maverick | Qwen2.5-Omni | Gemma 3 | Aya Vision | Phi-4-multimodal | SigLIP 2 | EVEv2 | Qwen2.5-VL | VideoLLaMA 3 | UI-TARS | MiniMax-01 | MiniCPM-o-2.6 | Eagle 2 | Sa2VA
2024: VideoChat-Flash | OmniVLM | Apollo | DeepSeek-VL2 | Maya | InternVL 2.5 | PaliGemma 2 | ShowUI | SmolVLM | AIMv2 | LLaVA-CoT | LLM2CLIP | Tarsier2 | Janus and Janus-Pro | ARIA | Emu3 | Molmo and PixMo | Llama 3.2-Vision | NVLM | Pixtral 12B | VILA-U | Qwen2-VL | EAGLE | Show-o | Idefics3-8B | Transfusion | mPLUG-Owl3 | VITA | LLaVA-OneVision | VILA² | INF-LLaVA | SlowFast-LLaVA | EVLM | InternLM-XComposer-2.5 | OMG-LLaVA | Cambrian-1 | EVE | Ovis | Parrot | ConvLLaVA | Phi-3-Vision and Phi-3.5-Vision | CogVLM2 | Chameleon | PaliGemma | xGen-MM (BLIP-3) | MANTIS | Moondream-next | Idefics2 | InternLM-XComposer2-4KHD | MM1 | DeepSeek-VL | AnyGPT | SPHINX-X | LLaVA 1.6 | MiniCPM-V | MouSi | InternLM-XComposer2 | MoE-LLaVA | moondream1 and moondream2 | FireLLaVA | COSMO
2023: TinyGPT-V | MobileVLM | Alpha-CLIP | Nous-Hermes-2-Vision - Mistral 7B | SPHINX | Florence-2 | u-LLaVA | LLaVA-Plus | OtterHD | CoVLM | GLaMM | Fuyu-8B | PaLI-3 Vision Language Models | MiniGPT-v2 | BakLLaVA | Ferret | LLaVA 1.5 | CogVLM | MetaCLIP | Qwen-VL | IDEFICS | BLIVA | KOSMOS-2 | LaVIN | InstructBLIP | ImageBind | LLaVA | MiniGPT-4 | SigLIP | OpenFlamingo | PaLM-E | KOSMOS-1 | BLIP-2
2022: MULTIINSTRUCT | PaLI | Flamingo | BLIP
2020: ViT
Release Timeline
Dates use the first documented official model release; when none is available, they use the paper's arXiv v1 submission or first technical report. Family point releases are folded into their first architecture release, and same-day entries retain catalog order.
🗓️ Release Timeline (155 architectures, newest first)
| Date | Architecture | Distinctive contribution |
|---|---|---|
| 2026-07-28 | MODUS | Decoder-only any-to-any modeling without modality-specific heads or losses |
| 2026-07-28 | Argus-Unified | Hybrid continuous and discrete visual tokens for economical understanding and generation |
| 2026-07-27 | Kimi K3 | Kimi Delta Attention, Attention Residuals, and extremely sparse LatentMoE routing |
| 2026-07-27 | Mage-VL | Codec-native selective video tokenization with a proactive event gate |
| 2026-07-15 | Inkling | Relative-position million-context multimodal MoE trained from scratch |
| 2026-07-15 | Hy-Embodied-VLM | Action-centric sparse-MoE reasoning for physical-world agents |
| 2026-07-11 | MonkeyOCRv2 | Joint image-to-text and pixel-reconstruction pretraining for document vision |
| 2026-06-11 | MiniMax M3 | Native multimodality with block-sparse grouped-query attention at million-token context |
| 2026-06-10 | InternVideo3 | Token-preserving latent KV compression and closed-loop video reasoning |
| 2026-06-09 | Keye-VL 2.0 | DeepSeek Sparse Attention adapted to GQA-based long-video multimodality |
| 2026-06-02 | Zamba2-VL | Hybrid Mamba-2 and shared-attention blocks for efficient VLM inference |
| 2026-05-31 | Cosmos 3 | Coupled autoregressive reasoner and diffusion generator for physical AI |
| 2026-05-18 | Lance | Shared-sequence understanding, generation, and editing with modality experts |
| 2026-05-08 | ZAYA1-VL | Vision-conditional LoRA and compressed convolutional attention in an open-data MoE |
| 2026-05-03 | Falcon Perception | Early fusion with hybrid attention and continuous mask heads |
| 2026-04-29 | GLM-5V-Turbo | Perception integrated into reasoning, planning, tools, and execution |
| 2026-04-21 | PLaMo 2.1-VL | Compact Japanese VQA and grounding for edge deployment |
| 2026-04-09 | EXAONE 4.5 | Native multimodal pretraining with document-focused data and 256K context |
| 2026-04-02 | BidirLM and BidirLM-Omni | Converting causal decoders into bidirectional multimodal encoders |
| 2026-03-31 | Gemma 4 | Dense and MoE native multimodality, including an encoder-free 12B design |
| 2026-03-06 | Penguin-VL | Text-LLM-initialized vision encoder and priority-aware token compression |
| 2026-03-04 | Phi-4-Reasoning-Vision | Mid-fusion compact VLM with explicit reasoning and direct-answer modes |
| 2026-03-01 | V-SONAR and V-LCM | Vision-language alignment and prediction in multilingual concept space |
| 2026-02-16 | Qwen3.5 | Native early fusion with hybrid linear/full attention and sparse MoE variants |
| 2026-01-27 | Youtu-VL | Unified autoregressive visual tokens that emit dense vision outputs without task heads |
| 2026-01-27 | Kimi K2.5 and K2.6 | Trillion-parameter native multimodal MoE for agents and computer use |
| 2026-01-14 | Step3-VL-10B | Language-aligned perception encoder with 16-fold visual-token compression |
| 2025-11-13 | ERNIE 5.0 | One autoregressive sparse MoE for text, images, video, audio, and generation |
| 2025-10-20 | DeepSeek-OCR | DeepEncoder compresses high-resolution documents into very short visual contexts |
| 2025-10-16 | PaddleOCR-VL | NaViT-style dynamic resolution with a compact ERNIE decoder for document parsing |
| 2025-09-22 | Qwen3-VL | DeepStack multi-level ViT fusion and explicit video timestamp alignment |
| 2025-07-25 | Step3 | Model-system co-design for communication-efficient sparse-MoE multimodality |
| 2025-07-01 | GLM-4.1V-Thinking | Curriculum-sampled reinforcement learning for multimodal reasoning |
| 2025-06-30 | ERNIE 4.5-VL | Heterogeneous shared and modality-specific experts with isolated routing |
| 2025-06-04 | MiMo-VL | Four-stage multimodal pretraining followed by mixed on-policy RL |
| 2025-05-20 | BAGEL | Mixture-of-Transformer-Experts for understanding and generation |
| 2025-05-11 | Seed1.5-VL | Compact vision encoder with a 20B-active MoE for reasoning and agents |
| 2025-04-11 | InternVL3 and InternVL3.5 | Native multimodal pretraining, later extended with adaptive resolution and cascade RL |
| 2025-04-10 | Kimi-VL | MoonViT native-resolution packing with a sparse MoE decoder |
| 2025-04-05 | Llama 4 Scout and Maverick | Early-fusion native multimodality in sparse-MoE Scout and Maverick models |
| 2025-03-26 | Qwen2.5-Omni | Streaming Thinker-Talker architecture for multimodal input and speech output |
| 2025-03-12 | Gemma 3 | Efficient local/global attention with long-context image understanding |
| 2025-03-04 | Aya Vision | Cross-modal model merging for multilingual multimodality without language forgetting |
| 2025-03-03 | Phi-4-multimodal | Mixture-of-LoRAs for text, vision, and speech |
| 2025-02-20 | SigLIP 2 | Multilingual, localization-aware, native-aspect-ratio vision-language encoding |
| 2025-02-08 | EVEv2 | Improved Baselines for Encoder-Free Vision-Language Models |
| 2025-01-26 | Qwen2.5-VL | Enhanced Vision-Language Capabilities in the Qwen Series |
| 2025-01-21 | VideoLLaMA 3 | Frontier Multimodal Foundation Models for Image and Video Understanding |
| 2025-01-20 | UI-TARS | Pioneering Automated GUI Interaction with Native Agents |
| 2025-01-14 | MiniMax-01 | Scaling Foundation Models with Lightning Attention |
| 2025-01-12 | MiniCPM-o-2.6 | A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming |
| 2025-01-10 | Eagle 2 | Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models |
| 2025-01-07 | Sa2VA | Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos |
| 2024-12-31 | VideoChat-Flash | Hierarchical Compression for Long-Context Video Modeling |
| 2024-12-16 | OmniVLM | A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference |
| 2024-12-13 | Apollo | An Exploration of Video Understanding in Large Multimodal Models |
| 2024-12-13 | DeepSeek-VL2 | Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding |
| 2024-12-10 | Maya | An Instruction Finetuned Multilingual Multimodal Model |
| 2024-12-05 | InternVL 2.5 | Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling |
| 2024-12-04 | PaliGemma 2 | A Family of Versatile VLMs for Transfer |
| 2024-11-26 | ShowUI | UI-guided visual-token selection and interleaved action histories |
| 2024-11-26 | SmolVLM | A Small, Efficient, and Open-Source Vision-Language Model |
| 2024-11-21 | AIMv2 | Multimodal Autoregressive Pre-training of Large Vision Encoders |
| 2024-11-15 | LLaVA-CoT | Let Vision Language Models Reason Step-by-Step |
| 2024-11-06 | LLM2CLIP | Powerful Language Model Unlocks Richer Visual Representation |
| 2024-11-05 | Tarsier2 | Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding |
| 2024-10-17 | Janus and Janus-Pro | Decoupled visual encoders for understanding and generation with one transformer |
| 2024-10-08 | ARIA | An Open Multimodal Native Mixture-of-Experts Model |
| 2024-09-27 | Emu3 | One next-token objective over discrete text, image, and video tokens |
| 2024-09-25 | Molmo and PixMo | Open data pipeline with human captions and grounded pointing supervision |
| 2024-09-25 | Llama 3.2-Vision | Enhanced Multimodal Capabilities Built on Llama 3 |
| 2024-09-17 | NVLM | Open Frontier-Class Multimodal LLMs |
| 2024-09-11 | Pixtral 12B | A Cutting-Edge Open Multimodal Language Model |
| 2024-09-06 | VILA-U | Shared discrete visual tokens for autoregressive understanding and generation |
| 2024-08-29 | Qwen2-VL | A Powerful Open-Source Vision-Language Model for Image and Video Understanding |
| 2024-08-28 | EAGLE | Exploring The Design Space for Multimodal LLMs with Mixture of Encoders |
| 2024-08-22 | Show-o | Autoregressive language and discrete-diffusion image generation in one transformer |
| 2024-08-22 | Idefics3-8B | Building and Better Understanding Vision-Language Models |
| 2024-08-20 | Transfusion | Autoregressive text and continuous image diffusion in one transformer |
| 2024-08-09 | mPLUG-Owl3 | Hyper-attention for long image sequences and video |
| 2024-08-09 | VITA | Towards Open-Source Interactive Omni Multimodal LLM |
| 2024-08-05 | LLaVA-OneVision | Easy Visual Task Transfer |
| 2024-07-24 | VILA² | VILA Augmented VILA |
| 2024-07-23 | INF-LLaVA | High-Resolution Image Perception for Multimodal Large Language Models |
| 2024-07-22 | SlowFast-LLaVA | A Strong Training-Free Baseline for Video Large Language Models |
| 2024-07-19 | EVLM | An Efficient Vision-Language Model for Visual Understanding |
| 2024-07-03 | InternLM-XComposer-2.5 | A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output |
| 2024-06-27 | OMG-LLaVA | Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding |
| 2024-06-24 | Cambrian-1 | Spatial Vision Aggregator and systematic multi-encoder study |
| 2024-06-17 | EVE | Unveiling Encoder-Free Vision-Language Models |
| 2024-06-14 | Ovis | Learnable visual vocabulary for structural visual-text embedding alignment |
| 2024-06-04 | Parrot | Multilingual Visual Instruction Tuning |
| 2024-05-24 | ConvLLaVA | Hierarchical Backbones as Visual Encoder for Large Multimodal Models |
| 2024-05-21 | Phi-3-Vision and Phi-3.5-Vision | Compact dynamic-resolution VLM with 128K context |
| 2024-05-20 | CogVLM2 | Enhanced Vision-Language Models for Image and Video Understanding |
| 2024-05-16 | Chameleon | Mixed-modal early fusion over a shared token sequence |
| 2024-05-14 | PaliGemma | A Versatile and Transferable 3B Vision-Language Model |
| 2024-05-06 | xGen-MM (BLIP-3) | An Open-Source Framework for Building Powerful and Responsible Large Multimodal Models |
| 2024-05-02 | MANTIS | Mastering Multi-Image Understanding Through Interleaved Instruction Tuning |
| 2024-04-19 | Moondream-next | Compact Vision-Language Model with Enhanced Capabilities |
| 2024-04-15 | Idefics2 | Open 8B VLM with native-resolution inputs and strong OCR and document understanding |
| 2024-04-09 | InternLM-XComposer2-4KHD | A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD |
| 2024-03-14 | MM1 | Controlled study of encoders, connectors, token counts, and data mixtures |
| 2024-03-08 | DeepSeek-VL | Towards Real-World Vision-Language Understanding |
| 2024-02-19 | AnyGPT | Any-to-any autoregression over discrete text, image, speech, and music tokens |
| 2024-02-08 | SPHINX-X | Scaling Data and Parameters for a Family of Multi-modal Large Language Models |
| 2024-01-30 | LLaVA 1.6 | LLaVA-NeXT Improved reasoning, OCR, and world knowledge |
| 2024-01-30 | MiniCPM-V | A GPT-4V Level MLLM on Your Phone |
| 2024-01-30 | MouSi | Poly-Visual-Expert Vision-Language Models |
| 2024-01-29 | InternLM-XComposer2 | Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model |
| 2024-01-29 | MoE-LLaVA | Mixture of Experts for Large Vision-Language Models |
| 2024-01-20 | moondream1 and moondream2 | Compact SigLIP–Phi VLMs optimized for efficient edge inference |
| 2024-01-05 | FireLLaVA | LLaVA derivative trained rapidly on a curated multimodal instruction mixture |
| 2024-01-01 | COSMO | COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training |
| 2023-12-28 | TinyGPT-V | Efficient Multimodal Large Language Model via Small Backbones |
| 2023-12-28 | MobileVLM | A Fast, Strong and Open Vision Language Assistant for Mobile Devices |
| 2023-12-06 | Alpha-CLIP | A CLIP Model Focusing on Wherever You Want |
| 2023-11-28 | Nous-Hermes-2-Vision - Mistral 7B | SigLIP-equipped Mistral VLM with OCR and function-calling data |
| 2023-11-13 | SPHINX | The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models |
| 2023-11-10 | Florence-2 | A Deep Dive into its Unified Architecture and Multi-Task Capabilities |
| 2023-11-09 | u-LLaVA | Unifying Multi-Modal Tasks via Large Language Model |
| 2023-11-09 | LLaVA-Plus | Learning to Use Tools for Creating Multimodal Agents |
| 2023-11-07 | OtterHD | A High-Resolution Multi-modality Model |
| 2023-11-06 | CoVLM | Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding |
| 2023-11-06 | GLaMM | Pixel Grounding Large Multimodal Model |
| 2023-10-17 | Fuyu-8B | A Multimodal Architecture for AI Agents |
| 2023-10-13 | PaLI-3 Vision Language Models | Smaller, Faster, Stronger |
| 2023-10-13 | MiniGPT-v2 | large language model as a unified interface for vision-language multi-task learning |
| 2023-10-12 | BakLLaVA | Mistral-based LLaVA variant with a CLIP vision encoder and projection adapter |
| 2023-10-11 | Ferret | Refer and Ground Anything Anywhere at Any Granularity |
| 2023-10-05 | LLaVA 1.5 | Improved Baselines with Visual Instruction Tuning |
| 2023-10-05 | CogVLM | Visual Expert for Pretrained Language Models |
| 2023-09-28 | MetaCLIP | Demystifying CLIP Data |
| 2023-08-24 | Qwen-VL | A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond |
| 2023-08-22 | IDEFICS | Open Flamingo-style model for interleaved image-text generation |
| 2023-08-19 | BLIVA | A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions |
| 2023-06-26 | KOSMOS-2 | Grounding Multimodal Large Language Models to the World |
| 2023-05-24 | LaVIN | Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models |
| 2023-05-11 | InstructBLIP | Towards General-purpose Vision-Language Models with Instruction Tuning |
| 2023-05-09 | ImageBind | One Embedding Space To Bind Them All |
| 2023-04-17 | LLaVA | Large Language and Vision Assistant - Visual Instruction Tuning |
| 2023-04-16 | MiniGPT-4 | Enhancing Vision-Language Understanding with Advanced Large Language Models |
| 2023-03-27 | SigLIP | Sigmoid Loss for Language Image Pre-Training |
| 2023-03-14 | OpenFlamingo | An Open-Source Framework for Training Large Autoregressive Vision-Language Models |
| 2023-03-06 | PaLM-E | An Embodied Multimodal Language Model |
| 2023-02-27 | KOSMOS-1 | Language Is Not All You Need: Aligning Perception with Language Models |
| 2023-01-30 | BLIP-2 | Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models |
| 2022-12-21 | MULTIINSTRUCT | Improving Multi-Modal Zero-Shot Learning via Instruction Tuning |
| 2022-09-14 | PaLI | A Jointly-Scaled Multilingual Language-Image Model |
| 2022-04-28 | Flamingo | a Visual Language Model for Few-Shot Learning |
| 2022-01-28 | BLIP | Bootstrapping Language-Image Pre-training |
| 2021-12-07 | GLIP | Grounded Language-Image Pre-training |
| 2021-06-25 | FROZEN | Multimodal Few-Shot Learning with Frozen Language Models |
| 2021-01-05 | CLIP | Contrastive Language-Image Pre-training |
| 2020-10-22 | ViT | An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale |
Architectures
MODUS: Decoder-Only Any-to-Any Multimodal Modeling
MODUS treats every modality symmetrically as both input and output, enabling chained generation and cross-modal self-verification without modality-specific heads, losses, or task pipelines.
Mingqiao Ye et al., EPFL
Released: 2026-07-28

Figure 2. Decoder-only any-to-any modeling across tokenized 1D and 2D modalities. Source paper, PDF p. 4. Figure notice.
ℹ️ More Information
MODUS is a decoder-only any-to-any model that represents diverse modalities inside one autoregressive architecture. Unlike encoder-decoder or diffusion systems assembled around modality-specific output paths, the same model predicts any supported modality from any combination of the others. This makes intermediate-modality chains and self-scoring through a second generated modality native behaviors rather than external workflows.
The design deliberately reuses strong pretrained decoder-only priors instead of training a bespoke multimodal stack from scratch. A single checkpoint is evaluated across heterogeneous tasks and modalities, making MODUS most notable as a general architectural formulation rather than a narrowly optimized VLM endpoint.
Argus-Unified: Economical Understanding and Generation
Argus-Unified combines continuous tokens for understanding with learned discrete tokens for generation, reusing a frozen unified visual encoder and pretrained VLM to lower the cost of unified modeling.
Weiming Zhuang et al.
Released: 2026-07-28

Figure 3. Two-stage hybrid-token training for unified image understanding and generation. Source paper, PDF p. 4. Figure notice.
ℹ️ More Information
Argus-Unified resolves the conflicting visual representations required by comprehension and synthesis through hybrid visual tokens. Continuous encoder features preserve semantic information for understanding, while a learned quantizer produces discrete tokens for image generation from the same frozen visual encoder.
Training first learns the quantizer and image decoder, then initializes the language component from a pretrained VLM and trains unified multimodal prediction. The reported recipe uses 15.6 million examples and roughly $2,000 of compute, positioning the model as a reproducible baseline for understanding-and-generation research rather than a scale demonstration.
Kimi K3: Kimi Delta Attention at Trillion-Parameter Scale
Kimi K3 is a native multimodal sparse MoE that combines Kimi Delta Attention, gated latent attention, Attention Residuals, and 896 routed experts for million-token reasoning and agency.
Kimi Team, Moonshot AI
Released: 2026-07-27

Figure 2. Kimi Delta Attention, Stable LatentMoE, Attention Residuals, and MoonViT-V2. Source paper, PDF p. 3. Figure notice.
ℹ️ More Information
Kimi K3 has 2.8T total and 104B active parameters across 93 layers. Sixty-nine layers use gated delta-rule linear attention and 24 use gated Multi-head Latent Attention; learned Attention Residuals create structured information paths across depth. Stable LatentMoE activates 16 of 896 routed experts plus two shared experts, making the model unusually sparse for its scale.
A 401M-parameter MoonViT-V2 encoder supplies native image and video input. The model supports a 1,048,576-token context and receives multimodal, long-context, agentic, reasoning, and tool-use training. Quantization-aware pretraining targets MXFP4 weights and MXFP8 activations rather than treating low-precision deployment as an afterthought.
Mage-VL: Codec-Native Streaming Multimodality
Mage-VL uses video-codec signals to encode only dynamic, information-rich regions and pairs a lightweight event gate with a causal decoder for proactive streaming perception.
Senqiao Yang et al., Microsoft Research
Released: 2026-07-27

Figure 3. Codec-native streaming perception with an event gate and causal language decoder. Source paper, PDF p. 8. Figure notice.
ℹ️ More Information
Mage-ViT replaces uniform frame sampling with selective 16×16-patch encoding guided by motion vectors and residual energy across anchor and predicted frames. This codec-native tokenization reduces visual-token use by more than 75% while retaining spatiotemporal context. The encoder is trained from scratch on roughly 560M unlabeled images and 100M unlabeled video frames.
The complete VLM uses a dual-system design: a small event gate decides when a stream warrants deeper processing, and a causal decoder performs contextual reasoning. This turns streaming perception into an event-driven process and yields up to a reported 3.5× wall-clock speedup while retaining static-image and spatial reasoning ability.
Inkling: Relative-Position Multimodal Mixture of Experts
Inkling is an open-weight multimodal MoE trained from scratch on text, images, audio, and video, using relative positions and alternating local/global attention for million-token contexts.
Thinking Machines Lab
Released: 2026-07-15
Architecture figure: The official Inkling model card contains no architecture figure.
ℹ️ More Information
Inkling contains 975B total and 41B active parameters. Its MoE layers activate six of 256 routed experts plus two shared experts using sigmoid routing and auxiliary-loss-free balancing. Attention alternates five local sliding-window layers with one global layer and uses learned relative-position representations rather than RoPE; short convolutions further refine key, value, attention, and MLP pathways.
Pretraining spans 45T text, image, audio, and video tokens. Post-training begins with a small synthetic supervised bootstrap and scales asynchronous reinforcement learning across multimodal understanding, tools, coding, mathematics, conversation, and safety. The released weights support controllable reasoning effort and contexts up to one million tokens.
Hy-Embodied-VLM: Sparse-MoE Reasoning for Physical Agents
Hy-Embodied-VLM-1.0 joins a native-aspect-ratio vision encoder to a compact sparse MoE and trains an action-centric reasoning hierarchy for perception, planning, reflection, and recovery.
Tencent Robotics X, Hy Vision Team and Futian Laboratory
Released: 2026-07-15

Figure 4. Self-evolving supervised fine-tuning, rejection sampling, and specialized reinforcement learning. Source paper, PDF p. 11. Figure notice.
ℹ️ More Information
Hy-Embodied-VLM-1.0 connects Hy-ViT2 to a roughly 30B-total, 3B-active language model. Each sparse-MoE layer selects eight of 128 routed experts plus one shared expert. It accepts as many as 128 images in a 32K context and exposes direct-response and explicit-thinking modes through the same checkpoint.
Its curriculum moves from action-relevant state understanding to transition reasoning and then sequential, adaptive reasoning. A self-evolving post-training loop alternates reinforcement learning with rejection-sampling fine-tuning, separately optimizing geometric precision and higher-level planning before policy fusion.
MonkeyOCRv2: Document-Native Visual-Text Pretraining
MonkeyOCRv2 learns document vision jointly through image-to-text generation and pixel reconstruction, preserving both semantic content and character-level strokes across 17 languages.
Yuliang Liu et al.
Released: 2026-07-11

Figure 1. Document-native pretraining through text generation and pixel reconstruction. Source paper, PDF p. 2. Figure notice.
ℹ️ More Information
MonkeyOCRv2 is a visual-text foundation encoder specialized for the properties that natural-image pretraining often misses: dense text, small glyphs, fine strokes, and layout. Its two complementary objectives align images with their transcribed content while forcing the encoder to retain pixel-level document structure.
The MonkeyDoc v2 corpus contains 113M document images in 17 languages. The frozen encoder improves recognition, detection, tamper analysis, segmentation, parsing, and understanding; paired with a lightweight language model, it forms a 0.7B document parser. Public weights were announced July 11, two days before the arXiv report, so the timeline follows the model release rather than the paper date.
MiniMax M3: Native Multimodality with Sparse Long-Context Attention
MiniMax M3 combines native multimodal pretraining with learned block-sparse grouped-query attention, making million-token contexts practical in a 109B-parameter model.
MiniMax
Released: 2026-06-11

Figure 1. MiniMax Sparse Attention index and exact-attention branches. Source paper, PDF p. 1. Figure notice.
ℹ️ More Information
MiniMax Sparse Attention adds a lightweight index branch that scores KV blocks independently for each grouped-query-attention group. The main branch performs exact attention only over selected blocks, while a KV-outer kernel groups queries retrieving the same block to improve memory locality and reuse.
The released 109B-parameter model is trained natively across modalities and evaluated at contexts up to one million tokens. The sparse-attention report focuses on architecture, kernel design, and scaling behavior; it does not publish a complete source-level training-data inventory.
InternVideo3: Multimodal Contextual Reasoning for Video Agents
InternVideo3 reframes long-video understanding as an evolving loop of observation, reasoning, tools, and memory, while compressing KV state without dropping the underlying token stream.
Ziang Yan et al.
Released: 2026-06-10

Figure 2. InternVideo3 with multimodal multi-head latent attention across long contexts. Source paper, PDF p. 7. Figure notice.
ℹ️ More Information
Multimodal Contextual Reasoning maintains one evolving context containing observations, instructions, intermediate reasoning, tool actions, and memory. This closed loop lets the model accumulate and verify evidence across long videos rather than treating video QA as a single static prompt.
Multimodal Multi-head Latent Attention reparameterizes and compresses KV-cache states while preserving the complete visual-token sequence. Training combines continued pretraining, short-to-long supervised tuning, rule-based reinforcement learning, and on-policy distillation; a retrieval-equipped video-agent implementation demonstrates the intended tool-using setting.
Keye-VL 2.0: Sparse Attention for Long-Video Agents
Keye-VL 2.0 adapts DeepSeek Sparse Attention to a GQA-based multimodal MoE, targeting lossless 256K contexts, hour-scale video, and self-correcting tool-using agents.
Kwai Keye Team
Released: 2026-06-09

Figure 2. Four-stage curriculum extending Keye-VL from alignment to 256K context. Source paper, PDF p. 8. Figure notice.
ℹ️ More Information
Keye-VL-2.0-30B-A3B is the first reported adaptation of DeepSeek Sparse Attention to grouped-query multimodal attention. Only 3B of 30B parameters activate per token, while the sparse retrieval mechanism keeps critical frames and long-range dependencies accessible across a 256K context. Heterogeneous ViT-language-model parallelism and custom sparse-attention kernels address the systems cost of hour-long video.
Cross-Modal Multi-Teacher On-Policy Distillation feeds dense token-level guidance from specialist teachers back into on-policy trajectories. Context-RL and Video-RL then target long-context reasoning, temporal localization, code, search, tools, and multimodal self-correction without collapsing the multi-task mixture.
Zamba2-VL: Hybrid State-Space Vision-Language Modeling
Zamba2-VL combines efficient Mamba-2 state-space layers with a small number of shared Transformer blocks, reducing long-context prefill and recurrent-state costs across compact VLM scales.
Zyphra
Released: 2026-06-02

Figure 1. Zamba2 hybrid state-space language backbone connected to a vision encoder. Source paper, PDF p. 3. Figure notice.
ℹ️ More Information
Zamba2-VL extends the Zamba hybrid backbone to vision-language inputs in 1.2B, 2.7B, and 7B variants. Most sequence processing occurs in Mamba-2 state-space layers, with a small set of shared attention blocks supplying global token interaction. This preserves transformer-like multimodal reasoning while moving more computation toward near-linear prefill and bounded recurrent state.
The architecture is particularly relevant for edge and long-context deployment: visual sequences can be large even when the language model is small, so reducing quadratic attention changes time-to-first-token and memory behavior more materially than ordinary parameter pruning.
Cosmos 3: Omnimodal World Modeling with Mixture of Transformers
Cosmos 3 couples an autoregressive reasoner with a diffusion generator through a shared representation for language, vision, audio, actions, simulation, and robot control.
NVIDIA
Released: 2026-05-31

Figure 5. Mixture-of-Transformers reasoner and generator with shared attention. Source paper, PDF p. 11. Figure notice.
ℹ️ More Information
Cosmos 3 uses a Mixture-of-Transformers with two communicating towers. An autoregressive transformer reasons over language, vision, and action sequences; a diffusion transformer generates continuous images, video, and world trajectories. Their shared representation lets the family serve as a VLM, forward or inverse dynamics model, simulator, generator, or policy without separate task pipelines.
Nano configurations combine 8B-class components and Super configurations combine 32B-class components. Training spans text, image, video, audio, simulated physical interaction, action trajectories, robot demonstrations, driving, and synthetic worlds. The July 20 Cosmos 3 Edge release preserves the same two-tower design in a smaller real-time deployment branch and is folded into this family.
Lance: Unified Image and Video Understanding, Generation, and Editing
Lance is a compact model trained from scratch to understand, generate, and edit images and video inside one interleaved sequence while preserving specialized semantic and synthesis capacity.
Lance Team
Released: 2026-05-18

Figure 6. Dual-expert sequence modeling for understanding and visual generation. Source paper, PDF p. 9. Figure notice.
ℹ️ More Information
Lance interleaves text, image, and video in a shared 3B-parameter sequence model while routing semantic understanding and visual synthesis through dedicated experts. Semantic ViT tokens coexist with clean or noised VAE latents, and generalized 3D causal attention plus modality-aware positional encoding preserve spatial and temporal relationships.
One checkpoint supports visual understanding, text-to-image and text-to-video generation, image-to-video generation, and editing. Its training-from-scratch recipe uses no more than 128 A100 GPUs, making the contribution as much about an attainable unified architecture as headline generation quality.
ZAYA1-VL: Vision-Specialized Compressed Convolutional Attention
ZAYA1-VL adds visual capacity to a sparse MoE through vision-only LoRA and compressed-convolutional-attention parameters, trained on an entirely open multimodal data mixture.
Zyphra
Released: 2026-05-08

Figure 2. Visual routing, compressed convolutional attention, and a hybrid language backbone. Source paper, PDF p. 5. Figure notice.
ℹ️ More Information
ZAYA1-VL-8B joins a Qwen2.5-VL visual encoder to ZAYA1's MoE decoder. Vision-only LoRA and Compressed Convolutional Attention parameters activate for visual tokens without duplicating the language backbone, while bidirectional attention within image-token spans improves spatial integration.
The model is trained on roughly 140B vision-language tokens assembled from open data and released under Apache 2.0. Its importance is the clean separation of visual specialization from shared language capacity in a compact sparse model.
Falcon Perception: Early-Fusion Grounding, Segmentation, and OCR
Falcon Perception uses early image-text fusion, hybrid attention, instance tokens, and continuous mask heads to unify open-vocabulary grounding, segmentation, and OCR in compact models.
Technology Innovation Institute
Released: 2026-05-03

Figure 1. Early-fusion perception Transformer with grounding, geometry, and segmentation pathways. Source paper, PDF p. 3. Figure notice.
ℹ️ More Information
Falcon Perception is a dense early-fusion Transformer in which image patches and text interact from the first layer. Its attention mask is bidirectional over image tokens and causal over prediction tokens. Variable-length instance tokens identify multiple objects, while lightweight continuous heads produce masks without turning segmentation into an unwieldy text sequence.
The family includes a 600M perception model and a 300M OCR model. The paper was submitted March 28, but the timeline uses the documented May 3 public model launch, following this repository's release-date policy.
GLM-5V-Turbo: Native Multimodal Agency
GLM-5V-Turbo integrates visual perception into reasoning, planning, tool use, execution, and verification instead of treating images and video as an auxiliary language-model interface.
GLM-V Team
Released: 2026-04-29

Figure 2. Multimodal multi-token prediction with image placeholders and shared Transformer blocks. Source paper, PDF p. 3. Figure notice.
ℹ️ More Information
GLM-5V-Turbo is organized around end-to-end multimodal agency across images, video, webpages, documents, GUIs, code, and external tools. Its contribution is the integration of perception with the agent loop: observations remain available during planning and execution, and outcomes can be visually verified rather than handed off to a detached captioning stage.
The report describes coordinated changes to model design, multimodal training, hierarchical optimization, reinforcement learning, toolchains, and agent-framework integration. Because this is a new native agent foundation model rather than a latency-only checkpoint, it is kept separate from the GLM-4.1V family.
PLaMo 2.1-VL: Lightweight Japanese Vision-Language Modeling
PLaMo 2.1-VL provides compact 2B and 8B Japanese-English VLMs that combine visual question answering and grounding for local, edge, and autonomous-device deployment.
Tommi Kerola et al., Preferred Networks
Released: 2026-04-21
Architecture figure: The PLaMo 2.1-VL paper contains application and data figures, but no model architecture diagram.
ℹ️ More Information
PLaMo 2.1-VL is designed around a constrained deployment envelope rather than frontier scale. Both sizes support VQA and bounding-box grounding, with evaluations emphasizing Japanese language operation and applications such as factory tool recognition and infrastructure anomaly detection.
A large synthetic-data pipeline supplies Japanese multimodal instructions, grounding examples, and domain-oriented supervision. The entry broadens the catalog's language and edge coverage; its contribution is a deployable bilingual VQA-and-grounding family rather than a new attention primitive.
EXAONE 4.5: Native Multimodal Pretraining for Documents
EXAONE 4.5 is LG AI Research's first open-weight VLM, integrating a dedicated visual encoder into EXAONE and emphasizing document-centric native multimodal pretraining and Korean context.
Eunbi Choi et al., LG AI Research
Released: 2026-04-09

Figure 1. Native-resolution vision encoding, projection, language decoding, and multi-token prediction. Source paper, PDF p. 2. Figure notice.
ℹ️ More Information
EXAONE 4.5 adds a dedicated vision encoder to the EXAONE 4.0 framework and trains the resulting stack natively over text and visual inputs. Its data mixture is deliberately document-heavy, targeting enterprise documents, OCR-rich layouts, and Korean contextual reasoning while retaining general language capability.
Context extends to 256K tokens for long documents and enterprise workflows. The combination of open weights, native multimodal training, long context, and focused document/Korean coverage makes it a distinct public VLM release rather than a post-training adapter.
BidirLM and BidirLM-Omni: Causal Decoders as Multimodal Encoders
BidirLM converts pretrained causal language decoders into bidirectional representation encoders; BidirLM-Omni extends the method to one contrastively aligned text, image, and audio embedding model.
BidirLM Team
Released: 2026-04-02

Figure 10. Specialist-backbone merging with frozen modality projection heads. Source paper, PDF p. 25. Figure notice.
ℹ️ More Information
BidirLM removes the causal mask and adapts a decoder with masked next-token prediction and contrastive learning. Weight merging and mixed-domain training preserve generative knowledge while recovering the full-context interaction needed by retrieval and embedding tasks.
BidirLM-Omni combines Qwen3 text, Qwen3-VL, and Qwen3-ASR specialist backbones with frozen visual and audio projectors, aligning all three modalities in one contrastive space. The entry belongs beside CLIP, SigLIP, and ImageBind as an encoder architecture, despite originating from decoder checkpoints.
Gemma 4: Open-Weight Native Multimodal Models
Gemma 4 is Google DeepMind's open-weight multimodal family spanning mobile-efficient, dense, sparse-MoE, and encoder-free designs, with configurable reasoning and deployment targets ranging from phones to servers.
Gemma Team, Google DeepMind
Released: 2026-03-31

Figure 2. Aspect-preserving image resizing, patch pooling, and soft-token production. Source paper, PDF p. 16. Figure notice.
ℹ️ More Information
Gemma 4 is a suite of related architectures rather than one uniform network. E2B and E4B target mobile and edge deployment; 31B is a dense server-class decoder; and 26B-A4B is a sparse Mixture-of-Experts model with roughly 4B active parameters. These branches accept images through improved visual processing, while the mobile models additionally support audio. The distinctive Gemma 4 12B Unified removes separate vision and audio encoders: raw image patches and audio features enter through direct linear projections, reducing duplicated encoder memory and latency. Context windows reach 128K or 256K depending on the variant.
Pretraining jointly develops language, code, reasoning, and multimodal understanding, followed by instruction tuning, safety alignment, and reasoning post-training. A configurable thinking mode lets the instruction-tuned models emit a reasoning trace before answering. The family also ships dedicated multi-token-prediction draft models for speculative decoding, while per-layer embeddings and shared key-value caching improve inference efficiency in applicable variants.
Google describes a large multilingual mixture containing web text, documents, code, mathematics, image-text pairs, OCR and document examples, audio, and video-derived supervision. The complete record-level training corpus is not released, so benchmark collections should not be mistaken for an exhaustive dataset list. Image support includes variable resolutions and aspect ratios; native audio and video coverage depends on the specific branch.
Penguin-VL: Efficient VLMs with LLM-Based Vision Encoders
Penguin-VL challenges the assumption that compact VLMs require contrastively pretrained vision backbones, adapting a text-only language model into a fine-grained visual encoder for efficient image and video understanding.
Boqiang Zhang, Lei Ke, Ruihan Yang, Qi Gao, Tianyuan Qu, Rossell Chen, Dong Yu, Leoweiliang
Released: 2026-03-06

Figure 3. An LLM-initialized vision encoder with priority-aware video-token compression. Source paper, PDF p. 4. Figure notice.
ℹ️ More Information
Penguin-VL builds its Penguin-Encoder by initializing a vision backbone from a text-only Qwen3-0.6B model. Causal attention is converted to bidirectional attention, and 2D rotary position embeddings permit variable-resolution visual tokenization. The encoder is connected to compact language decoders to form 2B- and 8B-class VLMs. For video, Temporal Redundancy-Aware compression assigns larger token budgets to informative key frames and compresses redundant intermediate frames, allowing longer sequences within a fixed visual-token budget.
Training begins with mixed-supervision encoder adaptation. A teacher vision encoder supplies amplitude, direction, and relational distillation targets while image reconstruction-style supervision stabilizes the LLM-to-vision conversion. A low-to-high-resolution curriculum then aligns the encoder for fine-grained perception, followed by unified image and video multimodal pretraining and instruction tuning. This recipe is designed to preserve OCR, spatial, temporal, and dense-captioning signals that category-oriented contrastive objectives can suppress.
The project releases Penguin-Recap-I, a reconstructed high-quality image training set, and Penguin-Recap-V, video annotations at dense timestamp, paragraph, and whole-video granularities. These resources complement broader image-text, document, OCR, mathematical reasoning, and video instruction mixtures described in the report. The authors' comparisons attribute the compact models' gains primarily to visual representation quality and token allocation rather than parameter scaling alone.
Phi-4-Reasoning-Vision: Compact Multimodal Reasoning
Phi-4-Reasoning-Vision-15B is a compact open-weight model that combines high-resolution perception with selectable direct and chain-of-thought modes, focusing on mathematical, scientific, document, and GUI reasoning.
Jyoti Aneja, Michael Harrison, Neel Joshi, Tyler LaBonte, John Langford, Eduardo Salinas
Released: 2026-03-04

Figure 3. SigLIP2 vision encoding, cross-modal projection, mid-fusion, and language reasoning. Source paper, PDF p. 4. Figure notice.
ℹ️ More Information
Phi-4-Reasoning-Vision-15B uses a mid-fusion architecture built from the Phi-4-Reasoning language backbone and a SigLIP 2 vision encoder. Visual features are projected into the language embedding space and inserted into the decoder sequence. Dynamic-resolution processing produces up to 3,600 visual tokens, supporting fine-grained documents and interface screenshots. Bidirectional attention is restricted to tokens within each image, improving spatial interaction while avoiding the overfitting Microsoft observed with broader bidirectional attention.
The model is trained through supervised fine-tuning on a deliberately mixed curriculum of reasoning and non-reasoning examples. Explicit and modes allow one checkpoint to use extended reasoning for mathematics and science or answer directly for perception-heavy tasks such as captioning, OCR, detection, and grounding. Microsoft reports four days of training on 240 NVIDIA B200 GPUs, followed by safety SFT and automated red-team evaluation.
Training data primarily comprises heavily filtered and corrected open-source vision-language datasets, augmented with synthetic examples, internally produced domain data, and targeted acquisitions. Microsoft emphasizes filtering, error correction, and synthetic augmentation as larger contributors than raw scale, but does not publish a complete source-level dataset inventory. Public evaluations span scientific diagrams, charts, mathematical reasoning, OCR, and GUI grounding; those benchmarks are evaluation references, not a claimed exhaustive training list.
V-SONAR and V-LCM: Vision-Language Modeling in Concept Space
V-SONAR aligns image and video representations with SONAR's multilingual concept space, enabling text-trained concept models to process vision. V-LCM then instruction-tunes a latent-diffusion Large Concept Model directly over unified visual and linguistic embeddings.
Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk
Released: 2026-03-01

Figure 1. Visual-semantic alignment and concept-space prediction. Source paper, PDF p. 3. Figure notice.
ℹ️ More Information
V-SONAR starts with Meta's Perception Encoder for image and video features and learns a lightweight projector into the pre-existing SONAR embedding space. Positional embeddings and a temporal-attention layer preserve frame order, while attention pooling produces a modality-level representation compatible with SONAR text concepts. Because the target space is shared with OMNISONAR, the representation can be decoded into text across SONAR's multilingual coverage. A text-only Large Concept Model can also consume these aligned visual embeddings without vision-specific pretraining.
V-SONAR is trained with a coarse-to-fine post-hoc alignment curriculum rather than retraining the foundation encoders from scratch. The alignment uses paired visual-caption supervision and SONAR text embeddings, with staged optimization and asynchronous learning rates for the newly initialized projector and pretrained encoder. V-LCM then concatenates V-SONAR visual concepts with SONAR language concepts and retains LCM's two-tower contextualizer and denoiser architecture. It predicts the next continuous concept embedding with the same latent-diffusion objective used during text-only LCM pretraining.
Vision-language instruction tuning uses M3IT, which spans image and video inputs, eight task categories, and 80 languages. Retrieval and captioning experiments additionally use VATEX, DREAM-1K, PE-Video, and VideoXum. The resulting system separates perception, multilingual semantic alignment, and generative concept modeling instead of converting every modality into a single discrete token vocabulary.
Qwen3.5: Native Multimodal Hybrid-Attention Models
Qwen3.5 unifies text, image, and video understanding in one model family while combining hybrid linear attention with sparse expert routing. Qwen3.6 continues the same native-multimodal direction with stronger agentic coding and updated dense and MoE checkpoints.
Qwen Team
Released: 2026-02-16
Architecture figure: Qwen3.5 has no public technical paper containing an architecture figure.
ℹ️ More Information
Qwen3.5 is a natively multimodal decoder family rather than a separate text model with a later VL edition. Its flagship 397B-A17B checkpoint uses sparse Mixture-of-Experts routing, activating about 17B parameters per token, and combines Gated DeltaNet linear-attention layers with standard softmax-attention layers. Visual inputs are handled inside the same model interface as text, supporting image and video reasoning, tool use, coding, and long-context agent workflows. Smaller dense and MoE variants extend the design across deployment scales.
The training pipeline combines large-scale multimodal pretraining with supervised fine-tuning and reinforcement learning. Qwen describes a scalable asynchronous RL system spanning text, multimodal, and multi-turn interaction, enabling a single checkpoint to switch between thinking and non-thinking behavior. The data mixture covers natural language, code, mathematics, images, video, documents, GUI interactions, tool calls, and multi-turn agent trajectories; the full source-level corpus is not publicly enumerated.
Qwen3.6 is best treated as a continuation of this architecture rather than a separate VLM family. Its open dense and sparse checkpoints retain native multimodality and unified thinking and non-thinking modes while prioritizing stability and agentic coding. Folding these releases into the Qwen3.5 entry keeps the catalog focused on architectural changes rather than model-version churn.
Youtu-VL: Unified Autoregressive Supervision for Dense Vision
Youtu-VL extends the language vocabulary with learned visual codes, making visual tokens prediction targets so one autoregressive model can emit segmentation, depth, pose, and detection without task-specific heads.
Tencent Youtu Lab
Released: 2026-01-27

Figure 3. Unified visual-text autoregressive supervision and dense-output decoding. Source paper, PDF p. 3. Figure notice.
ℹ️ More Information
Vision-Language Unified Autoregressive Supervision learns a visual codebook whose symbols expand the text vocabulary into one multimodal vocabulary. Visual tokens are not passive conditioning embeddings: the decoder predicts them alongside text through joint visual-token and text reconstruction.
Because dense outputs are serialized in the same vocabulary, a standard autoregressive VLM can produce segmentation, monocular depth, human pose, and object detections without separate heads. This turns dense perception into native generation while retaining ordinary visual question answering and instruction following.
Kimi K2.5 and K2.6: Native Multimodal Agentic MoE
Kimi K2.5 is a trillion-parameter native multimodal MoE for reasoning, coding, and computer use; K2.6 keeps the topology while extending context and agentic post-training.
Kimi Team, Moonshot AI
Released: 2026-01-27

Figure 10. Agentic reinforcement-learning environments, rollout management, and training services. Source paper, PDF p. 23. Figure notice.
ℹ️ More Information
Kimi K2.5 connects MoonViT to a 1T-total, 32B-active decoder with 61 layers, Multi-head Latent Attention, and 384 routed experts. Eight routed experts plus shared capacity activate per token. Native multimodal training supports images, video, reasoning, code, tools, and GUI interaction within the same model.
Released April 20, Kimi K2.6 retains the K2.5 architecture while extending context to 256K and strengthening coding, research, and multimodal-agent post-training. It is folded here rather than receiving a duplicate timeline row; Kimi K3 is separate because its attention, residual, sparsity, scale, and vision stack change materially.
Step3-VL-10B: Language-Aligned Perception with 16× Token Compression
Step3-VL-10B combines a 1.8B language-aligned perception encoder with a Qwen3-8B decoder and aggressively compresses high-resolution visual features before language reasoning.
StepFun
Released: 2026-01-14
Architecture figure: The Step3-VL-10B report contains performance and RL figures, but no architecture diagram.
ℹ️ More Information
Step3-VL-10B is architecturally distinct from the much larger 2025 Step3 sparse-MoE system. Its 1.8B perception encoder is aligned to language during pretraining, then a projector with two stride-2 stages reduces the spatial sequence by 16× before feeding a Qwen3-8B decoder.
Images combine one 728×728 global view with 504×504 local crops, treated as independent batch items. Explicit newline tokens mark patch rows while standard one-dimensional RoPE is retained. The PaCoRe parallel proposer-and-synthesizer reasoning method is a post-training and inference addition rather than the defining visual architecture.
ERNIE 5.0: Unified Autoregressive Omnimodal Mixture of Experts
ERNIE 5.0 unifies text, image, video, and audio understanding and generation in a 2.4T ultra-sparse MoE trained through modality-specific forms of grouped next-token prediction.
Baidu ERNIE Team
Released: 2025-11-13

Figure 2. Unified image understanding, image generation, and video generation objectives. Source paper, PDF p. 6. Figure notice.
ℹ️ More Information
ERNIE 5.0 maps language, images, video, and audio into one autoregressive framework using Next-Group-of-Tokens Prediction. Text uses next-token and multi-token prediction; visual generation predicts the next frame and scale; audio generation predicts codec tokens depth-wise. Modality-agnostic expert routing activates less than 3% of the 2.4T-parameter network.
Elastic depth, width, and top-k sparsity train a super-network from which smaller deployment configurations can be extracted. Baidu first unveiled ERNIE 5.0 at Baidu World on November 13, 2025, formally released it in January 2026, and published the technical report on February 4; the timeline therefore uses the earliest documented public disclosure.
DeepSeek-OCR: Visual Context Compression through DeepEncoder
DeepSeek-OCR treats document vision as a context-compression mechanism, representing thousands of text tokens with a much smaller sequence of visual tokens before decoding them with a sparse language model.
DeepSeek-AI
Released: 2025-10-20

Figure 3. A SAM-CLIP DeepEncoder connected to a sparse language decoder. Source paper, PDF p. 5. Figure notice.
ℹ️ More Information
DeepSeek-OCR consists of DeepEncoder and a DeepSeek-3B-MoE decoder with roughly 570M active parameters. DeepEncoder combines a SAM-derived local-perception pathway, a CLIP-derived semantic pathway, convolutional compression, and windowed and global attention to turn high-resolution pages into as few as 64 to 800 visual tokens. Multiple resolution modes trade recognition fidelity for compression.
Training first develops the encoder's visual-text representation and then jointly optimizes document decoding with the MoE language model. Data covers multilingual text pages, natural images, formulas, tables, charts, diagrams, and structured document conversion, supplemented with synthetic renderings. The work frames OCR as an experiment in optical compression for future long-context memory, rather than only a document benchmark model.
DeepSeek-OCR 2, released January 27, 2026, replaces the fixed raster-scan assumption with DeepEncoder V2 and Visual Causal Flow. The encoder dynamically reorders visual tokens according to document semantics before language decoding, exploring whether cascaded one-dimensional causal reasoning can better represent complex two-dimensional layouts. Paper · Repository
PaddleOCR-VL: Ultra-Compact Multilingual Document Parsing
PaddleOCR-VL combines a NaViT-style dynamic-resolution encoder with ERNIE-4.5-0.3B to parse multilingual documents using only 0.9B parameters.
PaddleOCR Team, Baidu
Released: 2025-10-16

Figure 2. Document layout analysis, compact VLM inference, instructions, and structured output. Source paper, PDF p. 5. Figure notice.
ℹ️ More Information
PaddleOCR-VL-0.9B connects a NaViT-style dynamic-resolution visual encoder to the compact ERNIE-4.5-0.3B language model. Variable-resolution packing preserves page detail and aspect ratio without forcing every document through a fixed square canvas. The decoder emits text and structured representations for paragraphs, tables, formulas, charts, and other document elements across 109 languages.
Its curriculum combines element-level recognition with page-level document parsing, followed by instruction tuning on complex layouts and structured outputs. Training data includes multilingual OCR, synthetic and scanned documents, tables, mathematical expressions, charts, reading-order annotations, and layout-rich pages. PaddleOCR-VL-1.5 (January 2026) adds robust physical-document distortions, seal recognition, and text spotting; 1.6 (June 2026) adds region-aware data optimization and progressive reinforcement-learning post-training without changing the core architecture.
Qwen3-VL: DeepStack Vision-Language Models
Qwen3-VL expands the Qwen vision-language family with dense and sparse-MoE checkpoints, native long multimodal context, and stronger spatial, video, OCR, visual-agent, and visual-coding capabilities.
Qwen Team
Released: 2025-09-22

Figure 1. Vision encoding, DeepStack injection, and dense or mixture-of-experts decoding. Source paper, PDF p. 3. Figure notice.
ℹ️ More Information
Qwen3-VL combines a native-dynamic-resolution Vision Transformer with Qwen3 language backbones in dense 2B, 4B, 8B, and 32B configurations and sparse 30B-A3B and 235B-A22B Mixture-of-Experts configurations. Its DeepStack mechanism injects features from multiple ViT depths into corresponding language-model layers instead of relying only on the visual encoder's final layer. Interleaved-MRoPE distributes temporal, height, and width position information across rotary-embedding frequencies, while explicit text-timestamp alignment improves event localization in video. The family supports interleaved text, images, and video in a native 256K-token context, with documented extrapolation for longer contexts.
Training proceeds through large-scale multimodal pretraining followed by supervised instruction tuning and reasoning-oriented post-training. The released Instruct checkpoints target direct response and agent interaction, whereas Thinking checkpoints generate extended reasoning for visual mathematics, spatial analysis, and other difficult multimodal tasks. The curriculum emphasizes recognition, multilingual OCR, document structure, 2D and 3D grounding, long-video understanding, GUI operation, and code generation from visual inputs.
The technical report describes a broad mixture of text, image-text, document, OCR, grounding, chart, multi-image, video, and agent-interaction data. Qwen does not publish a reproducible itemized list of every pretraining source, so individual benchmark datasets should not be presented as the complete training corpus.
Step3: Model-System Co-Design for Cost-Effective Multimodal Intelligence
Step3 is a 321B-total, 38B-active multimodal MoE whose architecture and serving system are co-designed to reduce the communication cost of sparse expert decoding.
StepFun
Released: 2025-07-25

Figure 6. Attention-FFN disaggregation across attention and expert instances. Source paper, PDF p. 11. Figure notice.
ℹ️ More Information
Step3 uses a 321B-parameter sparse-MoE decoder with approximately 38B parameters active per token. The language stack contains dense lower layers followed by routed expert layers, while its vision pathway supports images and video for multimodal reasoning. The model is designed together with an Attention-FFN Disaggregation serving architecture, separating attention and expert computation so hardware placement and communication topology match their different workloads.
Training combines large-scale text and multimodal pretraining with instruction tuning and reasoning-oriented reinforcement learning. The data mixture spans language, code, mathematics, image-text, OCR, documents, charts, multi-image, video, and multimodal reasoning. The system report discloses architecture and deployment mechanisms more clearly than the source-level training corpus, which remains categorically described.
GLM-4.1V-Thinking: General-Purpose Multimodal Reasoning through Curriculum-Sampled RL
GLM-4.1V-9B-Thinking combines a native-resolution AIMv2-based vision stack with long-chain-of-thought post-training and Reinforcement Learning with Curriculum Sampling to strengthen reasoning across visual, document, video, grounding, and agent tasks.
GLM-V Team, Zhipu AI and Tsinghua University
Released: 2025-07-01

Figure 2. Native-resolution vision encoding, projection, decoding, and timestamped video tokens. Source paper, PDF p. 4. Figure notice.
ℹ️ More Information
GLM-4.1V-9B-Thinking joins an AIMv2-Huge-initialized vision encoder, an MLP projector, and the GLM-4-9B-0414 language decoder. The vision tower replaces two-dimensional patch convolutions with 3D convolutions, temporally downsampling video while duplicating single images for consistent processing. Interpolated absolute position embeddings and 2D RoPE support native resolutions and extreme aspect ratios; the language decoder applies 3D RoPE to multimodal tokens. Explicit timestamp tokens after video frames supply temporal position and distance cues.
Training begins with broad multimodal pretraining, followed by continual training on video, higher-resolution imagery, and sequences up to 32K tokens. Supervised fine-tuning then teaches a standardized long-reasoning format with separate and spans. The final stage introduces Reinforcement Learning with Curriculum Sampling (RLCS). Rather than sampling a fixed task mixture, RLCS estimates changing model competence and prioritizes informative problems of suitable difficulty. Domain-specific rule, model, and hybrid rewards cover both verifiable and open-ended tasks.
Pretraining draws on curated caption pairs, knowledge-rich web and academic-book interleaving, 220M OCR images, natural-image and GUI grounding, instructional video, and text replay. SFT and RL span STEM reasoning, charts, documents, video, grounding, coding, GUI agents, and general visual instruction following. The later GLM-4.5V and GLM-4.6V families scale and extend this framework; they are successors, not separate architectures in this catalog.
ERNIE 4.5-VL: Heterogeneous Modality Mixture-of-Experts
ERNIE 4.5-VL combines shared and modality-specific experts with isolated routing and balancing losses so visual learning can reinforce rather than degrade language capability.
Baidu ERNIE Team
Released: 2025-06-30
Architecture figure: ERNIE 4.5-VL has no public paper containing an extractable architecture figure.
ℹ️ More Information
ERNIE 4.5-VL includes 424B-total/47B-active and 28B-total/3B-active multimodal MoE models, plus dense members of the broader ERNIE 4.5 family. Its heterogeneous modality MoE shares some experts across text and visual tokens while reserving other experts for individual modalities. Modality-isolated routing, router orthogonality loss, and multimodal token-balancing loss reduce competition between modalities.
The models are jointly pretrained on text, images, and video, then post-trained in both direct-response and thinking modes. PaddlePaddle infrastructure uses heterogeneous hybrid parallelism, hierarchical load balancing, FP8 mixed precision, and fine-grained recomputation to train and serve the sparse models efficiently. Baidu describes broad text, visual-knowledge, OCR, document, chart, video, and reasoning mixtures but does not provide a complete dataset manifest.
MiMo-VL: Multimodal Pretraining with Mixed On-Policy Reinforcement Learning
MiMo-VL combines a four-stage, 2.4T-token multimodal curriculum with mixed on-policy reinforcement learning across perception, reasoning, and GUI grounding.
Xiaomi MiMo Team
Released: 2025-06-04

Figure 2. Native-resolution vision encoding, projection, and language decoding. Source paper, PDF p. 6. Figure notice.
ℹ️ More Information
MiMo-VL-7B joins a high-resolution vision encoder and projector to a 7B language decoder and is released as supervised and reinforcement-learned checkpoints. Its primary contribution is the training system rather than an exotic connector: long-chain-of-thought examples are introduced during pretraining, not only during post-training, so multimodal reasoning is developed alongside perception and language.
Four pretraining stages consume approximately 2.4T tokens and progressively align vision, general multimodal knowledge, high-quality reasoning, and long-context capabilities. Mixed On-Policy Reinforcement Learning then trains one checkpoint across general QA, mathematics, grounding, and GUI tasks using domain-appropriate rule and model rewards. The mixture includes text replay, image-text, OCR, document, chart, video, spatial-grounding, GUI, and reasoning data; a complete source-level inventory is not published.
BAGEL: A Mixture-of-Transformer-Experts for Unified Understanding and Generation
BAGEL unifies visual understanding, autoregressive language modeling, and rectified-flow image generation in a decoder-only Mixture-of-Transformer-Experts whose modality-specific parameters communicate through shared self-attention.
Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan
Released: 2025-05-20

Figure 2. Shared self-attention with understanding and generation Transformer experts. Source paper, PDF p. 4. Figure notice.
ℹ️ More Information
BAGEL is a unified decoder-only model with 14B total and 7B active parameters. Its Mixture-of-Transformer-Experts architecture duplicates the Qwen2.5-initialized transformer into understanding and generation experts: text and SigLIP 2 ViT tokens use the understanding parameters, while FLUX-VAE latent tokens use the generation parameters. Both experts operate on one interleaved sequence through shared self-attention at every layer, avoiding a small connector bottleneck between comprehension and synthesis. Text is learned with next-token prediction; images are generated with rectified flow.
A generalized causal-attention scheme lets later text or images attend to earlier clean visual representations while preventing access to noised generation targets. Visual understanding uses a native-aspect-ratio-adapted SigLIP2-So400m encoder and MLP connector; generation uses a frozen FLUX VAE. Training progresses through alignment, large-scale pretraining, continued training, and supervised fine-tuning, jointly optimizing cross-entropy and flow-matching objectives.
BAGEL is pretrained on trillions of tokens assembled from text-only, paired image-text, and interleaved multimodal sources. The reported inventory includes 400M text samples, 500M understanding pairs, 1.6B generation pairs, 100M interleaved understanding samples, 45M video-derived sequences, and 20M web documents. OCR, charts, grounding, editing sets, and 500K reasoning-augmented generation and manipulation examples support document reading, spatial control, image editing, and long-context multimodal reasoning.
Seed1.5-VL: Sparse-MoE Multimodal Understanding and Agentic Reasoning
Seed1.5-VL couples a 532M-parameter vision encoder with a 20B-active sparse-MoE decoder for image, video, 2D and 3D grounding, GUI interaction, and long-chain-of-thought reasoning.
ByteDance Seed Team
Released: 2025-05-11

Figure 1. Native-resolution vision encoding, adaptation, sparse MoE decoding, and timestamped video. Source paper, PDF p. 5. Figure notice.
ℹ️ More Information
Seed1.5-VL uses a compact vision tower and a large sparse-MoE language model, maintaining 20B active decoder parameters while scaling total capacity substantially higher. Its visual pathway supports native-aspect-ratio images, multiple images, and video, with explicit support for spatial grounding and GUI control. The same model can provide short direct answers or extended multimodal reasoning.
Training progresses through vision-language pretraining, instruction tuning, long-chain-of-thought cold start, and reinforcement learning. The report emphasizes careful balancing of general text, visual knowledge, OCR and documents, charts, grounding, video, 3D understanding, GUI trajectories, games, and verifiable multimodal-reasoning problems. Source-level corpus provenance and all mixture proportions are not publicly reproducible. The arXiv v1 date is used because an earlier generally accessible official launch date is not unambiguously documented.
InternVL3 and InternVL3.5: Native Multimodal Pretraining and Adaptive Resolution
InternVL3 moves the InternVL family to native multimodal pretraining, while InternVL3.5 adds coarse-to-fine reinforcement learning and dynamically routed visual resolution.
InternVL Team, OpenGVLab
Released: 2025-04-11
Architecture figure: The InternVL3 paper contains evaluation figures, but no definitive model architecture diagram.
ℹ️ More Information
InternVL3 retains the InternViT-MLP-LLM structure and dynamic image tiling of earlier InternVL releases, but jointly learns language and multimodal knowledge during native pretraining instead of treating vision alignment primarily as a later adaptation stage. The family spans dense and sparse-MoE language backbones. Its training recipe combines multimodal pretraining, supervised fine-tuning, Mixed Preference Optimization, and test-time scaling for visual reasoning.
Released on August 26, 2025, InternVL3.5 is part of the same evolving family. It introduces Cascade RL, using offline RL for stable coarse alignment followed by online RL for refinement. A Visual Resolution Router selects visual-token resolution according to input complexity, while Decoupled Vision-Language Deployment places the vision encoder and language model on different devices to balance inference load. Training covers image, multi-image, document, chart, video, GUI, grounding, multilingual, and reasoning data; full pretraining provenance is not enumerated.
Kimi-VL: Native-Resolution Vision with a Sparse MoE Decoder
Kimi-VL couples the native-resolution MoonViT encoder to a sparse Mixture-of-Experts language decoder, activating 2.8B of 16B decoder parameters while supporting images, video, agent interaction, and 128K-token contexts.
Kimi Team
Released: 2025-04-10

Figure 3. MoonViT, multimodal projection, and a sparse mixture-of-experts decoder. Source paper, PDF p. 3. Figure notice.
ℹ️ More Information
Kimi-VL-A3B consists of MoonViT, a two-layer MLP projector, and a Moonlight-derived MoE language model with 16B total and 2.8B activated decoder parameters. MoonViT packs variable-sized patch sequences without tiling images into separately processed crops. It combines interpolated SigLIP-SO400M position embeddings with two-dimensional RoPE, preserving pretrained visual knowledge while representing fine spatial positions. A 2×2 pixel-shuffle operation compresses its output before projection into the language model.
After the language checkpoint's 5.2T-token text-only phase, Kimi-VL undergoes standalone vision training, vision-language alignment, joint pretraining, a high-quality cooldown, and long-context activation. These stages consume 4.4T additional tokens, including 2T for MoonViT training, 0.1T for alignment, and 2.3T across joint stages. Context is progressively extended from 8K to 128K. The instruct model receives joint supervised fine-tuning. Kimi-VL-Thinking adds long-chain-of-thought SFT and reinforcement learning; the later Thinking-2506 checkpoint is folded into this family.
The multimodal mixture is organized into six categories: caption, interleaved image-text, OCR, knowledge, video, and agent data. MoonViT learns from alt text, synthetic captions, grounding boxes, and OCR targets using both SigLIP-style contrastive and captioning losses. Long-video, long-document, academic visual QA, GUI-grounding, and text replay data provide long-context and agent capabilities while preserving language performance.
Llama 4 Scout and Maverick: Native Multimodal Mixture-of-Experts Models
Llama 4 introduces early-fusion multimodality to the Llama family through sparse Mixture-of-Experts decoders with 17B active parameters and contexts extending from one million to ten million tokens.
Meta
Released: 2025-04-05
Architecture figure: The official Llama 4 model card describes the architecture in text and tables only.
ℹ️ More Information
Llama 4 Scout and Llama 4 Maverick are natively multimodal, autoregressive MoE models rather than text decoders retrofitted only during instruction tuning. Scout has 109B total parameters across 16 experts and Maverick approximately 400B across 128 experts; both activate about 17B parameters per token. Early fusion places image and multilingual text tokens in a shared model sequence. Scout supports a 10M-token context, whereas Maverick supports 1M tokens.
Scout was trained on roughly 40T tokens and Maverick on roughly 22T. Meta describes a mixture of public data, licensed data, information from its products and services, and interactions with Meta AI, without releasing a complete corpus manifest. Post-training combines supervised fine-tuning, online reinforcement learning, and direct preference optimization. Maverick was co-distilled from the unreleased Llama 4 Behemoth teacher; Scout also uses distillation during training.
Qwen2.5-Omni: Streaming Multimodal Perception and Speech Generation
Qwen2.5-Omni understands text, images, audio, and video while generating text and natural speech through a streaming Thinker-Talker architecture.
Qwen Team
Released: 2025-03-26

Figure 2. Thinker-Talker architecture for multimodal perception and streaming speech. Source paper, PDF p. 3. Figure notice.
ℹ️ More Information
Qwen2.5-Omni is an end-to-end any-input model whose block-wise audio and vision encoders accept streaming speech, images, and video alongside text. Time-aligned Multimodal RoPE (TMRoPE) interleaves audio and video representations on a synchronized timeline. Its Thinker-Talker design separates semantic reasoning from speech realization: the Thinker produces text and hidden representations, while a dual-track autoregressive Talker conditions on those states to generate speech tokens concurrently.
A sliding-window diffusion transformer converts audio tokens into waveform output without waiting for a complete response, reducing first-packet latency. Training jointly aligns text, vision, video, audio, and speech-generation objectives, followed by multimodal instruction tuning. The report describes caption, OCR, video, audio-transcription, speech, and general text mixtures but does not publish a reproducible source-level inventory.
Gemma 3: Long-Context Multimodality with Efficient Interleaved Attention
Gemma 3 combines a frozen SigLIP vision encoder with a decoder-only language model whose five-local-to-one-global attention pattern reduces long-context KV-cache cost while retaining 128K-token multimodal context.
Gemma Team
Released: 2025-03-12
Architecture figure: The Gemma 3 report contains examples and analysis charts, but no architecture diagram.
ℹ️ More Information
Gemma 3 is a family of 1B, 4B, 12B, and 27B open-weight models; the 4B, 12B, and 27B variants a