← Open Source
gokayfem

awesome-vlm-architectures

Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.

ListsModel collectionsMarkdown
Open on GitHub
Momentum
+4stars in 24 hours+0.3%
1.33k
Stars
56
Forks
+6
This week
1
Contributors
Created 2024-02-15 · Updated 2026-10-07 · #2094 today
Top developers
README

👁️‍🗨️ Awesome VLM Architectures Awesome

VLM

Awesome VLM Architectures is a citation-first visual catalog of 155+ Vision-Language Model (VLM/MLLM) architectures, spanning contrastive encoders, multimodal LLMs, native multimodal models, unified understanding and generation, video, OCR, GUI agents, and embodied AI. Each entry links to primary sources and summarizes the model architecture, modality alignment or fusion, training stages, datasets, and distinctive design choices, with an architectural figure when available.

Use this repository to compare multimodal model families, trace architectural ideas over time, or retrieve grounded references for research and AI-agent workflows. The catalog includes a verified release timeline through July 2026 and covers foundational systems such as CLIP and Flamingo alongside current multimodal reasoning and agentic models. Expand any model panel for its detailed architecture summary.

Last reviewed: August 1, 2026.

Contents

Citation

If this repository is useful in your work, you may cite it below. Please also cite the original paper for claims about any individual model; this catalog is a guide to the literature, not a substitute for it.

Created and maintained by Gökay Aydoğan at fal.ai (ORCID; [email protected]).

Architecture images are credited individually in the figure credits and remain subject to the rights described in the figure notice.

📚 BibTeX

@misc{aydogan2024awesomevlmarchitectures,
  author       = {Gökay Aydoğan},
  title        = {Awesome VLM Architectures},
  year         = {2024},
  howpublished = {\url{https://github.com/gokayfem/awesome-vlm-architectures}},
  note         = {GitHub repository, fal.ai},
  url          = {https://github.com/gokayfem/awesome-vlm-architectures}
}

Models

All architecture panels are ordered by release date, newest first. Models released on the same day retain editorial catalog order.

🧭 Chronological Model Index (155 architectures, newest first)

2026: MODUS | Argus-Unified | Kimi K3 | Mage-VL | Inkling | Hy-Embodied-VLM | MonkeyOCRv2 | MiniMax M3 | InternVideo3 | Keye-VL 2.0 | Zamba2-VL | Cosmos 3 | Lance | ZAYA1-VL | Falcon Perception | GLM-5V-Turbo | PLaMo 2.1-VL | EXAONE 4.5 | BidirLM and BidirLM-Omni | Gemma 4 | Penguin-VL | Phi-4-Reasoning-Vision | V-SONAR and V-LCM | Qwen3.5 | Youtu-VL | Kimi K2.5 and K2.6 | Step3-VL-10B

2025: ERNIE 5.0 | DeepSeek-OCR | PaddleOCR-VL | Qwen3-VL | Step3 | GLM-4.1V-Thinking | ERNIE 4.5-VL | MiMo-VL | BAGEL | Seed1.5-VL | InternVL3 and InternVL3.5 | Kimi-VL | Llama 4 Scout and Maverick | Qwen2.5-Omni | Gemma 3 | Aya Vision | Phi-4-multimodal | SigLIP 2 | EVEv2 | Qwen2.5-VL | VideoLLaMA 3 | UI-TARS | MiniMax-01 | MiniCPM-o-2.6 | Eagle 2 | Sa2VA

2024: VideoChat-Flash | OmniVLM | Apollo | DeepSeek-VL2 | Maya | InternVL 2.5 | PaliGemma 2 | ShowUI | SmolVLM | AIMv2 | LLaVA-CoT | LLM2CLIP | Tarsier2 | Janus and Janus-Pro | ARIA | Emu3 | Molmo and PixMo | Llama 3.2-Vision | NVLM | Pixtral 12B | VILA-U | Qwen2-VL | EAGLE | Show-o | Idefics3-8B | Transfusion | mPLUG-Owl3 | VITA | LLaVA-OneVision | VILA² | INF-LLaVA | SlowFast-LLaVA | EVLM | InternLM-XComposer-2.5 | OMG-LLaVA | Cambrian-1 | EVE | Ovis | Parrot | ConvLLaVA | Phi-3-Vision and Phi-3.5-Vision | CogVLM2 | Chameleon | PaliGemma | xGen-MM (BLIP-3) | MANTIS | Moondream-next | Idefics2 | InternLM-XComposer2-4KHD | MM1 | DeepSeek-VL | AnyGPT | SPHINX-X | LLaVA 1.6 | MiniCPM-V | MouSi | InternLM-XComposer2 | MoE-LLaVA | moondream1 and moondream2 | FireLLaVA | COSMO

2023: TinyGPT-V | MobileVLM | Alpha-CLIP | Nous-Hermes-2-Vision - Mistral 7B | SPHINX | Florence-2 | u-LLaVA | LLaVA-Plus | OtterHD | CoVLM | GLaMM | Fuyu-8B | PaLI-3 Vision Language Models | MiniGPT-v2 | BakLLaVA | Ferret | LLaVA 1.5 | CogVLM | MetaCLIP | Qwen-VL | IDEFICS | BLIVA | KOSMOS-2 | LaVIN | InstructBLIP | ImageBind | LLaVA | MiniGPT-4 | SigLIP | OpenFlamingo | PaLM-E | KOSMOS-1 | BLIP-2

2022: MULTIINSTRUCT | PaLI | Flamingo | BLIP

2021: GLIP | FROZEN | CLIP

2020: ViT

Release Timeline

Dates use the first documented official model release; when none is available, they use the paper's arXiv v1 submission or first technical report. Family point releases are folded into their first architecture release, and same-day entries retain catalog order.

🗓️ Release Timeline (155 architectures, newest first)

Date Architecture Distinctive contribution
2026-07-28 MODUS Decoder-only any-to-any modeling without modality-specific heads or losses
2026-07-28 Argus-Unified Hybrid continuous and discrete visual tokens for economical understanding and generation
2026-07-27 Kimi K3 Kimi Delta Attention, Attention Residuals, and extremely sparse LatentMoE routing
2026-07-27 Mage-VL Codec-native selective video tokenization with a proactive event gate
2026-07-15 Inkling Relative-position million-context multimodal MoE trained from scratch
2026-07-15 Hy-Embodied-VLM Action-centric sparse-MoE reasoning for physical-world agents
2026-07-11 MonkeyOCRv2 Joint image-to-text and pixel-reconstruction pretraining for document vision
2026-06-11 MiniMax M3 Native multimodality with block-sparse grouped-query attention at million-token context
2026-06-10 InternVideo3 Token-preserving latent KV compression and closed-loop video reasoning
2026-06-09 Keye-VL 2.0 DeepSeek Sparse Attention adapted to GQA-based long-video multimodality
2026-06-02 Zamba2-VL Hybrid Mamba-2 and shared-attention blocks for efficient VLM inference
2026-05-31 Cosmos 3 Coupled autoregressive reasoner and diffusion generator for physical AI
2026-05-18 Lance Shared-sequence understanding, generation, and editing with modality experts
2026-05-08 ZAYA1-VL Vision-conditional LoRA and compressed convolutional attention in an open-data MoE
2026-05-03 Falcon Perception Early fusion with hybrid attention and continuous mask heads
2026-04-29 GLM-5V-Turbo Perception integrated into reasoning, planning, tools, and execution
2026-04-21 PLaMo 2.1-VL Compact Japanese VQA and grounding for edge deployment
2026-04-09 EXAONE 4.5 Native multimodal pretraining with document-focused data and 256K context
2026-04-02 BidirLM and BidirLM-Omni Converting causal decoders into bidirectional multimodal encoders
2026-03-31 Gemma 4 Dense and MoE native multimodality, including an encoder-free 12B design
2026-03-06 Penguin-VL Text-LLM-initialized vision encoder and priority-aware token compression
2026-03-04 Phi-4-Reasoning-Vision Mid-fusion compact VLM with explicit reasoning and direct-answer modes
2026-03-01 V-SONAR and V-LCM Vision-language alignment and prediction in multilingual concept space
2026-02-16 Qwen3.5 Native early fusion with hybrid linear/full attention and sparse MoE variants
2026-01-27 Youtu-VL Unified autoregressive visual tokens that emit dense vision outputs without task heads
2026-01-27 Kimi K2.5 and K2.6 Trillion-parameter native multimodal MoE for agents and computer use
2026-01-14 Step3-VL-10B Language-aligned perception encoder with 16-fold visual-token compression
2025-11-13 ERNIE 5.0 One autoregressive sparse MoE for text, images, video, audio, and generation
2025-10-20 DeepSeek-OCR DeepEncoder compresses high-resolution documents into very short visual contexts
2025-10-16 PaddleOCR-VL NaViT-style dynamic resolution with a compact ERNIE decoder for document parsing
2025-09-22 Qwen3-VL DeepStack multi-level ViT fusion and explicit video timestamp alignment
2025-07-25 Step3 Model-system co-design for communication-efficient sparse-MoE multimodality
2025-07-01 GLM-4.1V-Thinking Curriculum-sampled reinforcement learning for multimodal reasoning
2025-06-30 ERNIE 4.5-VL Heterogeneous shared and modality-specific experts with isolated routing
2025-06-04 MiMo-VL Four-stage multimodal pretraining followed by mixed on-policy RL
2025-05-20 BAGEL Mixture-of-Transformer-Experts for understanding and generation
2025-05-11 Seed1.5-VL Compact vision encoder with a 20B-active MoE for reasoning and agents
2025-04-11 InternVL3 and InternVL3.5 Native multimodal pretraining, later extended with adaptive resolution and cascade RL
2025-04-10 Kimi-VL MoonViT native-resolution packing with a sparse MoE decoder
2025-04-05 Llama 4 Scout and Maverick Early-fusion native multimodality in sparse-MoE Scout and Maverick models
2025-03-26 Qwen2.5-Omni Streaming Thinker-Talker architecture for multimodal input and speech output
2025-03-12 Gemma 3 Efficient local/global attention with long-context image understanding
2025-03-04 Aya Vision Cross-modal model merging for multilingual multimodality without language forgetting
2025-03-03 Phi-4-multimodal Mixture-of-LoRAs for text, vision, and speech
2025-02-20 SigLIP 2 Multilingual, localization-aware, native-aspect-ratio vision-language encoding
2025-02-08 EVEv2 Improved Baselines for Encoder-Free Vision-Language Models
2025-01-26 Qwen2.5-VL Enhanced Vision-Language Capabilities in the Qwen Series
2025-01-21 VideoLLaMA 3 Frontier Multimodal Foundation Models for Image and Video Understanding
2025-01-20 UI-TARS Pioneering Automated GUI Interaction with Native Agents
2025-01-14 MiniMax-01 Scaling Foundation Models with Lightning Attention
2025-01-12 MiniCPM-o-2.6 A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
2025-01-10 Eagle 2 Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
2025-01-07 Sa2VA Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
2024-12-31 VideoChat-Flash Hierarchical Compression for Long-Context Video Modeling
2024-12-16 OmniVLM A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference
2024-12-13 Apollo An Exploration of Video Understanding in Large Multimodal Models
2024-12-13 DeepSeek-VL2 Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
2024-12-10 Maya An Instruction Finetuned Multilingual Multimodal Model
2024-12-05 InternVL 2.5 Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
2024-12-04 PaliGemma 2 A Family of Versatile VLMs for Transfer
2024-11-26 ShowUI UI-guided visual-token selection and interleaved action histories
2024-11-26 SmolVLM A Small, Efficient, and Open-Source Vision-Language Model
2024-11-21 AIMv2 Multimodal Autoregressive Pre-training of Large Vision Encoders
2024-11-15 LLaVA-CoT Let Vision Language Models Reason Step-by-Step
2024-11-06 LLM2CLIP Powerful Language Model Unlocks Richer Visual Representation
2024-11-05 Tarsier2 Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
2024-10-17 Janus and Janus-Pro Decoupled visual encoders for understanding and generation with one transformer
2024-10-08 ARIA An Open Multimodal Native Mixture-of-Experts Model
2024-09-27 Emu3 One next-token objective over discrete text, image, and video tokens
2024-09-25 Molmo and PixMo Open data pipeline with human captions and grounded pointing supervision
2024-09-25 Llama 3.2-Vision Enhanced Multimodal Capabilities Built on Llama 3
2024-09-17 NVLM Open Frontier-Class Multimodal LLMs
2024-09-11 Pixtral 12B A Cutting-Edge Open Multimodal Language Model
2024-09-06 VILA-U Shared discrete visual tokens for autoregressive understanding and generation
2024-08-29 Qwen2-VL A Powerful Open-Source Vision-Language Model for Image and Video Understanding
2024-08-28 EAGLE Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
2024-08-22 Show-o Autoregressive language and discrete-diffusion image generation in one transformer
2024-08-22 Idefics3-8B Building and Better Understanding Vision-Language Models
2024-08-20 Transfusion Autoregressive text and continuous image diffusion in one transformer
2024-08-09 mPLUG-Owl3 Hyper-attention for long image sequences and video
2024-08-09 VITA Towards Open-Source Interactive Omni Multimodal LLM
2024-08-05 LLaVA-OneVision Easy Visual Task Transfer
2024-07-24 VILA² VILA Augmented VILA
2024-07-23 INF-LLaVA High-Resolution Image Perception for Multimodal Large Language Models
2024-07-22 SlowFast-LLaVA A Strong Training-Free Baseline for Video Large Language Models
2024-07-19 EVLM An Efficient Vision-Language Model for Visual Understanding
2024-07-03 InternLM-XComposer-2.5 A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
2024-06-27 OMG-LLaVA Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
2024-06-24 Cambrian-1 Spatial Vision Aggregator and systematic multi-encoder study
2024-06-17 EVE Unveiling Encoder-Free Vision-Language Models
2024-06-14 Ovis Learnable visual vocabulary for structural visual-text embedding alignment
2024-06-04 Parrot Multilingual Visual Instruction Tuning
2024-05-24 ConvLLaVA Hierarchical Backbones as Visual Encoder for Large Multimodal Models
2024-05-21 Phi-3-Vision and Phi-3.5-Vision Compact dynamic-resolution VLM with 128K context
2024-05-20 CogVLM2 Enhanced Vision-Language Models for Image and Video Understanding
2024-05-16 Chameleon Mixed-modal early fusion over a shared token sequence
2024-05-14 PaliGemma A Versatile and Transferable 3B Vision-Language Model
2024-05-06 xGen-MM (BLIP-3) An Open-Source Framework for Building Powerful and Responsible Large Multimodal Models
2024-05-02 MANTIS Mastering Multi-Image Understanding Through Interleaved Instruction Tuning
2024-04-19 Moondream-next Compact Vision-Language Model with Enhanced Capabilities
2024-04-15 Idefics2 Open 8B VLM with native-resolution inputs and strong OCR and document understanding
2024-04-09 InternLM-XComposer2-4KHD A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
2024-03-14 MM1 Controlled study of encoders, connectors, token counts, and data mixtures
2024-03-08 DeepSeek-VL Towards Real-World Vision-Language Understanding
2024-02-19 AnyGPT Any-to-any autoregression over discrete text, image, speech, and music tokens
2024-02-08 SPHINX-X Scaling Data and Parameters for a Family of Multi-modal Large Language Models
2024-01-30 LLaVA 1.6 LLaVA-NeXT Improved reasoning, OCR, and world knowledge
2024-01-30 MiniCPM-V A GPT-4V Level MLLM on Your Phone
2024-01-30 MouSi Poly-Visual-Expert Vision-Language Models
2024-01-29 InternLM-XComposer2 Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
2024-01-29 MoE-LLaVA Mixture of Experts for Large Vision-Language Models
2024-01-20 moondream1 and moondream2 Compact SigLIP–Phi VLMs optimized for efficient edge inference
2024-01-05 FireLLaVA LLaVA derivative trained rapidly on a curated multimodal instruction mixture
2024-01-01 COSMO COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
2023-12-28 TinyGPT-V Efficient Multimodal Large Language Model via Small Backbones
2023-12-28 MobileVLM A Fast, Strong and Open Vision Language Assistant for Mobile Devices
2023-12-06 Alpha-CLIP A CLIP Model Focusing on Wherever You Want
2023-11-28 Nous-Hermes-2-Vision - Mistral 7B SigLIP-equipped Mistral VLM with OCR and function-calling data
2023-11-13 SPHINX The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
2023-11-10 Florence-2 A Deep Dive into its Unified Architecture and Multi-Task Capabilities
2023-11-09 u-LLaVA Unifying Multi-Modal Tasks via Large Language Model
2023-11-09 LLaVA-Plus Learning to Use Tools for Creating Multimodal Agents
2023-11-07 OtterHD A High-Resolution Multi-modality Model
2023-11-06 CoVLM Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding
2023-11-06 GLaMM Pixel Grounding Large Multimodal Model
2023-10-17 Fuyu-8B A Multimodal Architecture for AI Agents
2023-10-13 PaLI-3 Vision Language Models Smaller, Faster, Stronger
2023-10-13 MiniGPT-v2 large language model as a unified interface for vision-language multi-task learning
2023-10-12 BakLLaVA Mistral-based LLaVA variant with a CLIP vision encoder and projection adapter
2023-10-11 Ferret Refer and Ground Anything Anywhere at Any Granularity
2023-10-05 LLaVA 1.5 Improved Baselines with Visual Instruction Tuning
2023-10-05 CogVLM Visual Expert for Pretrained Language Models
2023-09-28 MetaCLIP Demystifying CLIP Data
2023-08-24 Qwen-VL A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
2023-08-22 IDEFICS Open Flamingo-style model for interleaved image-text generation
2023-08-19 BLIVA A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions
2023-06-26 KOSMOS-2 Grounding Multimodal Large Language Models to the World
2023-05-24 LaVIN Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models
2023-05-11 InstructBLIP Towards General-purpose Vision-Language Models with Instruction Tuning
2023-05-09 ImageBind One Embedding Space To Bind Them All
2023-04-17 LLaVA Large Language and Vision Assistant - Visual Instruction Tuning
2023-04-16 MiniGPT-4 Enhancing Vision-Language Understanding with Advanced Large Language Models
2023-03-27 SigLIP Sigmoid Loss for Language Image Pre-Training
2023-03-14 OpenFlamingo An Open-Source Framework for Training Large Autoregressive Vision-Language Models
2023-03-06 PaLM-E An Embodied Multimodal Language Model
2023-02-27 KOSMOS-1 Language Is Not All You Need: Aligning Perception with Language Models
2023-01-30 BLIP-2 Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
2022-12-21 MULTIINSTRUCT Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
2022-09-14 PaLI A Jointly-Scaled Multilingual Language-Image Model
2022-04-28 Flamingo a Visual Language Model for Few-Shot Learning
2022-01-28 BLIP Bootstrapping Language-Image Pre-training
2021-12-07 GLIP Grounded Language-Image Pre-training
2021-06-25 FROZEN Multimodal Few-Shot Learning with Frozen Language Models
2021-01-05 CLIP Contrastive Language-Image Pre-training
2020-10-22 ViT An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Architectures

MODUS: Decoder-Only Any-to-Any Multimodal Modeling

MODUS treats every modality symmetrically as both input and output, enabling chained generation and cross-modal self-verification without modality-specific heads, losses, or task pipelines.

arXiv Project

Mingqiao Ye et al., EPFL

Released: 2026-07-28

MODUS: Decoder-Only Any-to-Any Multimodal Modeling architecture: Decoder-only any-to-any modeling across tokenized 1D and 2D modalities

Figure 2. Decoder-only any-to-any modeling across tokenized 1D and 2D modalities. Source paper, PDF p. 4. Figure notice.

ℹ️ More Information

MODUS is a decoder-only any-to-any model that represents diverse modalities inside one autoregressive architecture. Unlike encoder-decoder or diffusion systems assembled around modality-specific output paths, the same model predicts any supported modality from any combination of the others. This makes intermediate-modality chains and self-scoring through a second generated modality native behaviors rather than external workflows.

The design deliberately reuses strong pretrained decoder-only priors instead of training a bespoke multimodal stack from scratch. A single checkpoint is evaluated across heterogeneous tasks and modalities, making MODUS most notable as a general architectural formulation rather than a narrowly optimized VLM endpoint.

Argus-Unified: Economical Understanding and Generation

Argus-Unified combines continuous tokens for understanding with learned discrete tokens for generation, reusing a frozen unified visual encoder and pretrained VLM to lower the cost of unified modeling.

arXiv

Weiming Zhuang et al.

Released: 2026-07-28

Argus-Unified: Economical Understanding and Generation architecture: Two-stage hybrid-token training for unified image understanding and generation

Figure 3. Two-stage hybrid-token training for unified image understanding and generation. Source paper, PDF p. 4. Figure notice.

ℹ️ More Information

Argus-Unified resolves the conflicting visual representations required by comprehension and synthesis through hybrid visual tokens. Continuous encoder features preserve semantic information for understanding, while a learned quantizer produces discrete tokens for image generation from the same frozen visual encoder.

Training first learns the quantizer and image decoder, then initializes the language component from a pretrained VLM and trains unified multimodal prediction. The reported recipe uses 15.6 million examples and roughly $2,000 of compute, positioning the model as a reproducible baseline for understanding-and-generation research rather than a scale demonstration.

Kimi K3: Kimi Delta Attention at Trillion-Parameter Scale

Kimi K3 is a native multimodal sparse MoE that combines Kimi Delta Attention, gated latent attention, Attention Residuals, and 896 routed experts for million-token reasoning and agency.

arXiv GitHub HuggingFace

Kimi Team, Moonshot AI

Released: 2026-07-27

Kimi K3: Kimi Delta Attention at Trillion-Parameter Scale architecture: Kimi Delta Attention, Stable LatentMoE, Attention Residuals, and MoonViT-V2

Figure 2. Kimi Delta Attention, Stable LatentMoE, Attention Residuals, and MoonViT-V2. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

Kimi K3 has 2.8T total and 104B active parameters across 93 layers. Sixty-nine layers use gated delta-rule linear attention and 24 use gated Multi-head Latent Attention; learned Attention Residuals create structured information paths across depth. Stable LatentMoE activates 16 of 896 routed experts plus two shared experts, making the model unusually sparse for its scale.

A 401M-parameter MoonViT-V2 encoder supplies native image and video input. The model supports a 1,048,576-token context and receives multimodal, long-context, agentic, reasoning, and tool-use training. Quantization-aware pretraining targets MXFP4 weights and MXFP8 activations rather than treating low-precision deployment as an afterthought.

Mage-VL: Codec-Native Streaming Multimodality

Mage-VL uses video-codec signals to encode only dynamic, information-rich regions and pairs a lightweight event gate with a causal decoder for proactive streaming perception.

arXiv Project

Senqiao Yang et al., Microsoft Research

Released: 2026-07-27

Mage-VL: Codec-Native Streaming Multimodality architecture: Codec-native streaming perception with an event gate and causal language decoder

Figure 3. Codec-native streaming perception with an event gate and causal language decoder. Source paper, PDF p. 8. Figure notice.

ℹ️ More Information

Mage-ViT replaces uniform frame sampling with selective 16×16-patch encoding guided by motion vectors and residual energy across anchor and predicted frames. This codec-native tokenization reduces visual-token use by more than 75% while retaining spatiotemporal context. The encoder is trained from scratch on roughly 560M unlabeled images and 100M unlabeled video frames.

The complete VLM uses a dual-system design: a small event gate decides when a stream warrants deeper processing, and a causal decoder performs contextual reasoning. This turns streaming perception into an event-driven process and yields up to a reported 3.5× wall-clock speedup while retaining static-image and spatial reasoning ability.

Inkling: Relative-Position Multimodal Mixture of Experts

Inkling is an open-weight multimodal MoE trained from scratch on text, images, audio, and video, using relative positions and alternating local/global attention for million-token contexts.

Website Model Card HuggingFace

Thinking Machines Lab

Released: 2026-07-15

Architecture figure: The official Inkling model card contains no architecture figure.

ℹ️ More Information

Inkling contains 975B total and 41B active parameters. Its MoE layers activate six of 256 routed experts plus two shared experts using sigmoid routing and auxiliary-loss-free balancing. Attention alternates five local sliding-window layers with one global layer and uses learned relative-position representations rather than RoPE; short convolutions further refine key, value, attention, and MLP pathways.

Pretraining spans 45T text, image, audio, and video tokens. Post-training begins with a small synthetic supervised bootstrap and scales asynchronous reinforcement learning across multimodal understanding, tools, coding, mathematics, conversation, and safety. The released weights support controllable reasoning effort and contexts up to one million tokens.

Hy-Embodied-VLM: Sparse-MoE Reasoning for Physical Agents

Hy-Embodied-VLM-1.0 joins a native-aspect-ratio vision encoder to a compact sparse MoE and trains an action-centric reasoning hierarchy for perception, planning, reflection, and recovery.

arXiv GitHub HuggingFace

Tencent Robotics X, Hy Vision Team and Futian Laboratory

Released: 2026-07-15

Hy-Embodied-VLM: Sparse-MoE Reasoning for Physical Agents architecture: Self-evolving supervised fine-tuning, rejection sampling, and specialized reinforcement learning

Figure 4. Self-evolving supervised fine-tuning, rejection sampling, and specialized reinforcement learning. Source paper, PDF p. 11. Figure notice.

ℹ️ More Information

Hy-Embodied-VLM-1.0 connects Hy-ViT2 to a roughly 30B-total, 3B-active language model. Each sparse-MoE layer selects eight of 128 routed experts plus one shared expert. It accepts as many as 128 images in a 32K context and exposes direct-response and explicit-thinking modes through the same checkpoint.

Its curriculum moves from action-relevant state understanding to transition reasoning and then sequential, adaptive reasoning. A self-evolving post-training loop alternates reinforcement learning with rejection-sampling fine-tuning, separately optimizing geometric precision and higher-level planning before policy fusion.

MonkeyOCRv2: Document-Native Visual-Text Pretraining

MonkeyOCRv2 learns document vision jointly through image-to-text generation and pixel reconstruction, preserving both semantic content and character-level strokes across 17 languages.

arXiv GitHub

Yuliang Liu et al.

Released: 2026-07-11

MonkeyOCRv2: Document-Native Visual-Text Pretraining architecture: Document-native pretraining through text generation and pixel reconstruction

Figure 1. Document-native pretraining through text generation and pixel reconstruction. Source paper, PDF p. 2. Figure notice.

ℹ️ More Information

MonkeyOCRv2 is a visual-text foundation encoder specialized for the properties that natural-image pretraining often misses: dense text, small glyphs, fine strokes, and layout. Its two complementary objectives align images with their transcribed content while forcing the encoder to retain pixel-level document structure.

The MonkeyDoc v2 corpus contains 113M document images in 17 languages. The frozen encoder improves recognition, detection, tamper analysis, segmentation, parsing, and understanding; paired with a lightweight language model, it forms a 0.7B document parser. Public weights were announced July 11, two days before the arXiv report, so the timeline follows the model release rather than the paper date.

MiniMax M3: Native Multimodality with Sparse Long-Context Attention

MiniMax M3 combines native multimodal pretraining with learned block-sparse grouped-query attention, making million-token contexts practical in a 109B-parameter model.

arXiv Website HuggingFace

MiniMax

Released: 2026-06-11

MiniMax M3: Native Multimodality with Sparse Long-Context Attention architecture: MiniMax Sparse Attention index and exact-attention branches

Figure 1. MiniMax Sparse Attention index and exact-attention branches. Source paper, PDF p. 1. Figure notice.

ℹ️ More Information

MiniMax Sparse Attention adds a lightweight index branch that scores KV blocks independently for each grouped-query-attention group. The main branch performs exact attention only over selected blocks, while a KV-outer kernel groups queries retrieving the same block to improve memory locality and reuse.

The released 109B-parameter model is trained natively across modalities and evaluated at contexts up to one million tokens. The sparse-attention report focuses on architecture, kernel design, and scaling behavior; it does not publish a complete source-level training-data inventory.

InternVideo3: Multimodal Contextual Reasoning for Video Agents

InternVideo3 reframes long-video understanding as an evolving loop of observation, reasoning, tools, and memory, while compressing KV state without dropping the underlying token stream.

arXiv

Ziang Yan et al.

Released: 2026-06-10

InternVideo3: Multimodal Contextual Reasoning for Video Agents architecture: InternVideo3 with multimodal multi-head latent attention across long contexts

Figure 2. InternVideo3 with multimodal multi-head latent attention across long contexts. Source paper, PDF p. 7. Figure notice.

ℹ️ More Information

Multimodal Contextual Reasoning maintains one evolving context containing observations, instructions, intermediate reasoning, tool actions, and memory. This closed loop lets the model accumulate and verify evidence across long videos rather than treating video QA as a single static prompt.

Multimodal Multi-head Latent Attention reparameterizes and compresses KV-cache states while preserving the complete visual-token sequence. Training combines continued pretraining, short-to-long supervised tuning, rule-based reinforcement learning, and on-policy distillation; a retrieval-equipped video-agent implementation demonstrates the intended tool-using setting.

Keye-VL 2.0: Sparse Attention for Long-Video Agents

Keye-VL 2.0 adapts DeepSeek Sparse Attention to a GQA-based multimodal MoE, targeting lossless 256K contexts, hour-scale video, and self-correcting tool-using agents.

arXiv

Kwai Keye Team

Released: 2026-06-09

Keye-VL 2.0: Sparse Attention for Long-Video Agents architecture: Four-stage curriculum extending Keye-VL from alignment to 256K context

Figure 2. Four-stage curriculum extending Keye-VL from alignment to 256K context. Source paper, PDF p. 8. Figure notice.

ℹ️ More Information

Keye-VL-2.0-30B-A3B is the first reported adaptation of DeepSeek Sparse Attention to grouped-query multimodal attention. Only 3B of 30B parameters activate per token, while the sparse retrieval mechanism keeps critical frames and long-range dependencies accessible across a 256K context. Heterogeneous ViT-language-model parallelism and custom sparse-attention kernels address the systems cost of hour-long video.

Cross-Modal Multi-Teacher On-Policy Distillation feeds dense token-level guidance from specialist teachers back into on-policy trajectories. Context-RL and Video-RL then target long-context reasoning, temporal localization, code, search, tools, and multimodal self-correction without collapsing the multi-task mixture.

Zamba2-VL: Hybrid State-Space Vision-Language Modeling

Zamba2-VL combines efficient Mamba-2 state-space layers with a small number of shared Transformer blocks, reducing long-context prefill and recurrent-state costs across compact VLM scales.

arXiv Project

Zyphra

Released: 2026-06-02

Zamba2-VL: Hybrid State-Space Vision-Language Modeling architecture: Zamba2 hybrid state-space language backbone connected to a vision encoder

Figure 1. Zamba2 hybrid state-space language backbone connected to a vision encoder. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

Zamba2-VL extends the Zamba hybrid backbone to vision-language inputs in 1.2B, 2.7B, and 7B variants. Most sequence processing occurs in Mamba-2 state-space layers, with a small set of shared attention blocks supplying global token interaction. This preserves transformer-like multimodal reasoning while moving more computation toward near-linear prefill and bounded recurrent state.

The architecture is particularly relevant for edge and long-context deployment: visual sequences can be large even when the language model is small, so reducing quadratic attention changes time-to-first-token and memory behavior more materially than ordinary parameter pruning.

Cosmos 3: Omnimodal World Modeling with Mixture of Transformers

Cosmos 3 couples an autoregressive reasoner with a diffusion generator through a shared representation for language, vision, audio, actions, simulation, and robot control.

arXiv GitHub HuggingFace

NVIDIA

Released: 2026-05-31

Cosmos 3: Omnimodal World Modeling with Mixture of Transformers architecture: Mixture-of-Transformers reasoner and generator with shared attention

Figure 5. Mixture-of-Transformers reasoner and generator with shared attention. Source paper, PDF p. 11. Figure notice.

ℹ️ More Information

Cosmos 3 uses a Mixture-of-Transformers with two communicating towers. An autoregressive transformer reasons over language, vision, and action sequences; a diffusion transformer generates continuous images, video, and world trajectories. Their shared representation lets the family serve as a VLM, forward or inverse dynamics model, simulator, generator, or policy without separate task pipelines.

Nano configurations combine 8B-class components and Super configurations combine 32B-class components. Training spans text, image, video, audio, simulated physical interaction, action trajectories, robot demonstrations, driving, and synthetic worlds. The July 20 Cosmos 3 Edge release preserves the same two-tower design in a smaller real-time deployment branch and is folded into this family.

Lance: Unified Image and Video Understanding, Generation, and Editing

Lance is a compact model trained from scratch to understand, generate, and edit images and video inside one interleaved sequence while preserving specialized semantic and synthesis capacity.

arXiv Project

Lance Team

Released: 2026-05-18

Lance: Unified Image and Video Understanding, Generation, and Editing architecture: Dual-expert sequence modeling for understanding and visual generation

Figure 6. Dual-expert sequence modeling for understanding and visual generation. Source paper, PDF p. 9. Figure notice.

ℹ️ More Information

Lance interleaves text, image, and video in a shared 3B-parameter sequence model while routing semantic understanding and visual synthesis through dedicated experts. Semantic ViT tokens coexist with clean or noised VAE latents, and generalized 3D causal attention plus modality-aware positional encoding preserve spatial and temporal relationships.

One checkpoint supports visual understanding, text-to-image and text-to-video generation, image-to-video generation, and editing. Its training-from-scratch recipe uses no more than 128 A100 GPUs, making the contribution as much about an attainable unified architecture as headline generation quality.

ZAYA1-VL: Vision-Specialized Compressed Convolutional Attention

ZAYA1-VL adds visual capacity to a sparse MoE through vision-only LoRA and compressed-convolutional-attention parameters, trained on an entirely open multimodal data mixture.

arXiv Project HuggingFace

Zyphra

Released: 2026-05-08

ZAYA1-VL: Vision-Specialized Compressed Convolutional Attention architecture: Visual routing, compressed convolutional attention, and a hybrid language backbone

Figure 2. Visual routing, compressed convolutional attention, and a hybrid language backbone. Source paper, PDF p. 5. Figure notice.

ℹ️ More Information

ZAYA1-VL-8B joins a Qwen2.5-VL visual encoder to ZAYA1's MoE decoder. Vision-only LoRA and Compressed Convolutional Attention parameters activate for visual tokens without duplicating the language backbone, while bidirectional attention within image-token spans improves spatial integration.

The model is trained on roughly 140B vision-language tokens assembled from open data and released under Apache 2.0. Its importance is the clean separation of visual specialization from shared language capacity in a compact sparse model.

Falcon Perception: Early-Fusion Grounding, Segmentation, and OCR

Falcon Perception uses early image-text fusion, hybrid attention, instance tokens, and continuous mask heads to unify open-vocabulary grounding, segmentation, and OCR in compact models.

arXiv Project Release

Technology Innovation Institute

Released: 2026-05-03

Falcon Perception: Early-Fusion Grounding, Segmentation, and OCR architecture: Early-fusion perception Transformer with grounding, geometry, and segmentation pathways

Figure 1. Early-fusion perception Transformer with grounding, geometry, and segmentation pathways. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

Falcon Perception is a dense early-fusion Transformer in which image patches and text interact from the first layer. Its attention mask is bidirectional over image tokens and causal over prediction tokens. Variable-length instance tokens identify multiple objects, while lightweight continuous heads produce masks without turning segmentation into an unwieldy text sequence.

The family includes a 600M perception model and a 300M OCR model. The paper was submitted March 28, but the timeline uses the documented May 3 public model launch, following this repository's release-date policy.

GLM-5V-Turbo: Native Multimodal Agency

GLM-5V-Turbo integrates visual perception into reasoning, planning, tool use, execution, and verification instead of treating images and video as an auxiliary language-model interface.

arXiv

GLM-V Team

Released: 2026-04-29

GLM-5V-Turbo: Native Multimodal Agency architecture: Multimodal multi-token prediction with image placeholders and shared Transformer blocks

Figure 2. Multimodal multi-token prediction with image placeholders and shared Transformer blocks. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

GLM-5V-Turbo is organized around end-to-end multimodal agency across images, video, webpages, documents, GUIs, code, and external tools. Its contribution is the integration of perception with the agent loop: observations remain available during planning and execution, and outcomes can be visually verified rather than handed off to a detached captioning stage.

The report describes coordinated changes to model design, multimodal training, hierarchical optimization, reinforcement learning, toolchains, and agent-framework integration. Because this is a new native agent foundation model rather than a latency-only checkpoint, it is kept separate from the GLM-4.1V family.

PLaMo 2.1-VL: Lightweight Japanese Vision-Language Modeling

PLaMo 2.1-VL provides compact 2B and 8B Japanese-English VLMs that combine visual question answering and grounding for local, edge, and autonomous-device deployment.

arXiv

Tommi Kerola et al., Preferred Networks

Released: 2026-04-21

Architecture figure: The PLaMo 2.1-VL paper contains application and data figures, but no model architecture diagram.

ℹ️ More Information

PLaMo 2.1-VL is designed around a constrained deployment envelope rather than frontier scale. Both sizes support VQA and bounding-box grounding, with evaluations emphasizing Japanese language operation and applications such as factory tool recognition and infrastructure anomaly detection.

A large synthetic-data pipeline supplies Japanese multimodal instructions, grounding examples, and domain-oriented supervision. The entry broadens the catalog's language and edge coverage; its contribution is a deployable bilingual VQA-and-grounding family rather than a new attention primitive.

EXAONE 4.5: Native Multimodal Pretraining for Documents

EXAONE 4.5 is LG AI Research's first open-weight VLM, integrating a dedicated visual encoder into EXAONE and emphasizing document-centric native multimodal pretraining and Korean context.

arXiv

Eunbi Choi et al., LG AI Research

Released: 2026-04-09

EXAONE 4.5: Native Multimodal Pretraining for Documents architecture: Native-resolution vision encoding, projection, language decoding, and multi-token prediction

Figure 1. Native-resolution vision encoding, projection, language decoding, and multi-token prediction. Source paper, PDF p. 2. Figure notice.

ℹ️ More Information

EXAONE 4.5 adds a dedicated vision encoder to the EXAONE 4.0 framework and trains the resulting stack natively over text and visual inputs. Its data mixture is deliberately document-heavy, targeting enterprise documents, OCR-rich layouts, and Korean contextual reasoning while retaining general language capability.

Context extends to 256K tokens for long documents and enterprise workflows. The combination of open weights, native multimodal training, long context, and focused document/Korean coverage makes it a distinct public VLM release rather than a post-training adapter.

BidirLM and BidirLM-Omni: Causal Decoders as Multimodal Encoders

BidirLM converts pretrained causal language decoders into bidirectional representation encoders; BidirLM-Omni extends the method to one contrastively aligned text, image, and audio embedding model.

arXiv HuggingFace Release

BidirLM Team

Released: 2026-04-02

BidirLM and BidirLM-Omni: Causal Decoders as Multimodal Encoders architecture: Specialist-backbone merging with frozen modality projection heads

Figure 10. Specialist-backbone merging with frozen modality projection heads. Source paper, PDF p. 25. Figure notice.

ℹ️ More Information

BidirLM removes the causal mask and adapts a decoder with masked next-token prediction and contrastive learning. Weight merging and mixed-domain training preserve generative knowledge while recovering the full-context interaction needed by retrieval and embedding tasks.

BidirLM-Omni combines Qwen3 text, Qwen3-VL, and Qwen3-ASR specialist backbones with frozen visual and audio projectors, aligning all three modalities in one contrastive space. The entry belongs beside CLIP, SigLIP, and ImageBind as an encoder architecture, despite originating from decoder checkpoints.

Gemma 4: Open-Weight Native Multimodal Models

Gemma 4 is Google DeepMind's open-weight multimodal family spanning mobile-efficient, dense, sparse-MoE, and encoder-free designs, with configurable reasoning and deployment targets ranging from phones to servers.

arXiv Releases HuggingFace

Gemma Team, Google DeepMind

Released: 2026-03-31

Gemma 4: Open-Weight Native Multimodal Models architecture: Aspect-preserving image resizing, patch pooling, and soft-token production

Figure 2. Aspect-preserving image resizing, patch pooling, and soft-token production. Source paper, PDF p. 16. Figure notice.

ℹ️ More Information

Gemma 4 is a suite of related architectures rather than one uniform network. E2B and E4B target mobile and edge deployment; 31B is a dense server-class decoder; and 26B-A4B is a sparse Mixture-of-Experts model with roughly 4B active parameters. These branches accept images through improved visual processing, while the mobile models additionally support audio. The distinctive Gemma 4 12B Unified removes separate vision and audio encoders: raw image patches and audio features enter through direct linear projections, reducing duplicated encoder memory and latency. Context windows reach 128K or 256K depending on the variant.

Pretraining jointly develops language, code, reasoning, and multimodal understanding, followed by instruction tuning, safety alignment, and reasoning post-training. A configurable thinking mode lets the instruction-tuned models emit a reasoning trace before answering. The family also ships dedicated multi-token-prediction draft models for speculative decoding, while per-layer embeddings and shared key-value caching improve inference efficiency in applicable variants.

Google describes a large multilingual mixture containing web text, documents, code, mathematics, image-text pairs, OCR and document examples, audio, and video-derived supervision. The complete record-level training corpus is not released, so benchmark collections should not be mistaken for an exhaustive dataset list. Image support includes variable resolutions and aspect ratios; native audio and video coverage depends on the specific branch.

Penguin-VL: Efficient VLMs with LLM-Based Vision Encoders

Penguin-VL challenges the assumption that compact VLMs require contrastively pretrained vision backbones, adapting a text-only language model into a fine-grained visual encoder for efficient image and video understanding.

arXiv GitHub HuggingFace

Boqiang Zhang, Lei Ke, Ruihan Yang, Qi Gao, Tianyuan Qu, Rossell Chen, Dong Yu, Leoweiliang

Released: 2026-03-06

Penguin-VL: Efficient VLMs with LLM-Based Vision Encoders architecture: An LLM-initialized vision encoder with priority-aware video-token compression

Figure 3. An LLM-initialized vision encoder with priority-aware video-token compression. Source paper, PDF p. 4. Figure notice.

ℹ️ More Information

Penguin-VL builds its Penguin-Encoder by initializing a vision backbone from a text-only Qwen3-0.6B model. Causal attention is converted to bidirectional attention, and 2D rotary position embeddings permit variable-resolution visual tokenization. The encoder is connected to compact language decoders to form 2B- and 8B-class VLMs. For video, Temporal Redundancy-Aware compression assigns larger token budgets to informative key frames and compresses redundant intermediate frames, allowing longer sequences within a fixed visual-token budget.

Training begins with mixed-supervision encoder adaptation. A teacher vision encoder supplies amplitude, direction, and relational distillation targets while image reconstruction-style supervision stabilizes the LLM-to-vision conversion. A low-to-high-resolution curriculum then aligns the encoder for fine-grained perception, followed by unified image and video multimodal pretraining and instruction tuning. This recipe is designed to preserve OCR, spatial, temporal, and dense-captioning signals that category-oriented contrastive objectives can suppress.

The project releases Penguin-Recap-I, a reconstructed high-quality image training set, and Penguin-Recap-V, video annotations at dense timestamp, paragraph, and whole-video granularities. These resources complement broader image-text, document, OCR, mathematical reasoning, and video instruction mixtures described in the report. The authors' comparisons attribute the compact models' gains primarily to visual representation quality and token allocation rather than parameter scaling alone.

Phi-4-Reasoning-Vision: Compact Multimodal Reasoning

Phi-4-Reasoning-Vision-15B is a compact open-weight model that combines high-resolution perception with selectable direct and chain-of-thought modes, focusing on mathematical, scientific, document, and GUI reasoning.

arXiv GitHub HuggingFace

Jyoti Aneja, Michael Harrison, Neel Joshi, Tyler LaBonte, John Langford, Eduardo Salinas

Released: 2026-03-04

Phi-4-Reasoning-Vision: Compact Multimodal Reasoning architecture: SigLIP2 vision encoding, cross-modal projection, mid-fusion, and language reasoning

Figure 3. SigLIP2 vision encoding, cross-modal projection, mid-fusion, and language reasoning. Source paper, PDF p. 4. Figure notice.

ℹ️ More Information

Phi-4-Reasoning-Vision-15B uses a mid-fusion architecture built from the Phi-4-Reasoning language backbone and a SigLIP 2 vision encoder. Visual features are projected into the language embedding space and inserted into the decoder sequence. Dynamic-resolution processing produces up to 3,600 visual tokens, supporting fine-grained documents and interface screenshots. Bidirectional attention is restricted to tokens within each image, improving spatial interaction while avoiding the overfitting Microsoft observed with broader bidirectional attention.

The model is trained through supervised fine-tuning on a deliberately mixed curriculum of reasoning and non-reasoning examples. Explicit and modes allow one checkpoint to use extended reasoning for mathematics and science or answer directly for perception-heavy tasks such as captioning, OCR, detection, and grounding. Microsoft reports four days of training on 240 NVIDIA B200 GPUs, followed by safety SFT and automated red-team evaluation.

Training data primarily comprises heavily filtered and corrected open-source vision-language datasets, augmented with synthetic examples, internally produced domain data, and targeted acquisitions. Microsoft emphasizes filtering, error correction, and synthetic augmentation as larger contributors than raw scale, but does not publish a complete source-level dataset inventory. Public evaluations span scientific diagrams, charts, mathematical reasoning, OCR, and GUI grounding; those benchmarks are evaluation references, not a claimed exhaustive training list.

V-SONAR and V-LCM: Vision-Language Modeling in Concept Space

V-SONAR aligns image and video representations with SONAR's multilingual concept space, enabling text-trained concept models to process vision. V-LCM then instruction-tunes a latent-diffusion Large Concept Model directly over unified visual and linguistic embeddings.

arXiv OpenReview

Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk

Released: 2026-03-01

V-SONAR and V-LCM: Vision-Language Modeling in Concept Space architecture: Visual-semantic alignment and concept-space prediction

Figure 1. Visual-semantic alignment and concept-space prediction. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

V-SONAR starts with Meta's Perception Encoder for image and video features and learns a lightweight projector into the pre-existing SONAR embedding space. Positional embeddings and a temporal-attention layer preserve frame order, while attention pooling produces a modality-level representation compatible with SONAR text concepts. Because the target space is shared with OMNISONAR, the representation can be decoded into text across SONAR's multilingual coverage. A text-only Large Concept Model can also consume these aligned visual embeddings without vision-specific pretraining.

V-SONAR is trained with a coarse-to-fine post-hoc alignment curriculum rather than retraining the foundation encoders from scratch. The alignment uses paired visual-caption supervision and SONAR text embeddings, with staged optimization and asynchronous learning rates for the newly initialized projector and pretrained encoder. V-LCM then concatenates V-SONAR visual concepts with SONAR language concepts and retains LCM's two-tower contextualizer and denoiser architecture. It predicts the next continuous concept embedding with the same latent-diffusion objective used during text-only LCM pretraining.

Vision-language instruction tuning uses M3IT, which spans image and video inputs, eight task categories, and 80 languages. Retrieval and captioning experiments additionally use VATEX, DREAM-1K, PE-Video, and VideoXum. The resulting system separates perception, multilingual semantic alignment, and generative concept modeling instead of converting every modality into a single discrete token vocabulary.

Qwen3.5: Native Multimodal Hybrid-Attention Models

Qwen3.5 unifies text, image, and video understanding in one model family while combining hybrid linear attention with sparse expert routing. Qwen3.6 continues the same native-multimodal direction with stronger agentic coding and updated dense and MoE checkpoints.

Blog GitHub HuggingFace

Qwen Team

Released: 2026-02-16

Architecture figure: Qwen3.5 has no public technical paper containing an architecture figure.

ℹ️ More Information

Qwen3.5 is a natively multimodal decoder family rather than a separate text model with a later VL edition. Its flagship 397B-A17B checkpoint uses sparse Mixture-of-Experts routing, activating about 17B parameters per token, and combines Gated DeltaNet linear-attention layers with standard softmax-attention layers. Visual inputs are handled inside the same model interface as text, supporting image and video reasoning, tool use, coding, and long-context agent workflows. Smaller dense and MoE variants extend the design across deployment scales.

The training pipeline combines large-scale multimodal pretraining with supervised fine-tuning and reinforcement learning. Qwen describes a scalable asynchronous RL system spanning text, multimodal, and multi-turn interaction, enabling a single checkpoint to switch between thinking and non-thinking behavior. The data mixture covers natural language, code, mathematics, images, video, documents, GUI interactions, tool calls, and multi-turn agent trajectories; the full source-level corpus is not publicly enumerated.

Qwen3.6 is best treated as a continuation of this architecture rather than a separate VLM family. Its open dense and sparse checkpoints retain native multimodality and unified thinking and non-thinking modes while prioritizing stability and agentic coding. Folding these releases into the Qwen3.5 entry keeps the catalog focused on architectural changes rather than model-version churn.

Youtu-VL: Unified Autoregressive Supervision for Dense Vision

Youtu-VL extends the language vocabulary with learned visual codes, making visual tokens prediction targets so one autoregressive model can emit segmentation, depth, pose, and detection without task-specific heads.

arXiv GitHub HuggingFace

Tencent Youtu Lab

Released: 2026-01-27

Youtu-VL: Unified Autoregressive Supervision for Dense Vision architecture: Unified visual-text autoregressive supervision and dense-output decoding

Figure 3. Unified visual-text autoregressive supervision and dense-output decoding. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

Vision-Language Unified Autoregressive Supervision learns a visual codebook whose symbols expand the text vocabulary into one multimodal vocabulary. Visual tokens are not passive conditioning embeddings: the decoder predicts them alongside text through joint visual-token and text reconstruction.

Because dense outputs are serialized in the same vocabulary, a standard autoregressive VLM can produce segmentation, monocular depth, human pose, and object detections without separate heads. This turns dense perception into native generation while retaining ordinary visual question answering and instruction following.

Kimi K2.5 and K2.6: Native Multimodal Agentic MoE

Kimi K2.5 is a trillion-parameter native multimodal MoE for reasoning, coding, and computer use; K2.6 keeps the topology while extending context and agentic post-training.

arXiv HuggingFace Release

Kimi Team, Moonshot AI

Released: 2026-01-27

Kimi K2.5 and K2.6: Native Multimodal Agentic MoE architecture: Agentic reinforcement-learning environments, rollout management, and training services

Figure 10. Agentic reinforcement-learning environments, rollout management, and training services. Source paper, PDF p. 23. Figure notice.

ℹ️ More Information

Kimi K2.5 connects MoonViT to a 1T-total, 32B-active decoder with 61 layers, Multi-head Latent Attention, and 384 routed experts. Eight routed experts plus shared capacity activate per token. Native multimodal training supports images, video, reasoning, code, tools, and GUI interaction within the same model.

Released April 20, Kimi K2.6 retains the K2.5 architecture while extending context to 256K and strengthening coding, research, and multimodal-agent post-training. It is folded here rather than receiving a duplicate timeline row; Kimi K3 is separate because its attention, residual, sparsity, scale, and vision stack change materially.

Step3-VL-10B: Language-Aligned Perception with 16× Token Compression

Step3-VL-10B combines a 1.8B language-aligned perception encoder with a Qwen3-8B decoder and aggressively compresses high-resolution visual features before language reasoning.

arXiv Project

StepFun

Released: 2026-01-14

Architecture figure: The Step3-VL-10B report contains performance and RL figures, but no architecture diagram.

ℹ️ More Information

Step3-VL-10B is architecturally distinct from the much larger 2025 Step3 sparse-MoE system. Its 1.8B perception encoder is aligned to language during pretraining, then a projector with two stride-2 stages reduces the spatial sequence by 16× before feeding a Qwen3-8B decoder.

Images combine one 728×728 global view with 504×504 local crops, treated as independent batch items. Explicit newline tokens mark patch rows while standard one-dimensional RoPE is retained. The PaCoRe parallel proposer-and-synthesizer reasoning method is a post-training and inference addition rather than the defining visual architecture.

ERNIE 5.0: Unified Autoregressive Omnimodal Mixture of Experts

ERNIE 5.0 unifies text, image, video, and audio understanding and generation in a 2.4T ultra-sparse MoE trained through modality-specific forms of grouped next-token prediction.

arXiv Architecture

Baidu ERNIE Team

Released: 2025-11-13

ERNIE 5.0: Unified Autoregressive Omnimodal Mixture of Experts architecture: Unified image understanding, image generation, and video generation objectives

Figure 2. Unified image understanding, image generation, and video generation objectives. Source paper, PDF p. 6. Figure notice.

ℹ️ More Information

ERNIE 5.0 maps language, images, video, and audio into one autoregressive framework using Next-Group-of-Tokens Prediction. Text uses next-token and multi-token prediction; visual generation predicts the next frame and scale; audio generation predicts codec tokens depth-wise. Modality-agnostic expert routing activates less than 3% of the 2.4T-parameter network.

Elastic depth, width, and top-k sparsity train a super-network from which smaller deployment configurations can be extracted. Baidu first unveiled ERNIE 5.0 at Baidu World on November 13, 2025, formally released it in January 2026, and published the technical report on February 4; the timeline therefore uses the earliest documented public disclosure.

DeepSeek-OCR: Visual Context Compression through DeepEncoder

DeepSeek-OCR treats document vision as a context-compression mechanism, representing thousands of text tokens with a much smaller sequence of visual tokens before decoding them with a sparse language model.

arXiv GitHub HuggingFace

DeepSeek-AI

Released: 2025-10-20

DeepSeek-OCR: Visual Context Compression through DeepEncoder architecture: A SAM-CLIP DeepEncoder connected to a sparse language decoder

Figure 3. A SAM-CLIP DeepEncoder connected to a sparse language decoder. Source paper, PDF p. 5. Figure notice.

ℹ️ More Information

DeepSeek-OCR consists of DeepEncoder and a DeepSeek-3B-MoE decoder with roughly 570M active parameters. DeepEncoder combines a SAM-derived local-perception pathway, a CLIP-derived semantic pathway, convolutional compression, and windowed and global attention to turn high-resolution pages into as few as 64 to 800 visual tokens. Multiple resolution modes trade recognition fidelity for compression.

Training first develops the encoder's visual-text representation and then jointly optimizes document decoding with the MoE language model. Data covers multilingual text pages, natural images, formulas, tables, charts, diagrams, and structured document conversion, supplemented with synthetic renderings. The work frames OCR as an experiment in optical compression for future long-context memory, rather than only a document benchmark model.

DeepSeek-OCR 2, released January 27, 2026, replaces the fixed raster-scan assumption with DeepEncoder V2 and Visual Causal Flow. The encoder dynamically reorders visual tokens according to document semantics before language decoding, exploring whether cascaded one-dimensional causal reasoning can better represent complex two-dimensional layouts. Paper · Repository

PaddleOCR-VL: Ultra-Compact Multilingual Document Parsing

PaddleOCR-VL combines a NaViT-style dynamic-resolution encoder with ERNIE-4.5-0.3B to parse multilingual documents using only 0.9B parameters.

arXiv GitHub HuggingFace

PaddleOCR Team, Baidu

Released: 2025-10-16

PaddleOCR-VL: Ultra-Compact Multilingual Document Parsing architecture: Document layout analysis, compact VLM inference, instructions, and structured output

Figure 2. Document layout analysis, compact VLM inference, instructions, and structured output. Source paper, PDF p. 5. Figure notice.

ℹ️ More Information

PaddleOCR-VL-0.9B connects a NaViT-style dynamic-resolution visual encoder to the compact ERNIE-4.5-0.3B language model. Variable-resolution packing preserves page detail and aspect ratio without forcing every document through a fixed square canvas. The decoder emits text and structured representations for paragraphs, tables, formulas, charts, and other document elements across 109 languages.

Its curriculum combines element-level recognition with page-level document parsing, followed by instruction tuning on complex layouts and structured outputs. Training data includes multilingual OCR, synthetic and scanned documents, tables, mathematical expressions, charts, reading-order annotations, and layout-rich pages. PaddleOCR-VL-1.5 (January 2026) adds robust physical-document distortions, seal recognition, and text spotting; 1.6 (June 2026) adds region-aware data optimization and progressive reinforcement-learning post-training without changing the core architecture.

Qwen3-VL: DeepStack Vision-Language Models

Qwen3-VL expands the Qwen vision-language family with dense and sparse-MoE checkpoints, native long multimodal context, and stronger spatial, video, OCR, visual-agent, and visual-coding capabilities.

arXiv GitHub HuggingFace

Qwen Team

Released: 2025-09-22

Qwen3-VL: DeepStack Vision-Language Models architecture: Vision encoding, DeepStack injection, and dense or mixture-of-experts decoding

Figure 1. Vision encoding, DeepStack injection, and dense or mixture-of-experts decoding. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

Qwen3-VL combines a native-dynamic-resolution Vision Transformer with Qwen3 language backbones in dense 2B, 4B, 8B, and 32B configurations and sparse 30B-A3B and 235B-A22B Mixture-of-Experts configurations. Its DeepStack mechanism injects features from multiple ViT depths into corresponding language-model layers instead of relying only on the visual encoder's final layer. Interleaved-MRoPE distributes temporal, height, and width position information across rotary-embedding frequencies, while explicit text-timestamp alignment improves event localization in video. The family supports interleaved text, images, and video in a native 256K-token context, with documented extrapolation for longer contexts.

Training proceeds through large-scale multimodal pretraining followed by supervised instruction tuning and reasoning-oriented post-training. The released Instruct checkpoints target direct response and agent interaction, whereas Thinking checkpoints generate extended reasoning for visual mathematics, spatial analysis, and other difficult multimodal tasks. The curriculum emphasizes recognition, multilingual OCR, document structure, 2D and 3D grounding, long-video understanding, GUI operation, and code generation from visual inputs.

The technical report describes a broad mixture of text, image-text, document, OCR, grounding, chart, multi-image, video, and agent-interaction data. Qwen does not publish a reproducible itemized list of every pretraining source, so individual benchmark datasets should not be presented as the complete training corpus.

Step3: Model-System Co-Design for Cost-Effective Multimodal Intelligence

Step3 is a 321B-total, 38B-active multimodal MoE whose architecture and serving system are co-designed to reduce the communication cost of sparse expert decoding.

arXiv GitHub Website

StepFun

Released: 2025-07-25

Step3: Model-System Co-Design for Cost-Effective Multimodal Intelligence architecture: Attention-FFN disaggregation across attention and expert instances

Figure 6. Attention-FFN disaggregation across attention and expert instances. Source paper, PDF p. 11. Figure notice.

ℹ️ More Information

Step3 uses a 321B-parameter sparse-MoE decoder with approximately 38B parameters active per token. The language stack contains dense lower layers followed by routed expert layers, while its vision pathway supports images and video for multimodal reasoning. The model is designed together with an Attention-FFN Disaggregation serving architecture, separating attention and expert computation so hardware placement and communication topology match their different workloads.

Training combines large-scale text and multimodal pretraining with instruction tuning and reasoning-oriented reinforcement learning. The data mixture spans language, code, mathematics, image-text, OCR, documents, charts, multi-image, video, and multimodal reasoning. The system report discloses architecture and deployment mechanisms more clearly than the source-level training corpus, which remains categorically described.

GLM-4.1V-Thinking: General-Purpose Multimodal Reasoning through Curriculum-Sampled RL

GLM-4.1V-9B-Thinking combines a native-resolution AIMv2-based vision stack with long-chain-of-thought post-training and Reinforcement Learning with Curriculum Sampling to strengthen reasoning across visual, document, video, grounding, and agent tasks.

arXiv GitHub HuggingFace

GLM-V Team, Zhipu AI and Tsinghua University

Released: 2025-07-01

GLM-4.1V-Thinking: General-Purpose Multimodal Reasoning through Curriculum-Sampled RL architecture: Native-resolution vision encoding, projection, decoding, and timestamped video tokens

Figure 2. Native-resolution vision encoding, projection, decoding, and timestamped video tokens. Source paper, PDF p. 4. Figure notice.

ℹ️ More Information

GLM-4.1V-9B-Thinking joins an AIMv2-Huge-initialized vision encoder, an MLP projector, and the GLM-4-9B-0414 language decoder. The vision tower replaces two-dimensional patch convolutions with 3D convolutions, temporally downsampling video while duplicating single images for consistent processing. Interpolated absolute position embeddings and 2D RoPE support native resolutions and extreme aspect ratios; the language decoder applies 3D RoPE to multimodal tokens. Explicit timestamp tokens after video frames supply temporal position and distance cues.

Training begins with broad multimodal pretraining, followed by continual training on video, higher-resolution imagery, and sequences up to 32K tokens. Supervised fine-tuning then teaches a standardized long-reasoning format with separate and spans. The final stage introduces Reinforcement Learning with Curriculum Sampling (RLCS). Rather than sampling a fixed task mixture, RLCS estimates changing model competence and prioritizes informative problems of suitable difficulty. Domain-specific rule, model, and hybrid rewards cover both verifiable and open-ended tasks.

Pretraining draws on curated caption pairs, knowledge-rich web and academic-book interleaving, 220M OCR images, natural-image and GUI grounding, instructional video, and text replay. SFT and RL span STEM reasoning, charts, documents, video, grounding, coding, GUI agents, and general visual instruction following. The later GLM-4.5V and GLM-4.6V families scale and extend this framework; they are successors, not separate architectures in this catalog.

ERNIE 4.5-VL: Heterogeneous Modality Mixture-of-Experts

ERNIE 4.5-VL combines shared and modality-specific experts with isolated routing and balancing losses so visual learning can reinforce rather than degrade language capability.

Website GitHub HuggingFace

Baidu ERNIE Team

Released: 2025-06-30

Architecture figure: ERNIE 4.5-VL has no public paper containing an extractable architecture figure.

ℹ️ More Information

ERNIE 4.5-VL includes 424B-total/47B-active and 28B-total/3B-active multimodal MoE models, plus dense members of the broader ERNIE 4.5 family. Its heterogeneous modality MoE shares some experts across text and visual tokens while reserving other experts for individual modalities. Modality-isolated routing, router orthogonality loss, and multimodal token-balancing loss reduce competition between modalities.

The models are jointly pretrained on text, images, and video, then post-trained in both direct-response and thinking modes. PaddlePaddle infrastructure uses heterogeneous hybrid parallelism, hierarchical load balancing, FP8 mixed precision, and fine-grained recomputation to train and serve the sparse models efficiently. Baidu describes broad text, visual-knowledge, OCR, document, chart, video, and reasoning mixtures but does not provide a complete dataset manifest.

MiMo-VL: Multimodal Pretraining with Mixed On-Policy Reinforcement Learning

MiMo-VL combines a four-stage, 2.4T-token multimodal curriculum with mixed on-policy reinforcement learning across perception, reasoning, and GUI grounding.

arXiv GitHub HuggingFace

Xiaomi MiMo Team

Released: 2025-06-04

MiMo-VL: Multimodal Pretraining with Mixed On-Policy Reinforcement Learning architecture: Native-resolution vision encoding, projection, and language decoding

Figure 2. Native-resolution vision encoding, projection, and language decoding. Source paper, PDF p. 6. Figure notice.

ℹ️ More Information

MiMo-VL-7B joins a high-resolution vision encoder and projector to a 7B language decoder and is released as supervised and reinforcement-learned checkpoints. Its primary contribution is the training system rather than an exotic connector: long-chain-of-thought examples are introduced during pretraining, not only during post-training, so multimodal reasoning is developed alongside perception and language.

Four pretraining stages consume approximately 2.4T tokens and progressively align vision, general multimodal knowledge, high-quality reasoning, and long-context capabilities. Mixed On-Policy Reinforcement Learning then trains one checkpoint across general QA, mathematics, grounding, and GUI tasks using domain-appropriate rule and model rewards. The mixture includes text replay, image-text, OCR, document, chart, video, spatial-grounding, GUI, and reasoning data; a complete source-level inventory is not published.

BAGEL: A Mixture-of-Transformer-Experts for Unified Understanding and Generation

BAGEL unifies visual understanding, autoregressive language modeling, and rectified-flow image generation in a decoder-only Mixture-of-Transformer-Experts whose modality-specific parameters communicate through shared self-attention.

arXiv GitHub HuggingFace

Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan

Released: 2025-05-20

BAGEL: A Mixture-of-Transformer-Experts for Unified Understanding and Generation architecture: Shared self-attention with understanding and generation Transformer experts

Figure 2. Shared self-attention with understanding and generation Transformer experts. Source paper, PDF p. 4. Figure notice.

ℹ️ More Information

BAGEL is a unified decoder-only model with 14B total and 7B active parameters. Its Mixture-of-Transformer-Experts architecture duplicates the Qwen2.5-initialized transformer into understanding and generation experts: text and SigLIP 2 ViT tokens use the understanding parameters, while FLUX-VAE latent tokens use the generation parameters. Both experts operate on one interleaved sequence through shared self-attention at every layer, avoiding a small connector bottleneck between comprehension and synthesis. Text is learned with next-token prediction; images are generated with rectified flow.

A generalized causal-attention scheme lets later text or images attend to earlier clean visual representations while preventing access to noised generation targets. Visual understanding uses a native-aspect-ratio-adapted SigLIP2-So400m encoder and MLP connector; generation uses a frozen FLUX VAE. Training progresses through alignment, large-scale pretraining, continued training, and supervised fine-tuning, jointly optimizing cross-entropy and flow-matching objectives.

BAGEL is pretrained on trillions of tokens assembled from text-only, paired image-text, and interleaved multimodal sources. The reported inventory includes 400M text samples, 500M understanding pairs, 1.6B generation pairs, 100M interleaved understanding samples, 45M video-derived sequences, and 20M web documents. OCR, charts, grounding, editing sets, and 500K reasoning-augmented generation and manipulation examples support document reading, spatial control, image editing, and long-context multimodal reasoning.

Seed1.5-VL: Sparse-MoE Multimodal Understanding and Agentic Reasoning

Seed1.5-VL couples a 532M-parameter vision encoder with a 20B-active sparse-MoE decoder for image, video, 2D and 3D grounding, GUI interaction, and long-chain-of-thought reasoning.

arXiv GitHub

ByteDance Seed Team

Released: 2025-05-11

Seed1.5-VL: Sparse-MoE Multimodal Understanding and Agentic Reasoning architecture: Native-resolution vision encoding, adaptation, sparse MoE decoding, and timestamped video

Figure 1. Native-resolution vision encoding, adaptation, sparse MoE decoding, and timestamped video. Source paper, PDF p. 5. Figure notice.

ℹ️ More Information

Seed1.5-VL uses a compact vision tower and a large sparse-MoE language model, maintaining 20B active decoder parameters while scaling total capacity substantially higher. Its visual pathway supports native-aspect-ratio images, multiple images, and video, with explicit support for spatial grounding and GUI control. The same model can provide short direct answers or extended multimodal reasoning.

Training progresses through vision-language pretraining, instruction tuning, long-chain-of-thought cold start, and reinforcement learning. The report emphasizes careful balancing of general text, visual knowledge, OCR and documents, charts, grounding, video, 3D understanding, GUI trajectories, games, and verifiable multimodal-reasoning problems. Source-level corpus provenance and all mixture proportions are not publicly reproducible. The arXiv v1 date is used because an earlier generally accessible official launch date is not unambiguously documented.

InternVL3 and InternVL3.5: Native Multimodal Pretraining and Adaptive Resolution

InternVL3 moves the InternVL family to native multimodal pretraining, while InternVL3.5 adds coarse-to-fine reinforcement learning and dynamically routed visual resolution.

arXiv arXiv GitHub HuggingFace

InternVL Team, OpenGVLab

Released: 2025-04-11

Architecture figure: The InternVL3 paper contains evaluation figures, but no definitive model architecture diagram.

ℹ️ More Information

InternVL3 retains the InternViT-MLP-LLM structure and dynamic image tiling of earlier InternVL releases, but jointly learns language and multimodal knowledge during native pretraining instead of treating vision alignment primarily as a later adaptation stage. The family spans dense and sparse-MoE language backbones. Its training recipe combines multimodal pretraining, supervised fine-tuning, Mixed Preference Optimization, and test-time scaling for visual reasoning.

Released on August 26, 2025, InternVL3.5 is part of the same evolving family. It introduces Cascade RL, using offline RL for stable coarse alignment followed by online RL for refinement. A Visual Resolution Router selects visual-token resolution according to input complexity, while Decoupled Vision-Language Deployment places the vision encoder and language model on different devices to balance inference load. Training covers image, multi-image, document, chart, video, GUI, grounding, multilingual, and reasoning data; full pretraining provenance is not enumerated.

Kimi-VL: Native-Resolution Vision with a Sparse MoE Decoder

Kimi-VL couples the native-resolution MoonViT encoder to a sparse Mixture-of-Experts language decoder, activating 2.8B of 16B decoder parameters while supporting images, video, agent interaction, and 128K-token contexts.

arXiv GitHub HuggingFace

Kimi Team

Released: 2025-04-10

Kimi-VL: Native-Resolution Vision with a Sparse MoE Decoder architecture: MoonViT, multimodal projection, and a sparse mixture-of-experts decoder

Figure 3. MoonViT, multimodal projection, and a sparse mixture-of-experts decoder. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

Kimi-VL-A3B consists of MoonViT, a two-layer MLP projector, and a Moonlight-derived MoE language model with 16B total and 2.8B activated decoder parameters. MoonViT packs variable-sized patch sequences without tiling images into separately processed crops. It combines interpolated SigLIP-SO400M position embeddings with two-dimensional RoPE, preserving pretrained visual knowledge while representing fine spatial positions. A 2×2 pixel-shuffle operation compresses its output before projection into the language model.

After the language checkpoint's 5.2T-token text-only phase, Kimi-VL undergoes standalone vision training, vision-language alignment, joint pretraining, a high-quality cooldown, and long-context activation. These stages consume 4.4T additional tokens, including 2T for MoonViT training, 0.1T for alignment, and 2.3T across joint stages. Context is progressively extended from 8K to 128K. The instruct model receives joint supervised fine-tuning. Kimi-VL-Thinking adds long-chain-of-thought SFT and reinforcement learning; the later Thinking-2506 checkpoint is folded into this family.

The multimodal mixture is organized into six categories: caption, interleaved image-text, OCR, knowledge, video, and agent data. MoonViT learns from alt text, synthetic captions, grounding boxes, and OCR targets using both SigLIP-style contrastive and captioning losses. Long-video, long-document, academic visual QA, GUI-grounding, and text replay data provide long-context and agent capabilities while preserving language performance.

Llama 4 Scout and Maverick: Native Multimodal Mixture-of-Experts Models

Llama 4 introduces early-fusion multimodality to the Llama family through sparse Mixture-of-Experts decoders with 17B active parameters and contexts extending from one million to ten million tokens.

Model Card HuggingFace

Meta

Released: 2025-04-05

Architecture figure: The official Llama 4 model card describes the architecture in text and tables only.

ℹ️ More Information

Llama 4 Scout and Llama 4 Maverick are natively multimodal, autoregressive MoE models rather than text decoders retrofitted only during instruction tuning. Scout has 109B total parameters across 16 experts and Maverick approximately 400B across 128 experts; both activate about 17B parameters per token. Early fusion places image and multilingual text tokens in a shared model sequence. Scout supports a 10M-token context, whereas Maverick supports 1M tokens.

Scout was trained on roughly 40T tokens and Maverick on roughly 22T. Meta describes a mixture of public data, licensed data, information from its products and services, and interactions with Meta AI, without releasing a complete corpus manifest. Post-training combines supervised fine-tuning, online reinforcement learning, and direct preference optimization. Maverick was co-distilled from the unreleased Llama 4 Behemoth teacher; Scout also uses distillation during training.

Qwen2.5-Omni: Streaming Multimodal Perception and Speech Generation

Qwen2.5-Omni understands text, images, audio, and video while generating text and natural speech through a streaming Thinker-Talker architecture.

arXiv GitHub HuggingFace

Qwen Team

Released: 2025-03-26

Qwen2.5-Omni: Streaming Multimodal Perception and Speech Generation architecture: Thinker-Talker architecture for multimodal perception and streaming speech

Figure 2. Thinker-Talker architecture for multimodal perception and streaming speech. Source paper, PDF p. 3. Figure notice.

ℹ️ More Information

Qwen2.5-Omni is an end-to-end any-input model whose block-wise audio and vision encoders accept streaming speech, images, and video alongside text. Time-aligned Multimodal RoPE (TMRoPE) interleaves audio and video representations on a synchronized timeline. Its Thinker-Talker design separates semantic reasoning from speech realization: the Thinker produces text and hidden representations, while a dual-track autoregressive Talker conditions on those states to generate speech tokens concurrently.

A sliding-window diffusion transformer converts audio tokens into waveform output without waiting for a complete response, reducing first-packet latency. Training jointly aligns text, vision, video, audio, and speech-generation objectives, followed by multimodal instruction tuning. The report describes caption, OCR, video, audio-transcription, speech, and general text mixtures but does not publish a reproducible source-level inventory.

Gemma 3: Long-Context Multimodality with Efficient Interleaved Attention

Gemma 3 combines a frozen SigLIP vision encoder with a decoder-only language model whose five-local-to-one-global attention pattern reduces long-context KV-cache cost while retaining 128K-token multimodal context.

arXiv GitHub HuggingFace

Gemma Team

Released: 2025-03-12

Architecture figure: The Gemma 3 report contains examples and analysis charts, but no architecture diagram.

ℹ️ More Information

Gemma 3 is a family of 1B, 4B, 12B, and 27B open-weight models; the 4B, 12B, and 27B variants a