← Open Source
Yuliang-Liu

AWESOME-OCR-LLM

OCR in the Era of Large Language Models

ListsPaper collectionsDataset collections
Open on GitHub
Momentum
+0stars in 24 hours0.0%
725
Stars
54
Forks
+1
This week
13
Contributors
Created 2026-02-10 · Updated 2026-10-05 · #14226 today
Top developers
README

🔭 Daily Papers (2026-10-05)

Benchmarks and Evaluation

Venue Name Primary affiliation Title GitHub Date
Paper FoveDoc-Bench HIT Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models GitHub Stars Oct. 2026

FoveDoc-Bench — Shows that retrieved document pages are often found but never really read by vision-language models, and that supplying extracted text narrows this retrieval-reading gap for textual evidence while doing nothing for charts and figures.

📖 Contents

Overview

A curated, continuously updated reading list of OCR in the era of large language models, covering document parsing and understanding, visual text generation, benchmarks, challenges, and future perspectives, with a focus on research around the past five years (2021–now).

Scope. This list tracks OCR in the LLM era: work that applies large vision-language or multimodal models to text-rich images and documents (parsing, understanding, benchmarks, and specialized text tasks). It is not a general document-AI list, a generic MLLM list, or a classical OCR-1.0 list; such work appears only when it directly bears on text-rich visual understanding.

A note on evaluation. Most recent systems are released as technical reports with self-reported numbers, private test sets, and inconsistent protocols, so cross-paper scores are rarely comparable in a rigorous sense. We list results as reported and, where known, indicate the evaluation basis. The field still lacks a unified, contamination-resistant, reproducible benchmark, and we see building one as a prerequisite for trustworthy leaderboard claims.

🎉 News

  • [2026-2-11] 🔥 We release an open-source resource to help the community easily track recent OCR research!

Contributing. PRs welcome. One row per model, newest first; please include venue/date, affiliation, and a code or model link.

🔍 Emerging Trends

Emerging Trends (2023–2026)

  • End-to-end VLM-based parsing replaces modular OCR pipelines.
  • Reinforcement learning for layout and reading order modeling.
  • OCR-free document understanding models.
  • Scaling down: compact document VLMs under 1B parameters.
  • Long-doc OCR is a scaling problem of tokens and consistency, not just context length.
  • Structure (layout + logic) is the new accuracy.
  • Generative OCR shifts the core risk from “misrecognition” to “hallucination”.
  • Benchmarks are moving toward executable evaluation.
  • Document agents and autonomous reasoning over PDFs.
  • Document Agents need recoverability, not one-shot perfection.

📄 Document Parsing

Document parsing focuses on converting visually complex documents into structured, machine-readable representations. In the LLM era, parsing is no longer a pipeline of isolated modules, but increasingly unified within end-to-end VLM architectures.

Venue Name Primary affiliation Title GitHub Date
Paper PolyOCR-Venus Ant Group PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence - Sep. 2026
Paper Xiaomi-OCR-0 Xiaomi Xiaomi-OCR-0 Technical Report HuggingFace Sep. 2026
- Logics-Parsing-V3 Alibaba Logics-Parsing-V3: Structure-Aware Recurrent Parsing for Long Documents HuggingFace GitHub Stars Sep. 2026
Paper GravityOCR Trillion Labs Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding HuggingFace GitHub Stars Sep. 2026
PrismAlign Huawei PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR - Sep. 2026
Paper WeVisDoc Tencent WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing HuggingFace GitHub Stars Sep. 2026
Paper Jina-OCR-v1 Jina AI Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards HuggingFace Sep. 2026
Paper OCR-EDR Tencent OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement - Sep. 2026
Paper Wayu-Paxa-OCR-Zero Wayu Research How Far Can Synthetic Data Take Thai OCR? - Sep. 2026
Paper SCVER Fudan University State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models - Sep. 2026
Paper FinixDoc Ant Group FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks HuggingFace Aug. 2026
SmolDocling-KV-Link IBM Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images - Aug. 2026
Paper ArmorOCR Ant Group ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation GitHub Stars Aug. 2026
Paper NaviDC-OCR China Telecom AI NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents - Aug. 2026
Paper TongGuOCR SCUT TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents GitHub Stars Aug. 2026
Paper PaDoc Tsinghua University PaDoc: Layout-Grounded Parallel Decoding for Document Parsing GitHub Stars Aug. 2026
Logographic Pretraining QMUL Logographic Character Visual Pretraining via Semantic-based Contrastive Learning - Aug. 2026
SPIRAL SEU Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression GitHub Stars Aug. 2026
DocPO Tencent DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards - Aug. 2026
Paper DrawAI BUPT DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable HuggingFace GitHub Stars Aug. 2026
Paper LayoutLite Yuanli Technology LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR GitHub Stars Jul. 2026
Paper HPD-Parsing Baidu HPD-Parsing: Hierarchical Parallel Document Parsing HuggingFace GitHub Stars Jul. 2026
Paper OvisOCR2 Alibaba OvisOCR2 Technical Report HuggingFace GitHub Stars Jul. 2026
Paper DocOCR-Eval University of Melbourne DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth - Jul. 2026
Paper MonkeyOCRv2 HUST&Kingsoft MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Hugging Face GitHub Stars Jul. 2026
Paper Infinity-Parser2 INF Team Infinity-Parser2 Technical Report Hugging FaceGitHub Stars Jul. 2026
Paper HunyuanOCR-1.5 Tencent HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better HuggingFace Stars GitHub Stars Jul. 2026
Paper SAYRE Alibaba Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis - Jul. 2026
Paper P-MTP Baidu P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling - Jun. 2026
Paper RT-DocLayout Baidu RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild - Jun. 2026
Paper Unlimited-OCR Baidu Unlimited OCR Works GitHub Stars Jun. 2026
Paper Beaver Microsoft Research Building Agent Harnesses for Scientific Curation from Multimodal Sources - Jun. 2026
Paper Agents-K1 Shanghai AI Laboratory Agents-K1: Towards Agent-native Knowledge Orchestration HuggingFace Stars Jun. 2026
Paper PaddleOCR-VL-1.6 Baidu PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training GitHub Stars Jun. 2026
Paper PP-OCRv6 Baidu PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks GitHub Stars Jun. 2026
StrucTab IIE, CAS StrucTab: A Structured Optimization Framework for Table Parsing GitHub Stars Jun. 2026
Paper ExChart Zhejiang University Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework - Jun. 2026
Paper MinerU-Popo Shanghai AI Laboratory & OpenDataLab MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing GitHub Stars May. 2026
RTPrune DeepSeek RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference GitHub Stars May. 2026
Paper ABot-OCR Alibaba ABot-OCR Technical Report GitHub Stars May. 2026
Paper BabelDOC funstory.ai BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation GitHub Stars May. 2026
Paper Consensus Entropy Fudan University Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR GitHub Stars May. 2026
Paper FastOCR Tsinghua University & JD FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing - May. 2026
Paper MinerU2.5-Pro Shanghai AI Laboratory MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale GitHub Stars Apr. 2026
Paper PixelPrune OPPO PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding GitHub Stars Apr. 2026
Paper TexOCR Yale University & Zhejiang University TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction GitHub Stars Apr. 2026
Paper Falcon OCR Falcon Vision Team, TII Falcon Perception GitHub Stars Mar. 2026
Paper MinerU-Diffusion Shanghai AI Laboratory MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding GitHub Stars Mar. 2026
Paper Qianfan-OCR Baidu Qianfan-OCR: A Unified End-to-End Model for Document Intelligence GitHub Stars Mar. 2026
Paper dots.mocr HUST Multimodal OCR: Parse Anything from Documents GitHub Stars Mar. 2026

📄 See full list at Document-Parsing.md

📄 Document Understanding

Document understanding extends beyond structural parsing to semantic comprehension and reasoning over visually rich documents.

Venue Name Primary affiliation Title GitHub Date
Paper InterTab HKUST InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning - Sep. 2026
Paper D-RAC Yellow.ai Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion - Sep. 2026
Paper Chart-RVR University of Virginia Monitorable Chart Reasoning Agents via Verifiable Process Rewards HuggingFace GitHub Stars Sep. 2026
Paper ViSAR INSA Lyon ViSAR: Training-Free Adaptive-k Retrieval for Visual Document Question Answering - Sep. 2026
Paper DocIntent SCUT DocIntent: Answerability-Guided Agentic Restoration for Real-World Document Visual Question Answering - Sep. 2026
Doc-REFRAG Zhejiang University Doc-REFRAG: Rethinking Multimodal Document Retrieval-Augmented Generation GitHub Stars Sep. 2026
AWM University of Oslo AWM: Answerable Working Memory for Long-Document VQA Agents GitHub Stars Aug. 2026
Paper SAGE Fudan University SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding - Aug. 2026
Paper Q-Guide Amazon Question-Guided Evidence Acquisition for Multimodal Visual Question Answering - Aug. 2026
Paper DocClaw NTU DocClaw: A Unified Agentic System for Intelligent Document Processing GitHub Stars Aug. 2026
Paper ConceptFormer Northeastern University ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval GitHub Stars Aug. 2026
Hyper-M2RAG Hangzhou Dianzi University Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement GitHub Stars Aug. 2026
Paper Trident Emory University What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering - Aug. 2026
Paper D2-ScaleAgent Zhejiang University D2-ScaleAgent: Dual-Dimensional Scaling for Long Document Understanding - Aug. 2026
SEER UT Austin SEER: Long-Context Reasoning via Selective Visual-Text Compression GitHub Stars Aug. 2026
Paper HAM-RAG HKUST HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation GitHub Stars Aug. 2026
DRUF Shenzhen MSU-BIT University Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs GitHub Stars Aug. 2026
Paper DistilVDR Aalto University DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation GitHub Stars Aug. 2026
Paper InSight-doc HKUST InSight-doc: Agentic Visual Perception for Long-Document Understanding GitHub Stars Aug. 2026
Paper DocAtlas Wuhan University / Microsoft DocAtlas: Long-Document Understanding as Mutable-State Interaction - Aug. 2026
Paper DocMemo HIT, Shenzhen DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding GitHub Stars Aug. 2026
Paper ECF Beihang University Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models? GitHub Stars Aug. 2026
Paper ADOPD 2026 Georgia Tech Thinking with Anchors: Grounded and Efficient Document Reasoning HuggingFace GitHub Stars Aug. 2026
Paper VTS MBZUAI When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning - Aug. 2026
Paper Q-CueGraph The University of Tokyo Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning - Aug. 2026
Paper DocTrace Baidu DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning - Aug. 2026
Paper CURV William & Mary CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning - Aug. 2026
Paper RAGOCR Peking University RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation - Aug. 2026
Paper VaRS-Doc SJTU VaRS-Doc: Interpretation-Aware Variant Representations via Latent Self-Probing for Visual Document Retrieval GitHub Stars Aug. 2026
Paper ET-Prune SJTU ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs GitHub Stars Aug. 2026
Paper HierDoc USTC HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering - Aug. 2026
DualG-MRAG BUAA DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation - Jul. 2026
Paper MMLDSum-LLM OPPO MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware - Jul. 2026
Paper TAP-RAG Tianjin University TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering Anonymous Code Jul. 2026
Perception-RFT Quantiphi Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment - Jul. 2026
CMDR NTT CMDR: Contextual Multimodal Document Retrieval HuggingFace Stars GitHub Stars Jul. 2026
Paper HiEvi-RAG USTC Hierarchical Evidence-Driven Reasoning for Long Document Understanding - Jul. 2026
Paper MultAttnAttrib — MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering - Jul. 2026
Paper OracleAnalyser NUDT OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training - Jun. 2026
Paper DocArena Adobe Research DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents - Jun. 2026
ViTexQA Meituan ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering GitHub Stars HuggingFace Dataset Jun. 2026
Paper PreciseDoc Tsinghua University An LMM for Precisely Grounding Elements in Documents - Jun. 2026
LightSTAR SJTU LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement GitHub Stars Jun. 2026
Paper SciLens HKUST SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding - Jun. 2026
Paper SAFE-Cascade Walmart SAFE-Cascade: Cost-Adaptive Vision-Language Routing for Chart Question Answering - Jun. 2026
Paper UMG-RAG Purdue University Uncertainty-Aware Hybrid Retrieval for Long-Document RAG - Jun. 2026
Paper MAGE-RAG BIT MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA GitHub Stars Jun. 2026
Paper MINARD University of Maryland Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures - Jun. 2026
Paper MM-BizRAG JPMorgan Chase & Co. MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A - Jun. 2026
Paper KG4VD National Taiwan University Multimodal Graph RAG for Long-range Visually Rich Document Understanding GitHub Stars Jun. 2026

📄 See full list at Document-Understanding.md

📄 Visual Text Generation

Visual Text Generation focuses on generating or editing legible, visually harmonious, and semantically consistent text within images, serving as the creative inverse of OCR. In the LLM era, it is no longer a task reliant on specialized modules, but is emerging as a foundational skill for general-purpose generative models.

Venue Name Primary affiliation Title GitHub Date
Paper IDSpect Fudan University Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering - Sep. 2026
Paper DuetGen USTC Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation - Sep. 2026
LoGAN Netflix LoGAN: Multilingual Font Localization with Generative Agents - Sep. 2026
Paper GlyphAnchor Fudan University GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors - Sep. 2026
Paper TextRefine Kuaishou TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters - Aug. 2026
Paper PosterText Wuhan University PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster - Aug. 2026
Paper TransAnyText Wuhan University TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation - Aug. 2026
onoff GIST Bridging Online and Offline Handwriting via Differentiable Physical Rendering GitHub Stars Aug. 2026
Paper PosterMELD Tsinghua University PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs GitHub Stars Aug. 2026
InnoText SYSU InnoText: A Unified Model for Visual Text Generation and Editing - Jul. 2026
Paper Boogu-Image-0.1 Boogu Team,Huawei Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation HuggingFace GitHub Stars Jul. 2026
Paper SciForma Microsoft & Peking University SciForma: Structure-Faithful Generation of Scientific Diagrams GitHub Stars Jul. 2026
Paper VecFontLLM Fuzhou University & Peking University VecFontLLM: Anchor-Guided Direct Synthesis of Chinese Vector Fonts - Jul. 2026
Paper ArtChart Ant Group ArtChart: A Benchmark for Faithful Artistic Chart Generation with Integrated Text Rendering - Jul. 2026
Paper DataEvolver CSU DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation GitHub Stars Jul. 2026
Paper Qwen-Image-2.0-RL Alibaba Qwen-Image-2.0-RL Technical Report - Jun. 2026
UniTranslator IIE CAS UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation GitHub Stars Jun. 2026
Paper SteerVTE ByteDance & Peking University SteerVTE: Seamless Video Text Editing with Style and Glyph Control - Jun. 2026
Paper DiffMath SCUT & Huawei DiffMath: Symbol- and Graph-Aware Latent Diffusion Transformer for Handwritten Mathematical Expression Generation GitHub Stars Jun. 2026
Paper PhyDrawGen University of Dhaka PhyDrawGen: Physically Grounded Diagram Generation from Natural Language - Jun. 2026
Paper NIV Reichman University NIV: Neural Axis Variations for Variable Font Generation GitHub Stars Jun. 2026
Paper Qwen-Image-2.0 Alibaba Group Qwen-Image-2.0 Technical Report GitHub Stars May. 2026
Paper MangaFlow The University of Tokyo MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation - May. 2026
Paper Wan-Image Alibaba Group Wan-Image: Pushing the Boundaries of Generative Visual Intelligence - Apr. 2026
PosterIQ PolyU PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation GitHub Stars Mar. 2026
Paper EfficientPosterGen Tsinghua University EfficientPosterGen: Semantic-aware Efficient Poster Generation via Token Compression and Accurate Violation Detection GitHub Stars Mar. 2026
Paper TextFlow NJUST Towards Training-Free Scene Text Editing GitHub Stars Mar. 2026
Paper GlyphPrinter Fudan University GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering GitHub Stars Mar. 2026
Paper CTRL-S SJTU Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning - Mar. 2026
Paper LaDe Adobe Research LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition - Mar. 2026
Paper EchoGen USTC EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding - Mar. 2026
Paper WebVR StepFun WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics - Mar. 2026
Paper AutoFigure-Edit Westlake University AutoFigure-Edit: Generating Editable Scientific Illustration GitHub Stars Mar. 2026
Paper GlyphBanana SJTU & Xiaohongshu Inc. GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows GitHub Stars Mar. 2026
Paper InnoAds-Composer JD.com InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation - Mar. 2026
Paper Seeing is Improving USTC Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement GitHub Stars Mar. 2026
PosterOmni HKUST(GZ) PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback GitHub Stars Feb. 2026
TextPecker HUST TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering GitHub Stars Feb. 2026
Paper SSPT Tongji University Space Syntax-guided Post-training for Residential Floor Plan Generation - Feb. 2026
Paper ChatUMM Tsinghua University & Tencent Hunyuan ChatUMM: Robust Context Tracking for Conversational Interleaved Generation - Feb. 2026
PosterVerse SCUT PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography GitHub Stars Jan. 2026
- GPT-Image-1.5 OpenAI GPT-Image-1.5 --- Dec. 2025
- Gemini 3 Pro Image Google DeepMind Gemini 3 Pro Image (Nano Banana Pro) --- Nov. 2025
- Gemini 2.5 Flash Image Google DeepMind Gemini 2.5 Flash Image (Nano Banana) --- Oct. 2025
Paper Qwen-Image Qwen Team Qwen-Image Technical Report GitHub Stars Sep. 2025
Paper Seedream 4.0 ByteDance Seed Seedream 4.0: Toward Next-generation Multimodal Image Generation --- Aug. 2025
Paper Postergen Stony Brook University Postergen: Aesthetic-aware paper-to-poster generation via multi-agent llms GitHub Stars Aug. 2025
UniGlyph Tsinghua University UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis --- Jul. 2025
Paper X-Omni Tencent Hunyuan X X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again GitHub Stars Jul. 2025
Paper DreamPoster Intelligent Creation Lab, ByteDance DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design --- Jul. 2025

📄 See full list at Visual-Text-Generation.md

📄 Specialized Model

Beyond general document parsing and understanding, traditional OCR research remains essential for specialized visual-text structures, restoration, temporal analysis, and forensic reliability.

📄 Document Dewarping

Document dewarping restores photographed or scanned pages to a geometrically rectified and readable form.

Venue Name Primary affiliation Title GitHub Date
Paper BookNet HFUT BookNet: Dual-Page Book Image Rectification via Cross-Page Attention - Jan. 2026
Paper TADoc CAS TADoc: Robust Time-Aware Document Image Dewarping - Aug. 2025
Paper DocDewarpHV HIT Dual Dimensions Geometric Representation Learning Based Document Dewarping GitHub Stars Jul. 2025
DocMatcher FZI & KIT DocMatcher: Document Image Dewarping via Structural and Textual Line Matching GitHub Stars Mar. 2025
Paper - Shuya Branch of the Ivanovo State University Efficient Document Image Dewarping via Hybrid Deep Learning and Cubic Polynomial Geometry Restoration GitHub Stars Jan. 2025
DocScanner USTC DocScanner: Robust Document Image Rectification with Progressive Learning GitHub Stars Jan. 2025
DocRes SCUT DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks GitHub Stars Jun. 2024
DocNLC SCUT DocNLC: A Document Image Enhancement Framework with Normalized and Latent Contrastive Representation for Multiple Degradations GitHub Stars Feb. 2024
DocTr++ USTC Deep Unrestricted Document Image Rectification - 2024
LA-DocFlatten CAS Layout-aware Single-image Document Flattening GitHub Stars 2024
UVDoc ETH Zurich UVDoc: Neural Grid-based Document Unwarping GitHub Stars Oct. 2023
Foreground and Text-lines Aware Model HIT Shenzhen, China Foreground and text-lines aware document image rectification GitHub Stars Jun. 2023
DocMAE USTC & iFLYTEK DocMAE: Document Image Rectification via Self-supervised Representation Learning - 2023
Marior SCUT & IntSig Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the Wild GitHub Stars Oct. 2022

📄 Physical Structure Analysis

Physical structure analysis identifies document regions, layouts, and spatial relationships before semantic reasoning.

Venue Name Primary affiliation Title GitHub Date
Paper PorTEXTO NOVA School of Science and Technology PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction - Jun. 2026
Paper - Insiders Technologies GmbH Bounding Box Label Propagation for Re-Annotation of Document Layout Analysis Datasets - Jun. 2026
Paper IndustryBench-MIPU Multimodal and Industrial AI Team, Alibaba IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products GitHub Stars Jun. 2026
Paper ERN-Net Tamkang University ERN-Net : Evolving Reason Node-Net for Document Binarization - Jun. 2026
HiLEx IEM Kolkata HiLEx: Image-Based Hierarchical Layout Extraction from Question Papers GitHub Stars Mar. 2025
DCEM-ViT Sharda University, Greater Noida, India Devanagari character encoded mix-merge vision transformer for robust document layout analysis --- Mar. 2025
Efficient Additive Attention DLA Technical University of Kaiserslautern, Germany Efficient Additive Attention for Transformer-based Semi-supervised Document Layout Analysis --- Feb. 2025
DocSemi Technical University of Kaiserslautern, Germany DocSemi: Efficient Document Layout Analysis with Guided Queries --- Feb. 2025
FS-QCSNet China University of Mining and Technology Few-Shot Quaternion-valued Correlation Squeeze Network for Document Image Layout Segmentation --- Jan. 2025
LayoutDETR Salesforce Research LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer - Sep. 2024
LayoutLLM Alibaba Group LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding GitHub Stars Jun. 2024
RoDLA KIT & Univ. of Oxford RoDLA: Benchmarking the Robustness of Document Layout Analysis Models GitHub Stars Jun. 2024
DocLLM JPMorgan Chase & Co. DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding GitHub Stars May 2024
SemiDocSeg Computer Vision Center, Barcelona, Spain SemiDocSeg: Harnessing Semi-Supervised Learning for Document Layout Analysis --- Mar. 2024
VGT Alibaba Group Vision Grid Transformer for Document Layout Analysis GitHub Stars Oct. 2023
GeoLayoutLM Alibaba Group GeoLayoutLM: Geometric Pre-training for Visual Information Extraction GitHub Stars Jun. 2023
M6Doc SCUT M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis GitHub Stars Jun. 2023
HRDoc USTC & iFLYTEK HRDoc: Dataset and Baseline Method Toward Hierarchical Reconstruction of Document Structures GitHub Stars Feb. 2023
LayoutLMv3 Microsoft LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking GitHub Stars Oct. 2022

📄 Reading Order Prediction

Reading order prediction recovers the intended sequence of text and visual elements in complex page layouts.

Venue Name Primary affiliation Title GitHub Date
Paper Orli University of Konstanz End-to-End Text Line Detection and Ordering GitHub Stars Jun. 2026
[![Paper](https://img.shields.io/badge/paper-A42C25?style=for-the-badge&logo=arxiv&