🔭 Daily Papers (2026-10-05)
Benchmarks and Evaluation
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|---|---|---|---|---|
FoveDoc-Bench |
HIT | Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models | Oct. 2026 |
FoveDoc-Bench — Shows that retrieved document pages are often found but never really read by vision-language models, and that supplying extracted text narrows this retrieval-reading gap for textual evidence while doing nothing for charts and figures.
📖 Contents
- Overview
- 🎉 News
- 📖 Contents
- 🔭 Daily Papers
- 🔍 Emerging Trends
- 📄 Document Parsing
- 📄 Document Understanding
- 📄 Visual Text Generation
- 📄 Specialized Model
- 📄 Benchmarks and Evaluation
Overview
A curated, continuously updated reading list of OCR in the era of large language models, covering document parsing and understanding, visual text generation, benchmarks, challenges, and future perspectives, with a focus on research around the past five years (2021–now).
Scope. This list tracks OCR in the LLM era: work that applies large vision-language or multimodal models to text-rich images and documents (parsing, understanding, benchmarks, and specialized text tasks). It is not a general document-AI list, a generic MLLM list, or a classical OCR-1.0 list; such work appears only when it directly bears on text-rich visual understanding.
A note on evaluation. Most recent systems are released as technical reports with self-reported numbers, private test sets, and inconsistent protocols, so cross-paper scores are rarely comparable in a rigorous sense. We list results as reported and, where known, indicate the evaluation basis. The field still lacks a unified, contamination-resistant, reproducible benchmark, and we see building one as a prerequisite for trustworthy leaderboard claims.
🎉 News
- [2026-2-11] 🔥 We release an open-source resource to help the community easily track recent OCR research!
Contributing. PRs welcome. One row per model, newest first; please include venue/date, affiliation, and a code or model link.
🔍 Emerging Trends
Emerging Trends (2023–2026)
- End-to-end VLM-based parsing replaces modular OCR pipelines.
- Reinforcement learning for layout and reading order modeling.
- OCR-free document understanding models.
- Scaling down: compact document VLMs under 1B parameters.
- Long-doc OCR is a scaling problem of tokens and consistency, not just context length.
- Structure (layout + logic) is the new accuracy.
- Generative OCR shifts the core risk from “misrecognition” to “hallucination”.
- Benchmarks are moving toward executable evaluation.
- Document agents and autonomous reasoning over PDFs.
- Document Agents need recoverability, not one-shot perfection.
📄 Document Parsing
Document parsing focuses on converting visually complex documents into structured, machine-readable representations. In the LLM era, parsing is no longer a pipeline of isolated modules, but increasingly unified within end-to-end VLM architectures.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|---|---|---|---|---|
PolyOCR-Venus |
Ant Group | PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence | - | Sep. 2026 | |
Xiaomi-OCR-0 |
Xiaomi | Xiaomi-OCR-0 Technical Report | Sep. 2026 | ||
| - | Logics-Parsing-V3 |
Alibaba | Logics-Parsing-V3: Structure-Aware Recurrent Parsing for Long Documents | Sep. 2026 | |
GravityOCR |
Trillion Labs | Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding | Sep. 2026 | ||
![]() |
PrismAlign |
Huawei | PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR | - | Sep. 2026 |
WeVisDoc |
Tencent | WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing | Sep. 2026 | ||
Jina-OCR-v1 |
Jina AI | Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards | Sep. 2026 | ||
OCR-EDR |
Tencent | OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement | - | Sep. 2026 | |
Wayu-Paxa-OCR-Zero |
Wayu Research | How Far Can Synthetic Data Take Thai OCR? | - | Sep. 2026 | |
SCVER |
Fudan University | State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models | - | Sep. 2026 | |
FinixDoc |
Ant Group | FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks | Aug. 2026 | ||
![]() |
SmolDocling-KV-Link |
IBM | Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images | - | Aug. 2026 |
ArmorOCR |
Ant Group | ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation | Aug. 2026 | ||
NaviDC-OCR |
China Telecom AI | NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents | - | Aug. 2026 | |
TongGuOCR |
SCUT | TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents | Aug. 2026 | ||
PaDoc |
Tsinghua University | PaDoc: Layout-Grounded Parallel Decoding for Document Parsing | Aug. 2026 | ||
![]() |
Logographic Pretraining |
QMUL | Logographic Character Visual Pretraining via Semantic-based Contrastive Learning | - | Aug. 2026 |
![]() |
SPIRAL |
SEU | Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression | Aug. 2026 | |
![]() |
DocPO |
Tencent | DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards | - | Aug. 2026 |
DrawAI |
BUPT | DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable | Aug. 2026 | ||
LayoutLite |
Yuanli Technology | LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR | Jul. 2026 | ||
HPD-Parsing |
Baidu | HPD-Parsing: Hierarchical Parallel Document Parsing | Jul. 2026 | ||
OvisOCR2 |
Alibaba | OvisOCR2 Technical Report | Jul. 2026 | ||
DocOCR-Eval |
University of Melbourne | DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth | - | Jul. 2026 | |
MonkeyOCRv2 |
HUST&Kingsoft | MonkeyOCRv2: A Visual-Text Foundation Model for Document AI | Jul. 2026 | ||
Infinity-Parser2 |
INF Team | Infinity-Parser2 Technical Report | Jul. 2026 | ||
HunyuanOCR-1.5 |
Tencent | HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better | Jul. 2026 | ||
SAYRE |
Alibaba | Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis | - | Jul. 2026 | |
P-MTP |
Baidu | P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling | - | Jun. 2026 | |
RT-DocLayout |
Baidu | RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild | - | Jun. 2026 | |
Unlimited-OCR |
Baidu | Unlimited OCR Works | Jun. 2026 | ||
Beaver |
Microsoft Research | Building Agent Harnesses for Scientific Curation from Multimodal Sources | - | Jun. 2026 | |
Agents-K1 |
Shanghai AI Laboratory | Agents-K1: Towards Agent-native Knowledge Orchestration | Jun. 2026 | ||
PaddleOCR-VL-1.6 |
Baidu | PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training | Jun. 2026 | ||
PP-OCRv6 |
Baidu | PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks | Jun. 2026 | ||
StrucTab |
IIE, CAS | StrucTab: A Structured Optimization Framework for Table Parsing | Jun. 2026 | ||
ExChart |
Zhejiang University | Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework | - | Jun. 2026 | |
MinerU-Popo |
Shanghai AI Laboratory & OpenDataLab | MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing | May. 2026 | ||
![]() |
RTPrune |
DeepSeek | RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference | May. 2026 | |
ABot-OCR |
Alibaba | ABot-OCR Technical Report | May. 2026 | ||
BabelDOC |
funstory.ai | BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation | May. 2026 | ||
Consensus Entropy |
Fudan University | Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR | May. 2026 | ||
FastOCR |
Tsinghua University & JD | FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing | - | May. 2026 | |
MinerU2.5-Pro |
Shanghai AI Laboratory | MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale | Apr. 2026 | ||
PixelPrune |
OPPO | PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding | Apr. 2026 | ||
TexOCR |
Yale University & Zhejiang University | TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction | Apr. 2026 | ||
Falcon OCR |
Falcon Vision Team, TII | Falcon Perception | Mar. 2026 | ||
MinerU-Diffusion |
Shanghai AI Laboratory | MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding | Mar. 2026 | ||
Qianfan-OCR |
Baidu | Qianfan-OCR: A Unified End-to-End Model for Document Intelligence | Mar. 2026 | ||
dots.mocr |
HUST | Multimodal OCR: Parse Anything from Documents | Mar. 2026 |
📄 See full list at Document-Parsing.md
📄 Document Understanding
Document understanding extends beyond structural parsing to semantic comprehension and reasoning over visually rich documents.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|---|---|---|---|---|
InterTab |
HKUST | InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning | - | Sep. 2026 | |
D-RAC |
Yellow.ai | Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion | - | Sep. 2026 | |
Chart-RVR |
University of Virginia | Monitorable Chart Reasoning Agents via Verifiable Process Rewards | Sep. 2026 | ||
ViSAR |
INSA Lyon | ViSAR: Training-Free Adaptive-k Retrieval for Visual Document Question Answering | - | Sep. 2026 | |
DocIntent |
SCUT | DocIntent: Answerability-Guided Agentic Restoration for Real-World Document Visual Question Answering | - | Sep. 2026 | |
![]() |
Doc-REFRAG |
Zhejiang University | Doc-REFRAG: Rethinking Multimodal Document Retrieval-Augmented Generation | Sep. 2026 | |
![]() |
AWM |
University of Oslo | AWM: Answerable Working Memory for Long-Document VQA Agents | Aug. 2026 | |
SAGE |
Fudan University | SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding | - | Aug. 2026 | |
Q-Guide |
Amazon | Question-Guided Evidence Acquisition for Multimodal Visual Question Answering | - | Aug. 2026 | |
DocClaw |
NTU | DocClaw: A Unified Agentic System for Intelligent Document Processing | Aug. 2026 | ||
ConceptFormer |
Northeastern University | ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval | Aug. 2026 | ||
![]() |
Hyper-M2RAG |
Hangzhou Dianzi University | Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement | Aug. 2026 | |
Trident |
Emory University | What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering | - | Aug. 2026 | |
D2-ScaleAgent |
Zhejiang University | D2-ScaleAgent: Dual-Dimensional Scaling for Long Document Understanding | - | Aug. 2026 | |
![]() |
SEER |
UT Austin | SEER: Long-Context Reasoning via Selective Visual-Text Compression | Aug. 2026 | |
HAM-RAG |
HKUST | HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation | Aug. 2026 | ||
![]() |
DRUF |
Shenzhen MSU-BIT University | Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs | Aug. 2026 | |
DistilVDR |
Aalto University | DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation | Aug. 2026 | ||
InSight-doc |
HKUST | InSight-doc: Agentic Visual Perception for Long-Document Understanding | Aug. 2026 | ||
DocAtlas |
Wuhan University / Microsoft | DocAtlas: Long-Document Understanding as Mutable-State Interaction | - | Aug. 2026 | |
DocMemo |
HIT, Shenzhen | DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding | Aug. 2026 | ||
ECF |
Beihang University | Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models? | Aug. 2026 | ||
ADOPD 2026 |
Georgia Tech | Thinking with Anchors: Grounded and Efficient Document Reasoning | Aug. 2026 | ||
VTS |
MBZUAI | When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning | - | Aug. 2026 | |
Q-CueGraph |
The University of Tokyo | Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning | - | Aug. 2026 | |
DocTrace |
Baidu | DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning | - | Aug. 2026 | |
CURV |
William & Mary | CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning | - | Aug. 2026 | |
RAGOCR |
Peking University | RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation | - | Aug. 2026 | |
VaRS-Doc |
SJTU | VaRS-Doc: Interpretation-Aware Variant Representations via Latent Self-Probing for Visual Document Retrieval | Aug. 2026 | ||
ET-Prune |
SJTU | ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs | Aug. 2026 | ||
HierDoc |
USTC | HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering | - | Aug. 2026 | |
![]() |
DualG-MRAG |
BUAA | DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation | - | Jul. 2026 |
MMLDSum-LLM |
OPPO | MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware | - | Jul. 2026 | |
TAP-RAG |
Tianjin University | TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering | Jul. 2026 | ||
![]() |
Perception-RFT |
Quantiphi | Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment | - | Jul. 2026 |
CMDR |
NTT | CMDR: Contextual Multimodal Document Retrieval | Jul. 2026 | ||
HiEvi-RAG |
USTC | Hierarchical Evidence-Driven Reasoning for Long Document Understanding | - | Jul. 2026 | |
MultAttnAttrib |
— | MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering | - | Jul. 2026 | |
OracleAnalyser |
NUDT | OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training | - | Jun. 2026 | |
DocArena |
Adobe Research | DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents | - | Jun. 2026 | |
ViTexQA |
Meituan | ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering | Jun. 2026 | ||
PreciseDoc |
Tsinghua University | An LMM for Precisely Grounding Elements in Documents | - | Jun. 2026 | |
LightSTAR |
SJTU | LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement | Jun. 2026 | ||
SciLens |
HKUST | SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding | - | Jun. 2026 | |
SAFE-Cascade |
Walmart | SAFE-Cascade: Cost-Adaptive Vision-Language Routing for Chart Question Answering | - | Jun. 2026 | |
UMG-RAG |
Purdue University | Uncertainty-Aware Hybrid Retrieval for Long-Document RAG | - | Jun. 2026 | |
MAGE-RAG |
BIT | MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA | Jun. 2026 | ||
MINARD |
University of Maryland | Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures | - | Jun. 2026 | |
MM-BizRAG |
JPMorgan Chase & Co. | MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A | - | Jun. 2026 | |
KG4VD |
National Taiwan University | Multimodal Graph RAG for Long-range Visually Rich Document Understanding | Jun. 2026 |
📄 See full list at Document-Understanding.md
📄 Visual Text Generation
Visual Text Generation focuses on generating or editing legible, visually harmonious, and semantically consistent text within images, serving as the creative inverse of OCR. In the LLM era, it is no longer a task reliant on specialized modules, but is emerging as a foundational skill for general-purpose generative models.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|---|---|---|---|---|
IDSpect |
Fudan University | Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering | - | Sep. 2026 | |
DuetGen |
USTC | Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation | - | Sep. 2026 | |
LoGAN |
Netflix | LoGAN: Multilingual Font Localization with Generative Agents | - | Sep. 2026 | |
GlyphAnchor |
Fudan University | GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors | - | Sep. 2026 | |
TextRefine |
Kuaishou | TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters | - | Aug. 2026 | |
PosterText |
Wuhan University | PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster | - | Aug. 2026 | |
TransAnyText |
Wuhan University | TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation | - | Aug. 2026 | |
onoff |
GIST | Bridging Online and Offline Handwriting via Differentiable Physical Rendering | Aug. 2026 | ||
PosterMELD |
Tsinghua University | PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs | Aug. 2026 | ||
![]() |
InnoText |
SYSU | InnoText: A Unified Model for Visual Text Generation and Editing | - | Jul. 2026 |
Boogu-Image-0.1 |
Boogu Team,Huawei | Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation | Jul. 2026 | ||
SciForma |
Microsoft & Peking University | SciForma: Structure-Faithful Generation of Scientific Diagrams | Jul. 2026 | ||
VecFontLLM |
Fuzhou University & Peking University | VecFontLLM: Anchor-Guided Direct Synthesis of Chinese Vector Fonts | - | Jul. 2026 | |
ArtChart |
Ant Group | ArtChart: A Benchmark for Faithful Artistic Chart Generation with Integrated Text Rendering | - | Jul. 2026 | |
DataEvolver |
CSU | DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation | Jul. 2026 | ||
Qwen-Image-2.0-RL |
Alibaba | Qwen-Image-2.0-RL Technical Report | - | Jun. 2026 | |
UniTranslator |
IIE CAS | UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation | Jun. 2026 | ||
SteerVTE |
ByteDance & Peking University | SteerVTE: Seamless Video Text Editing with Style and Glyph Control | - | Jun. 2026 | |
DiffMath |
SCUT & Huawei | DiffMath: Symbol- and Graph-Aware Latent Diffusion Transformer for Handwritten Mathematical Expression Generation | Jun. 2026 | ||
PhyDrawGen |
University of Dhaka | PhyDrawGen: Physically Grounded Diagram Generation from Natural Language | - | Jun. 2026 | |
NIV |
Reichman University | NIV: Neural Axis Variations for Variable Font Generation | Jun. 2026 | ||
Qwen-Image-2.0 |
Alibaba Group | Qwen-Image-2.0 Technical Report | May. 2026 | ||
MangaFlow |
The University of Tokyo | MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation | - | May. 2026 | |
Wan-Image |
Alibaba Group | Wan-Image: Pushing the Boundaries of Generative Visual Intelligence | - | Apr. 2026 | |
![]() |
PosterIQ |
PolyU | PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation | Mar. 2026 | |
EfficientPosterGen |
Tsinghua University | EfficientPosterGen: Semantic-aware Efficient Poster Generation via Token Compression and Accurate Violation Detection | Mar. 2026 | ||
TextFlow |
NJUST | Towards Training-Free Scene Text Editing | Mar. 2026 | ||
GlyphPrinter |
Fudan University | GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering | Mar. 2026 | ||
CTRL-S |
SJTU | Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning | - | Mar. 2026 | |
LaDe |
Adobe Research | LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition | - | Mar. 2026 | |
EchoGen |
USTC | EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding | - | Mar. 2026 | |
WebVR |
StepFun | WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics | - | Mar. 2026 | |
AutoFigure-Edit |
Westlake University | AutoFigure-Edit: Generating Editable Scientific Illustration | Mar. 2026 | ||
GlyphBanana |
SJTU & Xiaohongshu Inc. | GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows | Mar. 2026 | ||
InnoAds-Composer |
JD.com | InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation | - | Mar. 2026 | |
Seeing is Improving |
USTC | Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement | Mar. 2026 | ||
![]() |
PosterOmni |
HKUST(GZ) | PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback | Feb. 2026 | |
![]() |
TextPecker |
HUST | TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering | Feb. 2026 | |
SSPT |
Tongji University | Space Syntax-guided Post-training for Residential Floor Plan Generation | - | Feb. 2026 | |
ChatUMM |
Tsinghua University & Tencent Hunyuan | ChatUMM: Robust Context Tracking for Conversational Interleaved Generation | - | Feb. 2026 | |
![]() |
PosterVerse |
SCUT | PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography | Jan. 2026 | |
| - | GPT-Image-1.5 |
OpenAI | GPT-Image-1.5 | --- | Dec. 2025 |
| - | Gemini 3 Pro Image |
Google DeepMind | Gemini 3 Pro Image (Nano Banana Pro) | --- | Nov. 2025 |
| - | Gemini 2.5 Flash Image |
Google DeepMind | Gemini 2.5 Flash Image (Nano Banana) | --- | Oct. 2025 |
Qwen-Image |
Qwen Team | Qwen-Image Technical Report | Sep. 2025 | ||
Seedream 4.0 |
ByteDance Seed | Seedream 4.0: Toward Next-generation Multimodal Image Generation | --- | Aug. 2025 | |
Postergen |
Stony Brook University | Postergen: Aesthetic-aware paper-to-poster generation via multi-agent llms | Aug. 2025 | ||
![]() |
UniGlyph |
Tsinghua University | UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis | --- | Jul. 2025 |
X-Omni |
Tencent Hunyuan X | X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again | Jul. 2025 | ||
DreamPoster |
Intelligent Creation Lab, ByteDance | DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design | --- | Jul. 2025 |
📄 See full list at Visual-Text-Generation.md
📄 Specialized Model
Beyond general document parsing and understanding, traditional OCR research remains essential for specialized visual-text structures, restoration, temporal analysis, and forensic reliability.
📄 Document Dewarping
Document dewarping restores photographed or scanned pages to a geometrically rectified and readable form.
📄 Physical Structure Analysis
Physical structure analysis identifies document regions, layouts, and spatial relationships before semantic reasoning.
📄 Reading Order Prediction
Reading order prediction recovers the intended sequence of text and visual elements in complex page layouts.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|---|---|---|---|---|
Orli |
University of Konstanz | End-to-End Text Line Detection and Ordering | Jun. 2026 | ||
| [














