← 开源
worldbench

awesome-ai-auto-research

🔥 A Survey on AI Auto-Research

ListsPaper collectionsHTML
在 GitHub 打开
增长势头
+024 小时新增 Star0.0%
541
Star
41
Fork
+3
本周
7
贡献者
创建于 2026-03-29 · 更新于 2026-10-06 · 今日第 16520 名
主要开发者
README

Awesome Logo arXiv Project Page Visitors PR's Welcome

:sunglasses: Awesome AI Auto-Research

This repository accompanies the survey paper "AI for Auto-Research: Roadmap & User Guide" and tracks papers on AI-assisted and automated scientific research, covering the full research lifecycle.

:robot: AI Auto-Research

We organize the academic research lifecycle as eight interconnected stages grouped into four epistemological phases. Each phase serves a distinct function in producing, scrutinizing, and communicating scientific knowledge.

Phase 1: Creation
Generating novel research ideas, searching and synthesizing literature, running coding experiments, and creating publication-quality tables and figures. This phase spans Idea Generation, Literature Review, Coding & Experiments, and Tables & Figures.
Phase 2: Writing
Drafting, editing, and polishing academic manuscripts. AI assistance ranges from semi-automated grammar and citation tools to fully automated paper generation — the most commercially mature yet ethically contested stage.
Phase 3: Validation
Automated peer review generation, reviewer-paper matching, review quality assessment, and AI-assisted author rebuttals. This phase covers Peer Review and Rebuttal & Revision.
Phase 4: Dissemination
Converting papers into slides, posters, videos, websites, and social media content. Each output format targets a different audience and demands its own design logic and AI tool chain.

For additional details, kindly refer to our :books: Paper and :earth_asia: Project Page.

:books: Citation

If you find this work helpful for your research, please kindly consider citing our paper:

@article{survey-ai-auto-research,
  title   = {{AI} for {Auto-Research}: Roadmap \& User Guide},
  author  = {Kong, Lingdong and Sun, Xian and Chow, Wei and Li, Linfeng and Lin, Kevin Qinghong and Zhang, Xuan Billy
             and Wang, Song and Li, Rong and Wu, Qing and Gao, Wei and Wang, Yingshuo and Xie, Shaoyuan
             and Liu, Jiachen and Qu, Leigang and Li, Shijie and Ng, Lai Xing and Cottereau, Benoit R.
             and Liu, Ziwei and Chua, Tat-Seng and Ooi, Wei Tsang},
  journal = {arXiv preprint arXiv:2605.18661},
  year    = {2026}
}

Table of Contents

1. Idea Generation

LLM Internal Knowledge-Based Generation

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
Chain of Ideas arXiv
Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents arXiv '24 - GitHub
ResearchAgent Website
ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models NAACL '25 - GitHub
SciMON arXiv
SciMON: Scientific Inspiration Machines Optimized for Novelty ACL '24 - GitHub
Idea Gen Agent arXiv
Can LLMs Generate Novel Research Ideas? A Large Scale Human Study with 100+ NLP Researchers arXiv '24 - -
IRIS Website
IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery ACL '25 - GitHub
Spark arXiv
Spark: A System for Scientifically Creative Idea Generation ICCC '25 - -
Diverse Hypo. Search arXiv
Towards Diverse Scientific Hypothesis Search with Large Language Models arXiv '26 - -
Tree-of-Ideas arXiv
Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution arXiv '26 - -
IDEAgent arXiv
IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation arXiv '26 - -

External Signal-Driven Generation

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
SGHA arXiv
SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models arXiv '26 - -
MAIL arXiv
MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in Chemistry arXiv '26 - -
MOOSE-Chem Website
MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses ICLR '25 - -
Nova arXiv
Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas arXiv '24 - -
SciAgents arXiv
SciAgents: Automating Scientific Discovery through Multi-Agent Intelligent Graph Reasoning arXiv '24 - GitHub
SciPIP arXiv
SciPIP: An LLM-based Scientific Paper Idea Proposer arXiv '24 - GitHub
IdeaSynth arXiv
IdeaSynth: Iterative Research Idea Development Through Evolving and Composing Idea Facets with Literature-Grounded Feedback CHI '25 - -
MOOSE-Chem2 Website
MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search NeurIPS '25 - -
HALO arXiv
HALO: Interactive Co-abductive Reasoning in Scientific Hypothesis Generation arXiv '26 - -
TCA-SIR arXiv
TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval arXiv '26 - -
ECLAIR arXiv
ECLAIR: A Causally-Grounded AI Framework for Scientific Discovery in Empirical Software Engineering arXiv '26 - -

Multi-Agent Collaborative Generation

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
HypoForge arXiv
HypoForge: A Self-Improving Multi-Agent Framework for Automated Hypothesis Generation and Testing via Scientific Skill Learning arXiv '26 - -
Combi. Creativity arXiv
Combi. Creativity arXiv '24 - -
Deep Ideation arXiv
Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network arXiv '25 - GitHub
VirSci Website
Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System ACL '25 - GitHub
Multi-Agent Dial. arXiv
Multi-Agent Dial. SIGDIAL '25 - -
Artificial Hivemind arXiv
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) NeurIPS '25 - -
Auditable AI Sci. arXiv
Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents arXiv '26 - -
Diverse Personalized Ideation arXiv
Diversifying Personalized Research Ideation against AI-Induced Homogenization arXiv '26 - -

Novelty and Feasibility Assessment

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
RATIO arXiv
RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature arXiv '26 - -
Lit2Test arXiv
What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation arXiv '26 - -
Energy Scoring arXiv
Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking arXiv '26 - -
Think-Probe-Respond arXiv
Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty EMNLP '26 - -
IdeaBench Website
LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context KDD '25 - -
LiveIdeaBench arXiv
LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context arXiv '24 - -
AI Idea Bench 2025 arXiv
AI Idea Bench 2025: AI Research Idea Generation Benchmark arXiv '25 - GitHub
HeurekaBench arXiv
HeurekaBench: A Benchmarking Framework for AI Co-scientist ICLR '26 - GitHub
ResearchBench arXiv
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition ACL '26 - -
HindSight arXiv
HindSight: Evaluating LLM-Generated Research Ideas via Future Impact arXiv '26 - -
Rubric Rewards arXiv
Training AI Co-Scientists Using Rubric Rewards arXiv '25 - -
DeepInnovator arXiv
DeepInnovator: Triggering the Innovative Capabilities of LLMs arXiv '26 - GitHub
FlowPIE arXiv
FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration arXiv '26 - -
SoundnessBench arXiv
SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? arXiv '26 - -
LLM-Judge Novelty arXiv
On the Limits of LLM-as-Judge for Scientific Novelty Assessment arXiv '26 - -
LigBench arXiv
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation arXiv '26 - -
Reconstruction arXiv
Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies arXiv '26 - -
AgentIdeaBench arXiv
AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era arXiv '26 - GitHub
IdeaAMBIG arXiv
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications arXiv '26 - -
NovGauge arXiv
NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment arXiv '26 - -

2. Literature Review & Paper Search

Literature Retrieval

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
ITER arXiv
ITER: Interaction-Aware Retrieval for Agentic Search arXiv '26 - -
Multi-Aspect Retrieval arXiv
Can Retrievers Find the Same Paper from Different Aspects? A Multi-Aspect Full-Paper Scientific Retrieval Benchmark arXiv '26 - -
CiteME arXiv
CiteME: Can Language Models Accurately Cite Scientific Claims? arXiv '24 - -
LitLLM arXiv
LitLLM: A Toolkit for Literature Review with Large Language Models arXiv '24 - -
LitSearch arXiv
LitSearch: A Retrieval Benchmark for Scientific Literature Search arXiv '24 - GitHub
PaperQA2 arXiv
Language Agents Achieve Superhuman Synthesis of Scientific Knowledge arXiv '24 - GitHub
OpenResearcher arXiv
OpenResearcher: Unleashing AI for Accelerated Scientific Research EMNLP '24 - -
PaSa arXiv
PaSa: An LLM Agent for Comprehensive Academic Paper Search arXiv '25 - GitHub
Self-Evolving Retrieval arXiv
Towards Self-Evolving Agentic Literature Retrieval arXiv '26 - -
MasterSet arXiv
MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature arXiv '26 - -
Search, Inspect, Fetch arXiv
Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents arXiv '26 - -
Rubric Reranker arXiv
Training Documents Reranker with Search Rubrics for Deep Research Agent arXiv '26 - -
Personalized DR Refinement arXiv
Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding arXiv '26 - -

Survey & Related Work Generation

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
DAS arXiv
Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation arXiv '26 - -
Tree-of-Concerns arXiv
Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique EMNLP '26 - -
ChatPaper Website
ChatPaper: Use LLM to summarize papers GitHub '23 - GitHub
PaperQA arXiv
PaperQA: Retrieval-Augmented Generative Agent for Scientific Research arXiv '23 - GitHub
AutoSurvey arXiv
AutoSurvey: Large Language Models Can Automatically Write Surveys arXiv '24 - GitHub
GPT Researcher Website
GPT Researcher: Autonomous Agent for Comprehensive Online Research GitHub '24 - GitHub
LLMs for Lit. Review arXiv
LLMs for Lit. Review arXiv '24 - -
STORM arXiv
Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models arXiv '24 - GitHub
Agentic AutoSurvey arXiv
Agentic AutoSurvey: Let LLMs Survey LLMs arXiv '25 - -
Citegeist arXiv
Citegeist: Automated Generation of Related Work Analysis on the arXiv Corpus arXiv '25 - -
IterSurvey arXiv
IterSurvey: Deep Literature Survey Automation with an Iterative Workflow arXiv '25 - GitHub
LiRA arXiv
LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation arXiv '25 - -
SurveyForge arXiv
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing arXiv '25 - GitHub
SurveyG arXiv
SurveyG: A Multi-Agent LLM Framework with Hierarchical Citation Graph for Automated Survey Generation arXiv '25 - -
SurveyX arXiv
SurveyX: Academic Survey Automation via Large Language Models arXiv '25 - -
InteractiveSurvey arXiv
InteractiveSurvey: An LLM-based Personalized and Interactive Survey Paper Generation System arXiv '25 - GitHub
CiteLLM arXiv
CiteLLM: An Agentic Platform for Trustworthy Scientific Reference Discovery arXiv '26 - -
DeepSurvey arXiv
DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation arXiv '26 - -
STRUCTSURVEY arXiv
STRUCTSURVEY: Structured Agentic Retrieval for Automated Survey Paper Generation arXiv '26 - -

Deep Research Agents

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
Crase arXiv
Structurally-bounded Agentic Graph Exploration for Evidence-Grounded Scholarly DeepSearch arXiv '26 - -
DeepWeaver arXiv
DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering arXiv '26 - -
AgentR arXiv
AgentR A Stateful and Recovery-Aware Software Architecture for LLM-based Auditable Workflows arXiv '26 - -
ASReview Website
An Open Source Machine Learning Framework for Efficient and Transparent Systematic Reviews Nature MI '21 - GitHub
CHIME arXiv
CHIME: LLM-Assisted Hierarchical Organization of Scientific Studies for Literature Review Support arXiv '24 - -
DeepResearch-Agent Website
DeepResearchAgent: A Hierarchical Multi-Agent System for Deep Research GitHub '25 - GitHub
DeerFlow Website
DeerFlow: A Deep Research Framework Orchestrating Sub-Agents, Memory, and Sandboxes GitHub '25 - GitHub
OpenScholar Website
OpenScholar: Synthesizing Scientific Literature with Retrieval-Augmented LMs Nature '26 - -
AutoAgent arXiv
AutoAgent arXiv '25 - -
Tongyi DeepResearch Website
Tongyi DeepResearch GitHub '25 - GitHub
O-Researcher arXiv
O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL arXiv '26 - -
OpenResearcher arXiv
OpenResearcher: Unleashing AI for Accelerated Scientific Research arXiv '26 - GitHub
AREX arXiv
AREX: Towards a Recursively Self-Improving Agent for Deep Research arXiv '26 - -
Predictive Navigation arXiv
Deep Research Pretraining via Predictive Navigation arXiv '26 - -
On-Device DR (4B) arXiv
On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage arXiv '26 - -
Carnot arXiv
Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries VLDB '26 - -
Marginal Value Est. arXiv
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents arXiv '26 - -
Retrieval-Aware Control arXiv
When Deep Research Agents Stagnate: Enhancing Reasoning with Retrieval-Aware Agent Control arXiv '26 - -
Analogical Deep Research arXiv
Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis arXiv '26 - -
Plato-Bio arXiv
Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks arXiv '26 - -
Albilich arXiv
Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration arXiv '26 - -

Retrieval and Synthesis Quality Assessment

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
IBIS arXiv
From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation EMNLP '26 - -
Agent to Blame arXiv
Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research EMNLP '26 - -
DeepScholar-Bench arXiv
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis arXiv '25 - GitHub
ReportBench arXiv
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks arXiv '25 - GitHub
IDRBench arXiv
IDRBench: Interactive Deep Research Benchmark arXiv '26 - -
ScholarGym arXiv
ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research arXiv '26 - -
SciNetBench arXiv
SciNetBench: A Relation-Aware Benchmark for Scientific Literature Retrieval Agents arXiv '26 - -
AutoResearchBench arXiv
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery arXiv '26 - -
PaperMind arXiv
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs arXiv '26 - -
DRNOISE arXiv
DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments arXiv '26 - -
HiEviDR-Bench arXiv
HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research arXiv '26 - -
SciExplore arXiv
SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration arXiv '26 - -
WANDR arXiv
WANDR: A Benchmark for Wide and Deep Research arXiv '26 - -
QA-to-DR Bench arXiv
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution arXiv '26 - -
PRISMA-LLM arXiv
PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews arXiv '26 - -
INSPIRE arXiv
Inspire: Benchmarking Scientific Literature Search for Open Research Problems arXiv '26 - -
CESS arXiv
Search Shapes Conclusions: Auditing Evidence Selection Bias in Deep Research Agents arXiv '26 - -
ScholarCatalyst arXiv
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research arXiv '26 - GitHub

3. Coding & Experimentation

Code Generation

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
SWE-bench arXiv
SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR '24 - GitHub
SWE-agent arXiv
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering arXiv '24 - GitHub
OpenHands arXiv
OpenHands: An Open Platform for AI Software Developers as Generalist Agents ICLR '25 - GitHub
SWE-bench Pro arXiv
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? arXiv '25 - -
SWE-EVO arXiv
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios arXiv '25 - -

Paper-to-Code

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
ReproAgent arXiv
ReproAgent: Contract-Guided Paper-to-Code Reproduction EMNLP '26 - -
DeepRepro arXiv
DeepRepro: State-Aware Subplanning for Paper-to-Code Reproduction in Evolving Repositories CIKM '26 - -
FunSearch Website
Mathematical Discoveries from Program Search with Large Language Models Nature '24 - GitHub
SciCode arXiv
SciCode: A Research Coding Benchmark Curated by Scientists arXiv '24 - GitHub
PaperBench arXiv
PaperBench: Evaluating AI's Ability to Replicate AI Research arXiv '25 - GitHub
PaperCoder arXiv
Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning arXiv '25 - GitHub
ResearchCodeBench arXiv
ResearchCodeBench: Benchmarking LLMs on Implementing Novel ML Research Code arXiv '25 - -
SciReplicate-Bench arXiv
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers arXiv '25 - GitHub
PaperCompiler arXiv
PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation arXiv '26 - -

Experiment Execution & Orchestration

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
Praxist arXiv
Praxist: From Experimental Artifacts to Solution Lineages arXiv '26 - -
Skill-Based Baselines arXiv
Can Coding Agents Build Robust Baselines? A Skill-Based Approach for Automating the Medical Imaging Model-Development Pipeline MICCAI '26 - -
BioPlanner arXiv
BioPlanner: Automatic Evaluation of LLMs on Protocol Planning arXiv '23 - GitHub
CRISPR-GPT arXiv
CRISPR-GPT for Agentic Automation of Gene-Editing Experiments arXiv '24 - -
DS-Agent arXiv
DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning arXiv '24 - GitHub
MLE-Bench arXiv
MLE-Bench: Evaluating Machine Learning Agents on Machine Learning Engineering arXiv '24 - -
MLAgentBench arXiv
MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation arXiv '24 - GitHub
MLR-Copilot arXiv
MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents arXiv '24 - -
AIDE arXiv
AIDE: AI-Driven Exploration in the Space of Code arXiv '25 - -
AlphaEvolve arXiv
AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery arXiv '25 - -
AutoReproduce arXiv
AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage arXiv '25 - GitHub
CURIE arXiv
Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents arXiv '25 - GitHub
MLGym arXiv
MLGym: A New Framework and Benchmark for Advancing AI Research Agents arXiv '25 - -
MLR-Bench arXiv
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research arXiv '25 - -
Execution-Grounded arXiv
Towards Execution-Grounded Automated AI Research arXiv '26 - -
Learn to Discover arXiv
Learning to Discover at Test Time arXiv '26 - -
AutoNumerics arXiv
AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing arXiv '26 - GitHub
SciNav arXiv
SciNav: A General Agent Framework for Scientific Coding Tasks arXiv '26 - -
FrontierScience arXiv
FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks arXiv '26 - -
EvoDS arXiv
EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management arXiv '26 - -
AutoTTS arXiv
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling arXiv '26 - GitHub
AutoScientists arXiv
AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation arXiv '26 - -
EurekAgent arXiv
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery arXiv '26 - -
Experimental Experience Modeling arXiv
Experimental Experience Modeling for Autonomous Research arXiv '26 - -

Code Correctness and Reproducibility Assessment

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
Experimental Fidelity arXiv
Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research arXiv '26 - -
DiscoveryBench arXiv
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models arXiv '24 - GitHub
DiscoveryWorld arXiv
DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents arXiv '24 - GitHub
InfiAgent-DABench arXiv
InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks arXiv '24 - -
ScienceAgentBench arXiv
ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery arXiv '24 - -
LAB-Bench arXiv
Lab-Bench: Measuring Capabilities of Language Models for Biology Research arXiv '24 - GitHub
KernelBench arXiv
KernelBench: Can LLMs Write Efficient GPU Kernels? arXiv '25 - GitHub
TritonBench arXiv
TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators arXiv '25 - GitHub
AstaBench arXiv
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite arXiv '25 - GitHub
ResearchClawBench arXiv
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows arXiv '25 - GitHub
EXP-Bench Website
EXP-Bench: Can AI Conduct AI Research Experiments? ICLR '26 - GitHub
PostTrainBench arXiv
PostTrainBench: Can LLM Agents Automate LLM Post-Training? arXiv '26 - GitHub
MLReplicate arXiv
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility arXiv '26 - -
BeyondSWE arXiv
BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? arXiv '26 Website GitHub
NatureBench arXiv
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? arXiv '26 - -
SciCoQA arXiv
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment ACL '26 Website GitHub
Dude arXiv
Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection arXiv '26 - GitHub
AgentActionBench arXiv
Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers arXiv '26 - -
AutoDataBench arXiv
AutoDataBench: A Data-centric Testbed for Accelerating Auto Research arXiv '26 - GitHub
EurekaBench arXiv
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights arXiv '26 - -

4. Tables & Figures

Scientific Figure Generation

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
ChartGPT arXiv
ChartGPT: Leveraging LLMs to Generate Charts from Abstract Natural Language arXiv '23 - -
MatPlotAgent arXiv
MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data Visualization arXiv '24 - -
CoDA arXiv
CoDA: Agentic Systems for Collaborative Data Visualization arXiv '25 - -
PlotGen arXiv
PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback arXiv '25 - -
VIS-Shepherd arXiv
VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation arXiv '25 - -
DiagramAgent arXiv
From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing CVPR '25 - -
StarVector arXiv
StarVector: Generating Scalable Vector Graphics Code from Images and Text CVPR '25 - -
VisCoder arXiv
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation EMNLP '25 - -
AI-Generated Figures arXiv
AI-Generated Figures arXiv '26 - -
AutoFigure-Edit arXiv
AutoFigure-Edit: Generating Editable Scientific Illustration arXiv '26 - GitHub
AutoFigure arXiv
AutoFigure-Edit: Generating Editable Scientific Illustration ICLR '26 - GitHub
PaperBanana arXiv
PaperBanana: Automating Academic Illustration for AI Scientists arXiv '26 - -
SAIL arXiv
Setting SAIL: Leveraging Scientist-AI-Loops for Rigorous Visualization Tools arXiv '26 - -
Crafter arXiv
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs arXiv '26 - -
DiagramRAG arXiv
DiagramRAG: A Lightweight Framework to Retrieve Scientific Diagram for Figure Generation arXiv '26 - -
GeoSVG-RL arXiv
GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation arXiv '26 - -
Can AI Draw Sci. arXiv
Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models arXiv '26 - -
SciDiagramEdit arXiv
SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions arXiv '26 - -
GenGA arXiv
GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers arXiv '26 - -
FigTree arXiv
Figures as Programs: Recursive Generation of Editable Scientific Figures arXiv '26 - -
EdiTikZ arXiv
EdiTikZ: Scientific Figure Editing from Revision Trajectories arXiv '26 - GitHub

Table Understanding & Generation

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
ArxivDIGESTables arXiv
ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models EMNLP '24 - -
Chain-of-Table arXiv
Chain-of-Table: Evolving Tables in Reasoning Chain for Table Understanding ICLR '24 - -
ShowTable arXiv
ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement CVPR '26 - -
Table2LaTeX-RL arXiv
Table2LaTeX-RL: Converting Table Images to High-Fidelity LaTeX Code Using Reinforced Multimodal Language Models arXiv '25 - -
CSPO arXiv
CSPO: Alleviating Reward Ambiguity for Structured Table-to-LaTeX Generation arXiv '26 - -

Mathematical Formulas & TikZ

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
AutomaTikZ arXiv
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ ICLR '24 - -
DeTikZify arXiv
DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ NeurIPS '24 - -
TikZilla arXiv
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning arXiv '26 - -
Edit2TikZ arXiv
Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ arXiv '26 - -

Visual Fidelity and Scientific Accuracy Assessment

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
StructEval Website
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs TMLR '25 Website GitHub
PlotCraft arXiv
PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization arXiv '25 - -
TeXpert Website
TeXpert: Multi-Level Benchmark for LaTeX Code Generation SDP '25 - -
AbGen arXiv
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research ACL '25 - -
SciFig arXiv
SciFig: Towards Automating Scientific Figure Generation arXiv '26 - -
SciFlow-Bench arXiv
SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing arXiv '26 - -
FigureBench arXiv
AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations ICLR '26 - GitHub
SciFigQual-Bench arXiv
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context arXiv '26 - -
SciFigAlign arXiv
SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence arXiv '26 - -
SciFigPlag-Bench arXiv
SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection arXiv '26 - -
VLM Blind/Misled arXiv
How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures arXiv '26 - -
SciFigure2Code arXiv
SciFigure2Code: An AI-Reconstructed Benchmark for Scientific Figure-to-Code arXiv '26 - -
ReFigBench arXiv
ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts arXiv '26 - -

5. Paper Writing

Semi-Automated Writing Assistance

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
CoAuthor arXiv
CoAuthor: Human-AI Collaborative Writing with Language Models arXiv '22 - -
AI Writing Study arXiv
AI Writing Study AIED '25 - -
DraftMarks arXiv
DraftMarks: Enhancing Transparency in Human-AI Co-Writing Through Interactive Skeuomorphic Process Traces arXiv '25 - -
PaperDebugger arXiv
PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing arXiv '25 - GitHub
ScholarCopilot arXiv
ScholarCopilot: Training LLMs for Academic Writing with Integrated Citation arXiv '25 - -
XtraGPT arXiv
XtraGPT: Context-Aware and Controllable Academic Paper Revision arXiv '25 - -
LimAgents arXiv
Multi-Agent LLMs for Generating Research Limitations arXiv '26 - -
PaperMentor arXiv
PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers on Overleaf arXiv '26 - -
AutoSupervision arXiv
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification arXiv '26 - -
ReasFlow arXiv
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System arXiv '26 - -

Fully Automated Paper Generation

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
CycleResearcher arXiv
CycleResearcher: Improving Automated Research via Automated Review ICLR '25 - -
Agent Laboratory Website
Agent Laboratory: Using LLM Agents as Research Assistants EMNLP '25 - -
FutureGen arXiv
FutureGen: A RAG-based Approach to Generate the Future Work of Scientific Article arXiv '25 - -
AI Scientist arXiv
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery Nature '26 - GitHub
APRES arXiv
APRES: An Agentic Paper Revision and Evaluation System arXiv '26 - -
LECTOR arXiv
LECTOR: Joint Optimization of Scientific Reasoning Graphs and Introduction Generation arXiv '26 - -
RWGBench arXiv
RWGBench: Evaluating Scholarly Positioning in Related Work Generation arXiv '26 - -

Societal Analysis

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
AI Writing Adoption Website
AI Writing Adoption Nature '26 - -
Nature AI Survey Website
More than Half of Researchers Now Use AI for Peer Review Nature '26 - -
Denial of Science arXiv
Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud arXiv '26 - -
AI Slop OSS arXiv
"AI Slop is DDoSing Open Source": Understanding the Impact of AI-Generated Contributions on Open Source Sustainability arXiv '26 - -

Writing Quality and AI Detection Assessment

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
Mapping LLM Use arXiv
Mapping the Increasing Use of LLMs in Scientific Papers arXiv '24 - -
CycleReviewer arXiv
CycleResearcher: Improving Automated Research via Automated Review ICLR '25 - -
Stanford Agentic Website
Stanford Agentic Web '25 - -
SciIG arXiv
Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper arXiv '25 - -
Watermarking arXiv
Detecting LLM-Generated Peer Reviews arXiv '25 - -
PaperWritingBench arXiv
PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing arXiv '26 - -
CiteTracer arXiv
Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection arXiv '26 - -
Process Eval arXiv
Process-Oriented Evaluation of AI-Assisted Scientific Writing arXiv '26 - -
SciSlopBench / SciSlopHarness arXiv
Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers arXiv '26 Website -

6. Peer Review

Automated Review Generation

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
ChatReviewer Website
ChatReviewer: ChatGPT-based Paper Reviewing and Response Generation GitHub '23 - GitHub
AI-Peer-Review Website
AI-Peer-Review GitHub '24 - GitHub
MARG arXiv
MARG: Multi-Agent Review Generation for Scientific Papers arXiv '24 - -
Reviewer2 arXiv
Reviewer2: Optimizing Review Generation Through Prompt Generation arXiv '24 - -
ReviewRL Website
ReviewRL: Towards Automated Scientific Review with RL EMNLP '25 - -
DeepReviewer arXiv
DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process arXiv '25 - -
OpenReviewer Website
OpenReviewer: A Specialized Large Language Model for Generating Critical Scientific Paper Reviews NAACL '25 - -
REMOR arXiv
REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning arXiv '25 - -
ScholarPeer arXiv
ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review arXiv '26 - -
ProReviewer arXiv
From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent arXiv '26 - -
PeerCheck arXiv
PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality arXiv '26 - -
Local Pre-Screening arXiv
Local AI pre-screening for human triple-blind peer review in health sciences arXiv '26 - -
ReVoicer arXiv
ReVoicer: Conversational Voice Annotation for Human-Centered, LLM-Assisted Peer Review arXiv '26 - -
ActReview arXiv
ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation arXiv '26 - -
PaperDoctor arXiv
PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress arXiv '26 - GitHub

Meta-Review & Reviewer Matching

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
Metag arXiv
Metag: A dataset to build agentic meta-reviewing capabilities arXiv '26 - -
AgentReview Website
AgentReview: Exploring Peer Review Dynamics with LLM Agents EMNLP '24 - -
Meta-Review LLMs Website
Meta-Review LLMs NAACL '25 - -
RATE arXiv
RATE: Reviewer Profiling and Annotation-free Training for Expertise Ranking in Peer Review Systems arXiv '26 - -
MERIT arXiv
MERIT: Matching Expertise via Rubric-Informed Training for Reviewer Assignment arXiv '26 - -

Adversarial Attacks & Bias Analysis

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
Raina etal arXiv
Raina etal EMNLP '24 - -
AI Review Lottery arXiv
The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates arXiv '24 - -
Ye etal arXiv
Ye etal arXiv '24 - -
Breaking the Reviewer arXiv
Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks arXiv '25 - -
LLM Reviewer Bias arXiv
LLM Reviewer Bias arXiv '25 - -
Prompt Injection arXiv
Prompt Injection Attacks on LLM Generated Reviews of Scientific Publications arXiv '25 - -
Sahoo etal arXiv
Sahoo etal arXiv '25 - -
Zhou etal arXiv
Zhou etal arXiv '25 - -
Presentation Gaming arXiv
No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions arXiv '26 - -
LLMs Favor LLMs? arXiv
Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review arXiv '26 - -
Gaming AI Reviews arXiv
Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community arXiv '26 - -
Phantom Refs arXiv
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences arXiv '26 - -
Rhetorical Reward-Hacking arXiv
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review arXiv '26 - -
SCOPE-Fuzzer arXiv
Breaking the Illusion of Review Reliability under Static Evaluation: SCOPE Fuzzing for LLM-based Scientific Reviewers arXiv '26 - -

Detection & Policy

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
AI Detection arXiv
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review arXiv '25 - -
AI Use Rejects Website
Major Conference Catches Illicit AI Use — and Rejects Hundreds of Papers Nature '26 - -
Nature AI Survey Website
More than Half of Researchers Now Use AI for Peer Review Nature '26 - -
Policy Enforcement arXiv
Policy Enforcement arXiv '26 - -
Reviewer Feedback Website
What Happens When Reviewers Receive AI Feedback in Their Reviews? CHI '26 - -
AAAI-26 Pilot arXiv
AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot arXiv '26 - -
Reviewer AI Policies arXiv
AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality arXiv '26 - -
ICML LLM Policy Study arXiv
Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026 arXiv '26 - -

Review Consistency and Bias Assessment

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
VERA-RL arXiv
Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper EMNLP '26 - -
Review Survey Website
More than Half of Researchers Now Use AI for Peer Review — often Against Guidance IF '25 - -
Stanford Agentic Website
Stanford Agentic Web '25 - -
ClaimCheck Website
ClaimCheck: How Grounded are LLM Critiques of Scientific Papers? EMNLP '25 - -
REFUTE Website
REFUTE: A Benchmark for Scientific Critique and Epistemic Calibration in Language Models HF '26 Website -
ReViewGraph arXiv
Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates AAAI '26 - -
ReviewAgents arXiv
ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews arXiv '25 - -
ICLR 2025 Study Website
ICLR 2025 Study NMI '26 - -
AI Reviewer Limits arXiv
On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists arXiv '26 - -
PRISM arXiv
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers arXiv '26 - -
LLM-Human Alignment arXiv
How Closely Do LLM Reviews Align with Human Peer Review? arXiv '26 - -
Epistemic Reliability arXiv
Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews arXiv '26 - -
SurveyReview arXiv
SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators arXiv '26 - -
Peerify arXiv
Peerify: Benchmarking Peer-Review Claim Verification arXiv '26 - -
HalluPeer arXiv
HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews arXiv '26 - GitHub
Judging a Review by its Cover arXiv
Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics arXiv '26 - -

7. Rebuttal

Reviewer Comment Analysis

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
ReviewMT arXiv
Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions arXiv '24 - -
ICLR Rebuttal Study arXiv
ICLR Rebuttal Study arXiv '25 - -
RbtAct arXiv
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation arXiv '26 - -
GoodPoint arXiv
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses arXiv '26 - -

Automated Rebuttal Generation

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
ReviewerToo arXiv
ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review arXiv '25 - -
RebuttalAgent arXiv
RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind ICLR '26 - GitHub
Author-in-the-Loop arXiv
Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review ACL '26 - -
DRPG arXiv
DRPG: An Agentic Framework for Academic Rebuttal arXiv '26 - GitHub
Paper2Rebuttal arXiv
Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance arXiv '26 - -
Defend arXiv
Defend: Automated Rebuttals for Peer Review with Minimal Author Guidance arXiv '26 - -

Rebuttal Effectiveness Assessment

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
Re$^2$ arXiv
Re$^2$ arXiv '25 - -
Commitment Checklist arXiv
Commitment Checklist: Auditing Author Commitments in Peer Review arXiv '26 - -
Re$^3$Align arXiv
Re$^3$Align ACL '26 - -
Rebuttals Move arXiv
Rebuttals Move Peer-Review Scores, but Initial-Review Structure Bounds the Movement arXiv '26 - -
Trust AI Reviews arXiv
To Trust or Not to Trust: Authors' Response to AI-based Reviews arXiv '26 - -
AppliedScientist arXiv
AppliedScientist: Automated Scientific Revision Through Iterative AI Reviewing arXiv '26 - -
Edit-Inducing Questions arXiv
Generating Edit-Inducing Questions for AI Research Manuscripts arXiv '26 - -

8. Dissemination (Paper2X)

Paper2Poster

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
P2P Website
P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark ICLR '26 - -
Paper2Poster Website
Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers NeurIPS '25 - GitHub
PosterForest arXiv
PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation arXiv '25 - -
PosterGen arXiv
PosterGen: Aesthetic-Aware Paper-to-Poster Generation via Multi-Agent LLMs arXiv '25 - -
APEX arXiv
APEX: Academic Poster Editing Agentic Expert arXiv '26 - GitHub
PosterOmni arXiv
PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback arXiv '26 - -
Any2Poster arXiv
Any2Poster: Any-Source Poster Generation Across Modalities and Domains arXiv '26 - -
PosterMELD arXiv
PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs arXiv '26 - -
PROS arXiv
Beyond Instruction-Driven Editing: Source-Grounded Problem Discovery with User-Governed Repair for Scientific Posters arXiv '26 - -
PosterVisor arXiv
From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts arXiv '26 - -

Paper2Slides

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
DOC2PPT Website
DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents AAAI '22 - -
PPTAgent arXiv
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides EMNLP '25 - GitHub
AutoPresent arXiv
AutoPresent: Designing Structured Visuals from Scratch CVPR '25 - -
Paper2Slides Website
Paper2Slides: From Paper to Presentation in One Click GitHub '25 - GitHub
Auto-Slides arXiv
Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations arXiv '25 - -
PASS arXiv
PASS: Presentation Automation for Slide Generation and Speech arXiv '25 - -
SlideGen arXiv
SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation arXiv '25 - -
Talk to Your Slides arXiv
Talk to Your Slides: Efficient Slide Editing Agent arXiv '25 - -
SlideTailor arXiv
SlideTailor: Personalized Presentation Slide Generation for Scientific Papers AAAI '26 - GitHub
DeepPresenter arXiv
DeepPresenter: Environment-Grounded Reflection for Agentic Presentation Generation arXiv '26 - GitHub
Office Raccoon Website
Office Raccoon Web '26 - -
X+Slides arXiv
X+Slides: Benchmarking Audience-Conditioned Slide Generation arXiv '26 - -
SeaSlides arXiv
SeaSlides: Semantic Abstraction Layer for Agentic Slide Generation arXiv '26 - -
SLIDEFORGE arXiv
SLIDEFORGE: An LLM Agent for Controllable Editing of Slides as Structured Artifacts arXiv '26 - GitHub
SlideLab arXiv
SlideLab: Audience-Centered Scientific Slide Generation and Evaluation arXiv '26 - -

Paper2Video

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
Preacher Website
Preacher: Paper-to-Video Agentic System ICCV '25 - GitHub
Paper2Video arXiv
Paper2Video: Automatic Video Generation from Scientific Papers arXiv '25 - GitHub
PresentAgent Website
PresentAgent: Multimodal Agent for Presentation Video Generation EMNLP '25 - GitHub
PresentAgent-2 arXiv
PresentAgent-2: Towards Generalist Multimodal Presentation Agents arXiv '26 - -
Paper2Video Talks arXiv
A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks arXiv '26 - -

Paper2Web & Social Media

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
Paper2Web arXiv
Paper2Web: Let's Make Your Paper Alive! arXiv '25 - GitHub
ResearchStudio-Reel arXiv
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog arXiv '26 - -
I-WebGenBench arXiv
I-WebGenBench: Evaluating Interactivity in LLM-Generated Scientific Web Applications arXiv '26 - -
SciForge arXiv
SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery arXiv '26 - -

Fidelity and Adoption Assessment

In chronological order, from the earliest to the latest.

Model Paper Venue Website GitHub
PPTEval arXiv
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides EMNLP '25 - GitHub
PresentQuiz arXiv
Paper2Video: Automatic Video Generation from Scientific Papers arXiv '25 - GitHub
PresentEval Website
PresentAgent: Multimodal Agent for Presentation Video Generation EMNLP '25 - GitHub
Sci. Comm. Correspondence arXiv
Unifying Scientific Communication: Fine-Grain