
:sunglasses: Awesome AI Auto-Research
This repository accompanies the survey paper "AI for Auto-Research: Roadmap & User Guide" and tracks papers on AI-assisted and automated scientific research, covering the full research lifecycle.
:robot: AI Auto-Research
We organize the academic research lifecycle as eight interconnected stages grouped into four epistemological phases. Each phase serves a distinct function in producing, scrutinizing, and communicating scientific knowledge.
|
|
 |
Phase 1: Creation |
| Generating novel research ideas, searching and synthesizing literature, running coding experiments, and creating publication-quality tables and figures. This phase spans Idea Generation, Literature Review, Coding & Experiments, and Tables & Figures. |
|
 |
Phase 2: Writing |
| Drafting, editing, and polishing academic manuscripts. AI assistance ranges from semi-automated grammar and citation tools to fully automated paper generation β the most commercially mature yet ethically contested stage. |
|
 |
Phase 3: Validation |
| Automated peer review generation, reviewer-paper matching, review quality assessment, and AI-assisted author rebuttals. This phase covers Peer Review and Rebuttal & Revision. |
|
 |
Phase 4: Dissemination |
| Converting papers into slides, posters, videos, websites, and social media content. Each output format targets a different audience and demands its own design logic and AI tool chain. |
|
|
|
For additional details, kindly refer to our :books: Paper and :earth_asia: Project Page.
:books: Citation
If you find this work helpful for your research, please kindly consider citing our paper:
@article{survey-ai-auto-research,
title = {{AI} for {Auto-Research}: Roadmap \& User Guide},
author = {Kong, Lingdong and Sun, Xian and Chow, Wei and Li, Linfeng and Lin, Kevin Qinghong and Zhang, Xuan Billy
and Wang, Song and Li, Rong and Wu, Qing and Gao, Wei and Wang, Yingshuo and Xie, Shaoyuan
and Liu, Jiachen and Qu, Leigang and Li, Shijie and Ng, Lai Xing and Cottereau, Benoit R.
and Liu, Ziwei and Chua, Tat-Seng and Ooi, Wei Tsang},
journal = {arXiv preprint arXiv:2605.18661},
year = {2026}
}
Table of Contents
 |
1. Idea Generation
LLM Internal Knowledge-Based Generation
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
Chain of Ideas |
 |
|
|
|
| Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents |
arXiv '24 |
- |
 |
|
ResearchAgent |
 |
|
|
|
| ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models |
NAACL '25 |
- |
 |
|
SciMON |
 |
|
|
|
| SciMON: Scientific Inspiration Machines Optimized for Novelty |
ACL '24 |
- |
 |
|
Idea Gen Agent |
 |
|
|
|
| Can LLMs Generate Novel Research Ideas? A Large Scale Human Study with 100+ NLP Researchers |
arXiv '24 |
- |
- |
|
IRIS |
 |
|
|
|
| IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery |
ACL '25 |
- |
 |
|
Spark |
 |
|
|
|
| Spark: A System for Scientifically Creative Idea Generation |
ICCC '25 |
- |
- |
|
Diverse Hypo. Search |
 |
|
|
|
| Towards Diverse Scientific Hypothesis Search with Large Language Models |
arXiv '26 |
- |
- |
|
Tree-of-Ideas |
 |
|
|
|
| Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution |
arXiv '26 |
- |
- |
|
IDEAgent |
 |
|
|
|
| IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation |
arXiv '26 |
- |
- |
|
|
|
|
|
|
External Signal-Driven Generation
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
SGHA |
 |
|
|
|
| SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models |
arXiv '26 |
- |
- |
|
MAIL |
 |
|
|
|
| MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in Chemistry |
arXiv '26 |
- |
- |
|
|
|
|
|
|
MOOSE-Chem |
 |
|
|
|
| MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses |
ICLR '25 |
- |
- |
|
Nova |
 |
|
|
|
| Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas |
arXiv '24 |
- |
- |
|
SciAgents |
 |
|
|
|
| SciAgents: Automating Scientific Discovery through Multi-Agent Intelligent Graph Reasoning |
arXiv '24 |
- |
 |
|
SciPIP |
 |
|
|
|
| SciPIP: An LLM-based Scientific Paper Idea Proposer |
arXiv '24 |
- |
 |
|
IdeaSynth |
 |
|
|
|
| IdeaSynth: Iterative Research Idea Development Through Evolving and Composing Idea Facets with Literature-Grounded Feedback |
CHI '25 |
- |
- |
|
MOOSE-Chem2 |
 |
|
|
|
| MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search |
NeurIPS '25 |
- |
- |
|
HALO |
 |
|
|
|
| HALO: Interactive Co-abductive Reasoning in Scientific Hypothesis Generation |
arXiv '26 |
- |
- |
|
TCA-SIR |
 |
|
|
|
| TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval |
arXiv '26 |
- |
- |
|
ECLAIR |
 |
|
|
|
| ECLAIR: A Causally-Grounded AI Framework for Scientific Discovery in Empirical Software Engineering |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Multi-Agent Collaborative Generation
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
HypoForge |
 |
|
|
|
| HypoForge: A Self-Improving Multi-Agent Framework for Automated Hypothesis Generation and Testing via Scientific Skill Learning |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Combi. Creativity |
 |
|
|
|
| Combi. Creativity |
arXiv '24 |
- |
- |
|
Deep Ideation |
 |
|
|
|
| Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network |
arXiv '25 |
- |
 |
|
VirSci |
 |
|
|
|
| Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System |
ACL '25 |
- |
 |
|
Multi-Agent Dial. |
 |
|
|
|
| Multi-Agent Dial. |
SIGDIAL '25 |
- |
- |
|
Artificial Hivemind |
 |
|
|
|
| Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) |
NeurIPS '25 |
- |
- |
|
Auditable AI Sci. |
 |
|
|
|
| Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents |
arXiv '26 |
- |
- |
|
Diverse Personalized Ideation |
 |
|
|
|
| Diversifying Personalized Research Ideation against AI-Induced Homogenization |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Novelty and Feasibility Assessment
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
RATIO |
 |
|
|
|
| RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature |
arXiv '26 |
- |
- |
|
Lit2Test |
 |
|
|
|
| What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation |
arXiv '26 |
- |
- |
|
Energy Scoring |
 |
|
|
|
| Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking |
arXiv '26 |
- |
- |
|
Think-Probe-Respond |
 |
|
|
|
| Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty |
EMNLP '26 |
- |
- |
|
|
|
|
|
|
IdeaBench |
 |
|
|
|
| LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context |
KDD '25 |
- |
- |
|
LiveIdeaBench |
 |
|
|
|
| LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context |
arXiv '24 |
- |
- |
|
AI Idea Bench 2025 |
 |
|
|
|
| AI Idea Bench 2025: AI Research Idea Generation Benchmark |
arXiv '25 |
- |
 |
|
HeurekaBench |
 |
|
|
|
| HeurekaBench: A Benchmarking Framework for AI Co-scientist |
ICLR '26 |
- |
 |
|
ResearchBench |
 |
|
|
|
| ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition |
ACL '26 |
- |
- |
|
HindSight |
 |
|
|
|
| HindSight: Evaluating LLM-Generated Research Ideas via Future Impact |
arXiv '26 |
- |
- |
|
Rubric Rewards |
 |
|
|
|
| Training AI Co-Scientists Using Rubric Rewards |
arXiv '25 |
- |
- |
|
DeepInnovator |
 |
|
|
|
| DeepInnovator: Triggering the Innovative Capabilities of LLMs |
arXiv '26 |
- |
 |
|
FlowPIE |
 |
|
|
|
| FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration |
arXiv '26 |
- |
- |
|
SoundnessBench |
 |
|
|
|
| SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? |
arXiv '26 |
- |
- |
|
LLM-Judge Novelty |
 |
|
|
|
| On the Limits of LLM-as-Judge for Scientific Novelty Assessment |
arXiv '26 |
- |
- |
|
LigBench |
 |
|
|
|
| LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation |
arXiv '26 |
- |
- |
|
Reconstruction |
 |
|
|
|
| Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies |
arXiv '26 |
- |
- |
|
AgentIdeaBench |
 |
|
|
|
| AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era |
arXiv '26 |
- |
 |
|
IdeaAMBIG |
 |
|
|
|
| IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications |
arXiv '26 |
- |
- |
|
NovGauge |
 |
|
|
|
| NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment |
arXiv '26 |
- |
- |
|
|
|
|
|
|
2. Literature Review & Paper Search
Literature Retrieval
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
ITER |
 |
|
|
|
| ITER: Interaction-Aware Retrieval for Agentic Search |
arXiv '26 |
- |
- |
|
Multi-Aspect Retrieval |
 |
|
|
|
| Can Retrievers Find the Same Paper from Different Aspects? A Multi-Aspect Full-Paper Scientific Retrieval Benchmark |
arXiv '26 |
- |
- |
|
|
|
|
|
|
CiteME |
 |
|
|
|
| CiteME: Can Language Models Accurately Cite Scientific Claims? |
arXiv '24 |
- |
- |
|
LitLLM |
 |
|
|
|
| LitLLM: A Toolkit for Literature Review with Large Language Models |
arXiv '24 |
- |
- |
|
LitSearch |
 |
|
|
|
| LitSearch: A Retrieval Benchmark for Scientific Literature Search |
arXiv '24 |
- |
 |
|
PaperQA2 |
 |
|
|
|
| Language Agents Achieve Superhuman Synthesis of Scientific Knowledge |
arXiv '24 |
- |
 |
|
OpenResearcher |
 |
|
|
|
| OpenResearcher: Unleashing AI for Accelerated Scientific Research |
EMNLP '24 |
- |
- |
|
PaSa |
 |
|
|
|
| PaSa: An LLM Agent for Comprehensive Academic Paper Search |
arXiv '25 |
- |
 |
|
Self-Evolving Retrieval |
 |
|
|
|
| Towards Self-Evolving Agentic Literature Retrieval |
arXiv '26 |
- |
- |
|
MasterSet |
 |
|
|
|
| MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature |
arXiv '26 |
- |
- |
|
Search, Inspect, Fetch |
 |
|
|
|
| Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents |
arXiv '26 |
- |
- |
|
Rubric Reranker |
 |
|
|
|
| Training Documents Reranker with Search Rubrics for Deep Research Agent |
arXiv '26 |
- |
- |
|
Personalized DR Refinement |
 |
|
|
|
| Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Survey & Related Work Generation
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
DAS |
 |
|
|
|
| Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation |
arXiv '26 |
- |
- |
|
Tree-of-Concerns |
 |
|
|
|
| Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique |
EMNLP '26 |
- |
- |
|
|
|
|
|
|
ChatPaper |
 |
|
|
|
| ChatPaper: Use LLM to summarize papers |
GitHub '23 |
- |
 |
|
PaperQA |
 |
|
|
|
| PaperQA: Retrieval-Augmented Generative Agent for Scientific Research |
arXiv '23 |
- |
 |
|
AutoSurvey |
 |
|
|
|
| AutoSurvey: Large Language Models Can Automatically Write Surveys |
arXiv '24 |
- |
 |
|
GPT Researcher |
 |
|
|
|
| GPT Researcher: Autonomous Agent for Comprehensive Online Research |
GitHub '24 |
- |
 |
|
LLMs for Lit. Review |
 |
|
|
|
| LLMs for Lit. Review |
arXiv '24 |
- |
- |
|
STORM |
 |
|
|
|
| Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models |
arXiv '24 |
- |
 |
|
Agentic AutoSurvey |
 |
|
|
|
| Agentic AutoSurvey: Let LLMs Survey LLMs |
arXiv '25 |
- |
- |
|
Citegeist |
 |
|
|
|
| Citegeist: Automated Generation of Related Work Analysis on the arXiv Corpus |
arXiv '25 |
- |
- |
|
IterSurvey |
 |
|
|
|
| IterSurvey: Deep Literature Survey Automation with an Iterative Workflow |
arXiv '25 |
- |
 |
|
LiRA |
 |
|
|
|
| LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation |
arXiv '25 |
- |
- |
|
SurveyForge |
 |
|
|
|
| SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing |
arXiv '25 |
- |
 |
|
SurveyG |
 |
|
|
|
| SurveyG: A Multi-Agent LLM Framework with Hierarchical Citation Graph for Automated Survey Generation |
arXiv '25 |
- |
- |
|
SurveyX |
 |
|
|
|
| SurveyX: Academic Survey Automation via Large Language Models |
arXiv '25 |
- |
- |
|
InteractiveSurvey |
 |
|
|
|
| InteractiveSurvey: An LLM-based Personalized and Interactive Survey Paper Generation System |
arXiv '25 |
- |
 |
|
CiteLLM |
 |
|
|
|
| CiteLLM: An Agentic Platform for Trustworthy Scientific Reference Discovery |
arXiv '26 |
- |
- |
|
DeepSurvey |
 |
|
|
|
| DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation |
arXiv '26 |
- |
- |
|
STRUCTSURVEY |
 |
|
|
|
| STRUCTSURVEY: Structured Agentic Retrieval for Automated Survey Paper Generation |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Deep Research Agents
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
Crase |
 |
|
|
|
| Structurally-bounded Agentic Graph Exploration for Evidence-Grounded Scholarly DeepSearch |
arXiv '26 |
- |
- |
|
DeepWeaver |
 |
|
|
|
| DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering |
arXiv '26 |
- |
- |
|
AgentR |
 |
|
|
|
| AgentR A Stateful and Recovery-Aware Software Architecture for LLM-based Auditable Workflows |
arXiv '26 |
- |
- |
|
|
|
|
|
|
ASReview |
 |
|
|
|
| An Open Source Machine Learning Framework for Efficient and Transparent Systematic Reviews |
Nature MI '21 |
- |
 |
|
CHIME |
 |
|
|
|
| CHIME: LLM-Assisted Hierarchical Organization of Scientific Studies for Literature Review Support |
arXiv '24 |
- |
- |
|
DeepResearch-Agent |
 |
|
|
|
| DeepResearchAgent: A Hierarchical Multi-Agent System for Deep Research |
GitHub '25 |
- |
 |
|
DeerFlow |
 |
|
|
|
| DeerFlow: A Deep Research Framework Orchestrating Sub-Agents, Memory, and Sandboxes |
GitHub '25 |
- |
 |
|
OpenScholar |
 |
|
|
|
| OpenScholar: Synthesizing Scientific Literature with Retrieval-Augmented LMs |
Nature '26 |
- |
- |
|
AutoAgent |
 |
|
|
|
| AutoAgent |
arXiv '25 |
- |
- |
|
Tongyi DeepResearch |
 |
|
|
|
| Tongyi DeepResearch |
GitHub '25 |
- |
 |
|
O-Researcher |
 |
|
|
|
| O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL |
arXiv '26 |
- |
- |
|
OpenResearcher |
 |
|
|
|
| OpenResearcher: Unleashing AI for Accelerated Scientific Research |
arXiv '26 |
- |
 |
|
AREX |
 |
|
|
|
| AREX: Towards a Recursively Self-Improving Agent for Deep Research |
arXiv '26 |
- |
- |
|
Predictive Navigation |
 |
|
|
|
| Deep Research Pretraining via Predictive Navigation |
arXiv '26 |
- |
- |
|
On-Device DR (4B) |
 |
|
|
|
| On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage |
arXiv '26 |
- |
- |
|
Carnot |
 |
|
|
|
| Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries |
VLDB '26 |
- |
- |
|
Marginal Value Est. |
 |
|
|
|
| Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents |
arXiv '26 |
- |
- |
|
Retrieval-Aware Control |
 |
|
|
|
| When Deep Research Agents Stagnate: Enhancing Reasoning with Retrieval-Aware Agent Control |
arXiv '26 |
- |
- |
|
Analogical Deep Research |
 |
|
|
|
| Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis |
arXiv '26 |
- |
- |
|
Plato-Bio |
 |
|
|
|
| Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks |
arXiv '26 |
- |
- |
|
Albilich |
 |
|
|
|
| Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Retrieval and Synthesis Quality Assessment
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
IBIS |
 |
|
|
|
| From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation |
EMNLP '26 |
- |
- |
|
Agent to Blame |
 |
|
|
|
| Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research |
EMNLP '26 |
- |
- |
|
|
|
|
|
|
DeepScholar-Bench |
 |
|
|
|
| DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis |
arXiv '25 |
- |
 |
|
ReportBench |
 |
|
|
|
| ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks |
arXiv '25 |
- |
 |
|
IDRBench |
 |
|
|
|
| IDRBench: Interactive Deep Research Benchmark |
arXiv '26 |
- |
- |
|
ScholarGym |
 |
|
|
|
| ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research |
arXiv '26 |
- |
- |
|
SciNetBench |
 |
|
|
|
| SciNetBench: A Relation-Aware Benchmark for Scientific Literature Retrieval Agents |
arXiv '26 |
- |
- |
|
AutoResearchBench |
 |
|
|
|
| AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery |
arXiv '26 |
- |
- |
|
PaperMind |
 |
|
|
|
| PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs |
arXiv '26 |
- |
- |
|
DRNOISE |
 |
|
|
|
| DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments |
arXiv '26 |
- |
- |
|
HiEviDR-Bench |
 |
|
|
|
| HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research |
arXiv '26 |
- |
- |
|
SciExplore |
 |
|
|
|
| SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration |
arXiv '26 |
- |
- |
|
WANDR |
 |
|
|
|
| WANDR: A Benchmark for Wide and Deep Research |
arXiv '26 |
- |
- |
|
QA-to-DR Bench |
 |
|
|
|
| From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution |
arXiv '26 |
- |
- |
|
PRISMA-LLM |
 |
|
|
|
| PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews |
arXiv '26 |
- |
- |
|
INSPIRE |
 |
|
|
|
| Inspire: Benchmarking Scientific Literature Search for Open Research Problems |
arXiv '26 |
- |
- |
|
CESS |
 |
|
|
|
| Search Shapes Conclusions: Auditing Evidence Selection Bias in Deep Research Agents |
arXiv '26 |
- |
- |
|
ScholarCatalyst |
 |
|
|
|
| ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research |
arXiv '26 |
- |
 |
|
|
|
|
|
|
3. Coding & Experimentation
Code Generation
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
SWE-bench |
 |
|
|
|
| SWE-bench: Can Language Models Resolve Real-World GitHub Issues? |
ICLR '24 |
- |
 |
|
SWE-agent |
 |
|
|
|
| SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering |
arXiv '24 |
- |
 |
|
OpenHands |
 |
|
|
|
| OpenHands: An Open Platform for AI Software Developers as Generalist Agents |
ICLR '25 |
- |
 |
|
SWE-bench Pro |
 |
|
|
|
| SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? |
arXiv '25 |
- |
- |
|
SWE-EVO |
 |
|
|
|
| SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios |
arXiv '25 |
- |
- |
|
|
|
|
|
|
Paper-to-Code
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
ReproAgent |
 |
|
|
|
| ReproAgent: Contract-Guided Paper-to-Code Reproduction |
EMNLP '26 |
- |
- |
|
DeepRepro |
 |
|
|
|
| DeepRepro: State-Aware Subplanning for Paper-to-Code Reproduction in Evolving Repositories |
CIKM '26 |
- |
- |
|
|
|
|
|
|
FunSearch |
 |
|
|
|
| Mathematical Discoveries from Program Search with Large Language Models |
Nature '24 |
- |
 |
|
SciCode |
 |
|
|
|
| SciCode: A Research Coding Benchmark Curated by Scientists |
arXiv '24 |
- |
 |
|
PaperBench |
 |
|
|
|
| PaperBench: Evaluating AI's Ability to Replicate AI Research |
arXiv '25 |
- |
 |
|
PaperCoder |
 |
|
|
|
| Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning |
arXiv '25 |
- |
 |
|
ResearchCodeBench |
 |
|
|
|
| ResearchCodeBench: Benchmarking LLMs on Implementing Novel ML Research Code |
arXiv '25 |
- |
- |
|
SciReplicate-Bench |
 |
|
|
|
| SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers |
arXiv '25 |
- |
 |
|
PaperCompiler |
 |
|
|
|
| PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Experiment Execution & Orchestration
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
Praxist |
 |
|
|
|
| Praxist: From Experimental Artifacts to Solution Lineages |
arXiv '26 |
- |
- |
|
Skill-Based Baselines |
 |
|
|
|
| Can Coding Agents Build Robust Baselines? A Skill-Based Approach for Automating the Medical Imaging Model-Development Pipeline |
MICCAI '26 |
- |
- |
|
|
|
|
|
|
BioPlanner |
 |
|
|
|
| BioPlanner: Automatic Evaluation of LLMs on Protocol Planning |
arXiv '23 |
- |
 |
|
CRISPR-GPT |
 |
|
|
|
| CRISPR-GPT for Agentic Automation of Gene-Editing Experiments |
arXiv '24 |
- |
- |
|
DS-Agent |
 |
|
|
|
| DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning |
arXiv '24 |
- |
 |
|
MLE-Bench |
 |
|
|
|
| MLE-Bench: Evaluating Machine Learning Agents on Machine Learning Engineering |
arXiv '24 |
- |
- |
|
MLAgentBench |
 |
|
|
|
| MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation |
arXiv '24 |
- |
 |
|
MLR-Copilot |
 |
|
|
|
| MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents |
arXiv '24 |
- |
- |
|
AIDE |
 |
|
|
|
| AIDE: AI-Driven Exploration in the Space of Code |
arXiv '25 |
- |
- |
|
AlphaEvolve |
 |
|
|
|
| AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery |
arXiv '25 |
- |
- |
|
AutoReproduce |
 |
|
|
|
| AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage |
arXiv '25 |
- |
 |
|
CURIE |
 |
|
|
|
| Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents |
arXiv '25 |
- |
 |
|
MLGym |
 |
|
|
|
| MLGym: A New Framework and Benchmark for Advancing AI Research Agents |
arXiv '25 |
- |
- |
|
MLR-Bench |
 |
|
|
|
| MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research |
arXiv '25 |
- |
- |
|
Execution-Grounded |
 |
|
|
|
| Towards Execution-Grounded Automated AI Research |
arXiv '26 |
- |
- |
|
Learn to Discover |
 |
|
|
|
| Learning to Discover at Test Time |
arXiv '26 |
- |
- |
|
AutoNumerics |
 |
|
|
|
| AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing |
arXiv '26 |
- |
 |
|
SciNav |
 |
|
|
|
| SciNav: A General Agent Framework for Scientific Coding Tasks |
arXiv '26 |
- |
- |
|
FrontierScience |
 |
|
|
|
| FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks |
arXiv '26 |
- |
- |
|
EvoDS |
 |
|
|
|
| EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management |
arXiv '26 |
- |
- |
|
AutoTTS |
 |
|
|
|
| LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling |
arXiv '26 |
- |
 |
|
AutoScientists |
 |
|
|
|
| AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation |
arXiv '26 |
- |
- |
|
EurekAgent |
 |
|
|
|
| EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery |
arXiv '26 |
- |
- |
|
Experimental Experience Modeling |
 |
|
|
|
| Experimental Experience Modeling for Autonomous Research |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Code Correctness and Reproducibility Assessment
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
Experimental Fidelity |
 |
|
|
|
| Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research |
arXiv '26 |
- |
- |
|
|
|
|
|
|
DiscoveryBench |
 |
|
|
|
| DiscoveryBench: Towards Data-Driven Discovery with Large Language Models |
arXiv '24 |
- |
 |
|
DiscoveryWorld |
 |
|
|
|
| DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents |
arXiv '24 |
- |
 |
|
InfiAgent-DABench |
 |
|
|
|
| InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks |
arXiv '24 |
- |
- |
|
ScienceAgentBench |
 |
|
|
|
| ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery |
arXiv '24 |
- |
- |
|
LAB-Bench |
 |
|
|
|
| Lab-Bench: Measuring Capabilities of Language Models for Biology Research |
arXiv '24 |
- |
 |
|
KernelBench |
 |
|
|
|
| KernelBench: Can LLMs Write Efficient GPU Kernels? |
arXiv '25 |
- |
 |
|
TritonBench |
 |
|
|
|
| TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators |
arXiv '25 |
- |
 |
|
AstaBench |
 |
|
|
|
| AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite |
arXiv '25 |
- |
 |
|
ResearchClawBench |
 |
|
|
|
| Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows |
arXiv '25 |
- |
 |
|
EXP-Bench |
 |
|
|
|
| EXP-Bench: Can AI Conduct AI Research Experiments? |
ICLR '26 |
- |
 |
|
PostTrainBench |
 |
|
|
|
| PostTrainBench: Can LLM Agents Automate LLM Post-Training? |
arXiv '26 |
- |
 |
|
MLReplicate |
 |
|
|
|
| MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility |
arXiv '26 |
- |
- |
|
BeyondSWE |
 |
|
|
|
| BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? |
arXiv '26 |
 |
 |
|
NatureBench |
 |
|
|
|
| NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? |
arXiv '26 |
- |
- |
|
SciCoQA |
 |
|
|
|
| SciCoQA: Quality Assurance for Scientific Paper--Code Alignment |
ACL '26 |
 |
 |
|
Dude |
 |
|
|
|
| Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection |
arXiv '26 |
- |
 |
|
AgentActionBench |
 |
|
|
|
| Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers |
arXiv '26 |
- |
- |
|
AutoDataBench |
 |
|
|
|
| AutoDataBench: A Data-centric Testbed for Accelerating Auto Research |
arXiv '26 |
- |
 |
|
EurekaBench |
 |
|
|
|
| EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights |
arXiv '26 |
- |
- |
|
|
|
|
|
|
4. Tables & Figures
Scientific Figure Generation
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
ChartGPT |
 |
|
|
|
| ChartGPT: Leveraging LLMs to Generate Charts from Abstract Natural Language |
arXiv '23 |
- |
- |
|
MatPlotAgent |
 |
|
|
|
| MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data Visualization |
arXiv '24 |
- |
- |
|
CoDA |
 |
|
|
|
| CoDA: Agentic Systems for Collaborative Data Visualization |
arXiv '25 |
- |
- |
|
PlotGen |
 |
|
|
|
| PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback |
arXiv '25 |
- |
- |
|
VIS-Shepherd |
 |
|
|
|
| VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation |
arXiv '25 |
- |
- |
|
DiagramAgent |
 |
|
|
|
| From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing |
CVPR '25 |
- |
- |
|
StarVector |
 |
|
|
|
| StarVector: Generating Scalable Vector Graphics Code from Images and Text |
CVPR '25 |
- |
- |
|
VisCoder |
 |
|
|
|
| VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation |
EMNLP '25 |
- |
- |
|
AI-Generated Figures |
 |
|
|
|
| AI-Generated Figures |
arXiv '26 |
- |
- |
|
AutoFigure-Edit |
 |
|
|
|
| AutoFigure-Edit: Generating Editable Scientific Illustration |
arXiv '26 |
- |
 |
|
AutoFigure |
 |
|
|
|
| AutoFigure-Edit: Generating Editable Scientific Illustration |
ICLR '26 |
- |
 |
|
PaperBanana |
 |
|
|
|
| PaperBanana: Automating Academic Illustration for AI Scientists |
arXiv '26 |
- |
- |
|
SAIL |
 |
|
|
|
| Setting SAIL: Leveraging Scientist-AI-Loops for Rigorous Visualization Tools |
arXiv '26 |
- |
- |
|
Crafter |
 |
|
|
|
| Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs |
arXiv '26 |
- |
- |
|
DiagramRAG |
 |
|
|
|
| DiagramRAG: A Lightweight Framework to Retrieve Scientific Diagram for Figure Generation |
arXiv '26 |
- |
- |
|
GeoSVG-RL |
 |
|
|
|
| GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation |
arXiv '26 |
- |
- |
|
Can AI Draw Sci. |
 |
|
|
|
| Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models |
arXiv '26 |
- |
- |
|
SciDiagramEdit |
 |
|
|
|
| SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions |
arXiv '26 |
- |
- |
|
GenGA |
 |
|
|
|
| GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers |
arXiv '26 |
- |
- |
|
FigTree |
 |
|
|
|
| Figures as Programs: Recursive Generation of Editable Scientific Figures |
arXiv '26 |
- |
- |
|
EdiTikZ |
 |
|
|
|
| EdiTikZ: Scientific Figure Editing from Revision Trajectories |
arXiv '26 |
- |
 |
|
|
|
|
|
|
Table Understanding & Generation
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
ArxivDIGESTables |
 |
|
|
|
| ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models |
EMNLP '24 |
- |
- |
|
Chain-of-Table |
 |
|
|
|
| Chain-of-Table: Evolving Tables in Reasoning Chain for Table Understanding |
ICLR '24 |
- |
- |
|
ShowTable |
 |
|
|
|
| ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement |
CVPR '26 |
- |
- |
|
Table2LaTeX-RL |
 |
|
|
|
| Table2LaTeX-RL: Converting Table Images to High-Fidelity LaTeX Code Using Reinforced Multimodal Language Models |
arXiv '25 |
- |
- |
|
CSPO |
 |
|
|
|
| CSPO: Alleviating Reward Ambiguity for Structured Table-to-LaTeX Generation |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Mathematical Formulas & TikZ
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
AutomaTikZ |
 |
|
|
|
| AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ |
ICLR '24 |
- |
- |
|
DeTikZify |
 |
|
|
|
| DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ |
NeurIPS '24 |
- |
- |
|
TikZilla |
 |
|
|
|
| TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning |
arXiv '26 |
- |
- |
|
Edit2TikZ |
 |
|
|
|
| Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Visual Fidelity and Scientific Accuracy Assessment
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
StructEval |
 |
|
|
|
| StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs |
TMLR '25 |
 |
 |
|
PlotCraft |
 |
|
|
|
| PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization |
arXiv '25 |
- |
- |
|
TeXpert |
 |
|
|
|
| TeXpert: Multi-Level Benchmark for LaTeX Code Generation |
SDP '25 |
- |
- |
|
AbGen |
 |
|
|
|
| AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research |
ACL '25 |
- |
- |
|
SciFig |
 |
|
|
|
| SciFig: Towards Automating Scientific Figure Generation |
arXiv '26 |
- |
- |
|
SciFlow-Bench |
 |
|
|
|
| SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing |
arXiv '26 |
- |
- |
|
FigureBench |
 |
|
|
|
| AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations |
ICLR '26 |
- |
 |
|
SciFigQual-Bench |
 |
|
|
|
| SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context |
arXiv '26 |
- |
- |
|
SciFigAlign |
 |
|
|
|
| SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence |
arXiv '26 |
- |
- |
|
SciFigPlag-Bench |
 |
|
|
|
| SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection |
arXiv '26 |
- |
- |
|
VLM Blind/Misled |
 |
|
|
|
| How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures |
arXiv '26 |
- |
- |
|
SciFigure2Code |
 |
|
|
|
| SciFigure2Code: An AI-Reconstructed Benchmark for Scientific Figure-to-Code |
arXiv '26 |
- |
- |
|
ReFigBench |
 |
|
|
|
| ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts |
arXiv '26 |
- |
- |
|
|
|
|
|
|
5. Paper Writing
Semi-Automated Writing Assistance
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
CoAuthor |
 |
|
|
|
| CoAuthor: Human-AI Collaborative Writing with Language Models |
arXiv '22 |
- |
- |
|
AI Writing Study |
 |
|
|
|
| AI Writing Study |
AIED '25 |
- |
- |
|
DraftMarks |
 |
|
|
|
| DraftMarks: Enhancing Transparency in Human-AI Co-Writing Through Interactive Skeuomorphic Process Traces |
arXiv '25 |
- |
- |
|
PaperDebugger |
 |
|
|
|
| PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing |
arXiv '25 |
- |
 |
|
ScholarCopilot |
 |
|
|
|
| ScholarCopilot: Training LLMs for Academic Writing with Integrated Citation |
arXiv '25 |
- |
- |
|
XtraGPT |
 |
|
|
|
| XtraGPT: Context-Aware and Controllable Academic Paper Revision |
arXiv '25 |
- |
- |
|
LimAgents |
 |
|
|
|
| Multi-Agent LLMs for Generating Research Limitations |
arXiv '26 |
- |
- |
|
PaperMentor |
 |
|
|
|
| PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers on Overleaf |
arXiv '26 |
- |
- |
|
AutoSupervision |
 |
|
|
|
| AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification |
arXiv '26 |
- |
- |
|
ReasFlow |
 |
|
|
|
| ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Fully Automated Paper Generation
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
CycleResearcher |
 |
|
|
|
| CycleResearcher: Improving Automated Research via Automated Review |
ICLR '25 |
- |
- |
|
Agent Laboratory |
 |
|
|
|
| Agent Laboratory: Using LLM Agents as Research Assistants |
EMNLP '25 |
- |
- |
|
FutureGen |
 |
|
|
|
| FutureGen: A RAG-based Approach to Generate the Future Work of Scientific Article |
arXiv '25 |
- |
- |
|
AI Scientist |
 |
|
|
|
| The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery |
Nature '26 |
- |
 |
|
APRES |
 |
|
|
|
| APRES: An Agentic Paper Revision and Evaluation System |
arXiv '26 |
- |
- |
|
LECTOR |
 |
|
|
|
| LECTOR: Joint Optimization of Scientific Reasoning Graphs and Introduction Generation |
arXiv '26 |
- |
- |
|
RWGBench |
 |
|
|
|
| RWGBench: Evaluating Scholarly Positioning in Related Work Generation |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Societal Analysis
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
AI Writing Adoption |
 |
|
|
|
| AI Writing Adoption |
Nature '26 |
- |
- |
|
Nature AI Survey |
 |
|
|
|
| More than Half of Researchers Now Use AI for Peer Review |
Nature '26 |
- |
- |
|
Denial of Science |
 |
|
|
|
| Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud |
arXiv '26 |
- |
- |
|
AI Slop OSS |
 |
|
|
|
| "AI Slop is DDoSing Open Source": Understanding the Impact of AI-Generated Contributions on Open Source Sustainability |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Writing Quality and AI Detection Assessment
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
Mapping LLM Use |
 |
|
|
|
| Mapping the Increasing Use of LLMs in Scientific Papers |
arXiv '24 |
- |
- |
|
CycleReviewer |
 |
|
|
|
| CycleResearcher: Improving Automated Research via Automated Review |
ICLR '25 |
- |
- |
|
Stanford Agentic |
 |
|
|
|
| Stanford Agentic |
Web '25 |
- |
- |
|
SciIG |
 |
|
|
|
| Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper |
arXiv '25 |
- |
- |
|
Watermarking |
 |
|
|
|
| Detecting LLM-Generated Peer Reviews |
arXiv '25 |
- |
- |
|
PaperWritingBench |
 |
|
|
|
| PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing |
arXiv '26 |
- |
- |
|
CiteTracer |
 |
|
|
|
| Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection |
arXiv '26 |
- |
- |
|
Process Eval |
 |
|
|
|
| Process-Oriented Evaluation of AI-Assisted Scientific Writing |
arXiv '26 |
- |
- |
|
SciSlopBench / SciSlopHarness |
 |
|
|
|
| Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers |
arXiv '26 |
 |
- |
|
|
|
|
|
|
6. Peer Review
Automated Review Generation
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
ChatReviewer |
 |
|
|
|
| ChatReviewer: ChatGPT-based Paper Reviewing and Response Generation |
GitHub '23 |
- |
 |
|
AI-Peer-Review |
 |
|
|
|
| AI-Peer-Review |
GitHub '24 |
- |
 |
|
MARG |
 |
|
|
|
| MARG: Multi-Agent Review Generation for Scientific Papers |
arXiv '24 |
- |
- |
|
Reviewer2 |
 |
|
|
|
| Reviewer2: Optimizing Review Generation Through Prompt Generation |
arXiv '24 |
- |
- |
|
ReviewRL |
 |
|
|
|
| ReviewRL: Towards Automated Scientific Review with RL |
EMNLP '25 |
- |
- |
|
DeepReviewer |
 |
|
|
|
| DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process |
arXiv '25 |
- |
- |
|
OpenReviewer |
 |
|
|
|
| OpenReviewer: A Specialized Large Language Model for Generating Critical Scientific Paper Reviews |
NAACL '25 |
- |
- |
|
REMOR |
 |
|
|
|
| REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning |
arXiv '25 |
- |
- |
|
ScholarPeer |
 |
|
|
|
| ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review |
arXiv '26 |
- |
- |
|
ProReviewer |
 |
|
|
|
| From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent |
arXiv '26 |
- |
- |
|
PeerCheck |
 |
|
|
|
| PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality |
arXiv '26 |
- |
- |
|
Local Pre-Screening |
 |
|
|
|
| Local AI pre-screening for human triple-blind peer review in health sciences |
arXiv '26 |
- |
- |
|
ReVoicer |
 |
|
|
|
| ReVoicer: Conversational Voice Annotation for Human-Centered, LLM-Assisted Peer Review |
arXiv '26 |
- |
- |
|
ActReview |
 |
|
|
|
| ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation |
arXiv '26 |
- |
- |
|
PaperDoctor |
 |
|
|
|
| PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress |
arXiv '26 |
- |
 |
|
|
|
|
|
|
Meta-Review & Reviewer Matching
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
Metag |
 |
|
|
|
| Metag: A dataset to build agentic meta-reviewing capabilities |
arXiv '26 |
- |
- |
|
|
|
|
|
|
AgentReview |
 |
|
|
|
| AgentReview: Exploring Peer Review Dynamics with LLM Agents |
EMNLP '24 |
- |
- |
|
Meta-Review LLMs |
 |
|
|
|
| Meta-Review LLMs |
NAACL '25 |
- |
- |
|
RATE |
 |
|
|
|
| RATE: Reviewer Profiling and Annotation-free Training for Expertise Ranking in Peer Review Systems |
arXiv '26 |
- |
- |
|
MERIT |
 |
|
|
|
| MERIT: Matching Expertise via Rubric-Informed Training for Reviewer Assignment |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Adversarial Attacks & Bias Analysis
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
Raina etal |
 |
|
|
|
| Raina etal |
EMNLP '24 |
- |
- |
|
AI Review Lottery |
 |
|
|
|
| The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates |
arXiv '24 |
- |
- |
|
Ye etal |
 |
|
|
|
| Ye etal |
arXiv '24 |
- |
- |
|
Breaking the Reviewer |
 |
|
|
|
| Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks |
arXiv '25 |
- |
- |
|
LLM Reviewer Bias |
 |
|
|
|
| LLM Reviewer Bias |
arXiv '25 |
- |
- |
|
Prompt Injection |
 |
|
|
|
| Prompt Injection Attacks on LLM Generated Reviews of Scientific Publications |
arXiv '25 |
- |
- |
|
Sahoo etal |
 |
|
|
|
| Sahoo etal |
arXiv '25 |
- |
- |
|
Zhou etal |
 |
|
|
|
| Zhou etal |
arXiv '25 |
- |
- |
|
Presentation Gaming |
 |
|
|
|
| No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions |
arXiv '26 |
- |
- |
|
LLMs Favor LLMs? |
 |
|
|
|
| Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review |
arXiv '26 |
- |
- |
|
Gaming AI Reviews |
 |
|
|
|
| Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community |
arXiv '26 |
- |
- |
|
Phantom Refs |
 |
|
|
|
| Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences |
arXiv '26 |
- |
- |
|
Rhetorical Reward-Hacking |
 |
|
|
|
| How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review |
arXiv '26 |
- |
- |
|
SCOPE-Fuzzer |
 |
|
|
|
| Breaking the Illusion of Review Reliability under Static Evaluation: SCOPE Fuzzing for LLM-based Scientific Reviewers |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Detection & Policy
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
AI Detection |
 |
|
|
|
| Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review |
arXiv '25 |
- |
- |
|
AI Use Rejects |
 |
|
|
|
| Major Conference Catches Illicit AI Use β and Rejects Hundreds of Papers |
Nature '26 |
- |
- |
|
Nature AI Survey |
 |
|
|
|
| More than Half of Researchers Now Use AI for Peer Review |
Nature '26 |
- |
- |
|
Policy Enforcement |
 |
|
|
|
| Policy Enforcement |
arXiv '26 |
- |
- |
|
Reviewer Feedback |
 |
|
|
|
| What Happens When Reviewers Receive AI Feedback in Their Reviews? |
CHI '26 |
- |
- |
|
AAAI-26 Pilot |
 |
|
|
|
| AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot |
arXiv '26 |
- |
- |
|
Reviewer AI Policies |
 |
|
|
|
| AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality |
arXiv '26 |
- |
- |
|
ICML LLM Policy Study |
 |
|
|
|
| Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026 |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Review Consistency and Bias Assessment
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
VERA-RL |
 |
|
|
|
| Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper |
EMNLP '26 |
- |
- |
|
|
|
|
|
|
Review Survey |
 |
|
|
|
| More than Half of Researchers Now Use AI for Peer Review β often Against Guidance |
IF '25 |
- |
- |
|
Stanford Agentic |
 |
|
|
|
| Stanford Agentic |
Web '25 |
- |
- |
|
ClaimCheck |
 |
|
|
|
| ClaimCheck: How Grounded are LLM Critiques of Scientific Papers? |
EMNLP '25 |
- |
- |
|
REFUTE |
 |
|
|
|
| REFUTE: A Benchmark for Scientific Critique and Epistemic Calibration in Language Models |
HF '26 |
Website |
- |
|
ReViewGraph |
 |
|
|
|
| Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates |
AAAI '26 |
- |
- |
|
ReviewAgents |
 |
|
|
|
| ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews |
arXiv '25 |
- |
- |
|
ICLR 2025 Study |
 |
|
|
|
| ICLR 2025 Study |
NMI '26 |
- |
- |
|
AI Reviewer Limits |
 |
|
|
|
| On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists |
arXiv '26 |
- |
- |
|
PRISM |
 |
|
|
|
| PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers |
arXiv '26 |
- |
- |
|
LLM-Human Alignment |
 |
|
|
|
| How Closely Do LLM Reviews Align with Human Peer Review? |
arXiv '26 |
- |
- |
|
Epistemic Reliability |
 |
|
|
|
| Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews |
arXiv '26 |
- |
- |
|
SurveyReview |
 |
|
|
|
| SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators |
arXiv '26 |
- |
- |
|
Peerify |
 |
|
|
|
| Peerify: Benchmarking Peer-Review Claim Verification |
arXiv '26 |
- |
- |
|
HalluPeer |
 |
|
|
|
| HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews |
arXiv '26 |
- |
 |
|
Judging a Review by its Cover |
 |
|
|
|
| Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics |
arXiv '26 |
- |
- |
|
|
|
|
|
|
7. Rebuttal
Reviewer Comment Analysis
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
ReviewMT |
 |
|
|
|
| Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions |
arXiv '24 |
- |
- |
|
ICLR Rebuttal Study |
 |
|
|
|
| ICLR Rebuttal Study |
arXiv '25 |
- |
- |
|
RbtAct |
 |
|
|
|
| RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation |
arXiv '26 |
- |
- |
|
GoodPoint |
 |
|
|
|
| GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Automated Rebuttal Generation
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
ReviewerToo |
 |
|
|
|
| ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review |
arXiv '25 |
- |
- |
|
RebuttalAgent |
 |
|
|
|
| RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind |
ICLR '26 |
- |
 |
|
Author-in-the-Loop |
 |
|
|
|
| Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review |
ACL '26 |
- |
- |
|
DRPG |
 |
|
|
|
| DRPG: An Agentic Framework for Academic Rebuttal |
arXiv '26 |
- |
 |
|
Paper2Rebuttal |
 |
|
|
|
| Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance |
arXiv '26 |
- |
- |
|
Defend |
 |
|
|
|
| Defend: Automated Rebuttals for Peer Review with Minimal Author Guidance |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Rebuttal Effectiveness Assessment
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
Re$^2$ |
 |
|
|
|
| Re$^2$ |
arXiv '25 |
- |
- |
|
Commitment Checklist |
 |
|
|
|
| Commitment Checklist: Auditing Author Commitments in Peer Review |
arXiv '26 |
- |
- |
|
Re$^3$Align |
 |
|
|
|
| Re$^3$Align |
ACL '26 |
- |
- |
|
Rebuttals Move |
 |
|
|
|
| Rebuttals Move Peer-Review Scores, but Initial-Review Structure Bounds the Movement |
arXiv '26 |
- |
- |
|
Trust AI Reviews |
 |
|
|
|
| To Trust or Not to Trust: Authors' Response to AI-based Reviews |
arXiv '26 |
- |
- |
|
AppliedScientist |
 |
|
|
|
| AppliedScientist: Automated Scientific Revision Through Iterative AI Reviewing |
arXiv '26 |
- |
- |
|
Edit-Inducing Questions |
 |
|
|
|
| Generating Edit-Inducing Questions for AI Research Manuscripts |
arXiv '26 |
- |
- |
|
|
|
|
|
|
8. Dissemination (Paper2X)
Paper2Poster
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
P2P |
 |
|
|
|
| P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark |
ICLR '26 |
- |
- |
|
Paper2Poster |
 |
|
|
|
| Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers |
NeurIPS '25 |
- |
 |
|
PosterForest |
 |
|
|
|
| PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation |
arXiv '25 |
- |
- |
|
PosterGen |
 |
|
|
|
| PosterGen: Aesthetic-Aware Paper-to-Poster Generation via Multi-Agent LLMs |
arXiv '25 |
- |
- |
|
APEX |
 |
|
|
|
| APEX: Academic Poster Editing Agentic Expert |
arXiv '26 |
- |
 |
|
PosterOmni |
 |
|
|
|
| PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback |
arXiv '26 |
- |
- |
|
Any2Poster |
 |
|
|
|
| Any2Poster: Any-Source Poster Generation Across Modalities and Domains |
arXiv '26 |
- |
- |
|
PosterMELD |
 |
|
|
|
| PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs |
arXiv '26 |
- |
- |
|
PROS |
 |
|
|
|
| Beyond Instruction-Driven Editing: Source-Grounded Problem Discovery with User-Governed Repair for Scientific Posters |
arXiv '26 |
- |
- |
|
PosterVisor |
 |
|
|
|
| From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Paper2Slides
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
DOC2PPT |
 |
|
|
|
| DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents |
AAAI '22 |
- |
- |
|
PPTAgent |
 |
|
|
|
| PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides |
EMNLP '25 |
- |
 |
|
AutoPresent |
 |
|
|
|
| AutoPresent: Designing Structured Visuals from Scratch |
CVPR '25 |
- |
- |
|
Paper2Slides |
 |
|
|
|
| Paper2Slides: From Paper to Presentation in One Click |
GitHub '25 |
- |
 |
|
Auto-Slides |
 |
|
|
|
| Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations |
arXiv '25 |
- |
- |
|
PASS |
 |
|
|
|
| PASS: Presentation Automation for Slide Generation and Speech |
arXiv '25 |
- |
- |
|
SlideGen |
 |
|
|
|
| SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation |
arXiv '25 |
- |
- |
|
Talk to Your Slides |
 |
|
|
|
| Talk to Your Slides: Efficient Slide Editing Agent |
arXiv '25 |
- |
- |
|
SlideTailor |
 |
|
|
|
| SlideTailor: Personalized Presentation Slide Generation for Scientific Papers |
AAAI '26 |
- |
 |
|
DeepPresenter |
 |
|
|
|
| DeepPresenter: Environment-Grounded Reflection for Agentic Presentation Generation |
arXiv '26 |
- |
 |
|
Office Raccoon |
 |
|
|
|
| Office Raccoon |
Web '26 |
- |
- |
|
X+Slides |
 |
|
|
|
| X+Slides: Benchmarking Audience-Conditioned Slide Generation |
arXiv '26 |
- |
- |
|
SeaSlides |
 |
|
|
|
| SeaSlides: Semantic Abstraction Layer for Agentic Slide Generation |
arXiv '26 |
- |
- |
|
SLIDEFORGE |
 |
|
|
|
| SLIDEFORGE: An LLM Agent for Controllable Editing of Slides as Structured Artifacts |
arXiv '26 |
- |
 |
|
SlideLab |
 |
|
|
|
| SlideLab: Audience-Centered Scientific Slide Generation and Evaluation |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Paper2Video
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
Preacher |
 |
|
|
|
| Preacher: Paper-to-Video Agentic System |
ICCV '25 |
- |
 |
|
Paper2Video |
 |
|
|
|
| Paper2Video: Automatic Video Generation from Scientific Papers |
arXiv '25 |
- |
 |
|
PresentAgent |
 |
|
|
|
| PresentAgent: Multimodal Agent for Presentation Video Generation |
EMNLP '25 |
- |
 |
|
PresentAgent-2 |
 |
|
|
|
| PresentAgent-2: Towards Generalist Multimodal Presentation Agents |
arXiv '26 |
- |
- |
|
Paper2Video Talks |
 |
|
|
|
| A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Paper2Web & Social Media
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
Paper2Web |
 |
|
|
|
| Paper2Web: Let's Make Your Paper Alive! |
arXiv '25 |
- |
 |
|
ResearchStudio-Reel |
 |
|
|
|
| ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog |
arXiv '26 |
- |
- |
|
I-WebGenBench |
 |
|
|
|
| I-WebGenBench: Evaluating Interactivity in LLM-Generated Scientific Web Applications |
arXiv '26 |
- |
- |
|
SciForge |
 |
|
|
|
| SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery |
arXiv '26 |
- |
- |
|
|
|
|
|
|
Fidelity and Adoption Assessment
In chronological order, from the earliest to the latest.
| Model |
Paper |
Venue |
Website |
GitHub |
|
|
|
|
|
PPTEval |
 |
|
|
|
| PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides |
EMNLP '25 |
- |
 |
|
PresentQuiz |
 |
|
|
|
| Paper2Video: Automatic Video Generation from Scientific Papers |
arXiv '25 |
- |
 |
|
PresentEval |
 |
|
|
|
| PresentAgent: Multimodal Agent for Presentation Video Generation |
EMNLP '25 |
- |
 |
|
Sci. Comm. Correspondence |
 |
|
|
|
| Unifying Scientific Communication: Fine-Grain |
|
|
|
|