Maintained by Yingpeng Ma and Yan Ma
A curated list of papers on story generation and storytelling in the era of large language models: long-form fiction, screenplays and drama, games, narrative world models, visual stories, and how to evaluate and co-create them. Every paper comes with a one-line summary, and papers we consider essential reading are marked with 🌟.
Thank you for the stars! Contributions are very welcome: open an issue or PR for missing papers or mistakes. Contact: mayingpeng33 [AT] gmail [DOT] com
📰 News
- [2026-09] 🎉 Major update: 199 papers, a new taxonomy, one-line summaries and 🌟 must-read picks.
- [2026-05] 🔥 Our paper on long-horizon consistency in interactive narratives is accepted to ICML 2026! See it here.
🗺️ Overview
- Beyond Text: interactive drama, games, narrative world models, screenplays, and visual stories.
- Text Stories: planning, coherence, characters, creativity, and training for written stories.
- Evaluation: benchmarks, metrics, and analyses of written stories.
- Co-creation: tools for creators, and studies of how people write with AI.
- Surveys: overviews of the whole field.
Each paper appears exactly once. Human-centered systems and studies go to Co-creation; work on other media goes to Beyond Text (including its evaluation); remaining work on written stories goes to Evaluation or to the Text Stories topic it mainly addresses. Visual work is included only when it operates at the story level (plot, script, shot planning, narrative reasoning), not when it only improves rendering quality or character consistency. Within a section, papers are sorted by year, with 🌟 must-reads first.
Data table for the chart (2026 counts through September; the 4 surveys are not shown)
| Year | Beyond Text | Text Stories | Evaluation | Co-creation | Total |
|---|---|---|---|---|---|
| 2023 | 15 | 6 | 5 | 4 | 30 |
| 2024 | 12 | 11 | 8 | 3 | 34 |
| 2025 | 21 | 15 | 8 | 6 | 50 |
| 2026* | 28 | 29 | 10 | 14 | 81 |
📑 Table of Contents
- 📰 News
- 🗺️ Overview
- 📄 Papers
- 🎭 Beyond Text
- 🎪 Interactive Drama (10)
- 🎲 Games (16)
- 🌐 World Models (6)
- 🎞️ Screenplays (9)
- 🖼️ Visual2Story (14)
- 🌄 Story2Visual (21)
- ✍️ Text Stories
- 🗺️ Planning (12)
- 🧵 Coherence (8)
- 🧑🤝🧑 Characters (8)
- 🎨 Creativity (17)
- 🎯 Training (16)
- 📏 Evaluation
- 🧪 Benchmarks (11)
- 📐 Metrics (7)
- 🔍 Analyses (13)
- 🤝 Co-creation
- 🛠️ Tools (19)
- 👥 User Studies (8)
- 📚 Surveys (4)
- 🎭 Beyond Text
- 🧰 Public Resources
- 🤝 Contributing
- 📝 Citation
📄 Papers
How to read an entry: venue · citation count (refreshed weekly) · 🌟 must-read · title · [paper] · GitHub stars of the official code, when available, followed by authors and a one-line summary.
Venue colors:
🎭 Beyond Text
🎪 Interactive Drama
🌟 Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives [paper]
Yingpeng Ma, Jianhao Yan, Bei-Ning Shi, Karim Kam, Runnan Wang, Xue-Bo Liu, Yulong Chen, Yue Zhang, Derek F. WongIntroduces a 100-environment benchmark on whether LLM narrators keep story commitments under user interventions; even GPT-5.2 survives only 42% after 20 turns.
NARRA-Gym for Evaluating Interactive Narrative Agents [paper]
Yue Huang, Yu-Chen Ma, Jiayi Ye, Wen-Jie Wang, Zi-Peng Ling, Xing Hu, Yuexing Hao, Zi-Chen Chen, Zhangchen Xu, Yun-Hong He, et al.Introduces an executable environment growing emotional seeds into full interactive story episodes, showing fluent LLMs still fail on robustness and personalization.
AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing [paper]
Zhenhua Xu, Dongsheng Chen, Shuo Wang, Jian Li, Chengjie Wang, Meng Han, Ya-Biao WangProposes a multi-agent role-play framework whose scene manager selects speakers, switches scenes, and introduces roles, with training data and a benchmark.
HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics [paper]
[dataset]
Shu-Fan Jiang, Si-Zhou Chen, Chios Chen, Chi Zhang, Xiao-Lei Zhang, Xue-Long LiBuilds HAMLET, a multi-agent framework that turns a topic into a narrative blueprint and performs live embodied theatre with adaptive actor agents.
🌟 Towards Enhanced Immersion and Agency for LLM-based Interactive Drama [paper]
Hongqiu Wu, Weiqi Wu, Tianyang Xu, Jiameng Zhang, Hai ZhaoProposes Playwriting-guided Generation and Plot-based Reflection to improve player immersion and agency in LLM-based interactive drama.
🌟 CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds [paper]
Lei Wang, Jian-Xun Lian, Yi Huang, Yanqi Dai, Haoxuan Li, Xu Chen, Xing Xie, Ji-Rong WenIntroduces a simulation sandbox with character and narrator agents that produces behavior trajectories for fine-grained evaluation of LLM role-playing.
OPEN-THEATRE: An Open-Source Toolkit for LLM-based Interactive Drama [paper]
Tianyang Xu, Hongqiu Wu, Weiqi Wu, Hai ZhaoReleases Open-Theatre, an open-source toolkit for LLM interactive drama with multi-agent architecture and hierarchical retrieval-based memory for coherent long-term behavior.
RolePlot: A Systematic Framework for Evaluating and Enhancing the Plot-Progression Capabilities of Role-Playing Agents [paper]
Pinyi Zhang, Si-Yu An, Lingfeng Qiao, Yi-Fei Yu, Jing-Yang Chen, Jie Wang, Di Yin, Xing Sun, Kai ZhangProposes a plot-progression dataset and method for role-playing agents, detecting an LLM embedding trigger subspace to prompt timely plot advances.
🌟 From Role-Play to Drama-Interaction: An LLM Solution [paper]
Weiqi Wu, Hongqiu Wu, Lai Jiang, Xing-Chen Liu, Jiale Hong, Haizhen Zhao, Min ZhangDefines LLM-based interactive drama and trains a drama LLM using Narrative Chain control, Auto-Drama script synthesis, and Sparse Instruction Tuning.
NarrativePlay: An Automated System for Crafting Visual Worlds in Novels for Role-Playing [paper]
Run-Cong Zhao, Wenjia Zhang, Jiazheng Li, Lixing Zhu, Yanran Li, Yulan He, Lin GuiPresents a demo system that lets users role-play a novel character in LLM-generated narrative environments with generated visuals and speech.
🎲 Games
When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations [paper]
Yuqi Chen, Sixuan Li, Yunfeng Cai, Xueai Li, Kaiwen Yan, Ying LiIntroduces a process benchmark for storytelling in evolving world simulations, finding generation length, canonical consistency, and narrative richness are distinct, competing capacities.
Generating Clue-Driven Investigative Game Narratives with Large Language Models [paper]
Vikram Kumaran, A. Smith, Wookhee Min, Randall Spain, Bradford W. Mott, James C. LesterBuilds an LLM framework that generates solvable clue-driven investigative 3D game episodes around a deductive solution model guiding characters, clues, and dialogue.
IVIE: A Neuro-symbolic Approach to Incremental and Validated Generation of Interactive Fiction Worlds [paper]
Micaela Vaucher, Santiago Silveira, Santiago Góngora, Luis ChiruzzoGenerates playable interactive fiction worlds in four incremental stages, letting LLMs make creative choices while symbolic validation keeps world state coherent.
Guiding, Not Railroading: Design and Evaluation of a Multi-Agent System for Narrative Redirection in Role-playing Games [paper]
Nicolai Hejlesen Jørgensen, Sarmilan Tharmabalan, Ilhan Aslan, Nicolai Brodersen Hansen, Timothy MerrittBuilds a multi-agent RPG game master with a narrative graph and tests six redirection strategies; players prefer in-world redirection over hard denials.
STORY2GAME: Generating (Almost) Everything in an Interactive Fiction Game [paper]
E. Zhou, Shreyas Basavatia, M. Siam, Zexin Chen, Mark O. RiedlBuilds STORY2GAME, which generates a story, populates a world, and writes action code from LLM-derived preconditions and effects for playable interactive fiction.
NarrativeGenie: Generating Narrative Beats and Dynamic Storytelling with Large Language Models [paper]
Vikram Kumaran, Jonathan Rowe, James C. LesterBuilds NarrativeGenie, which turns a designer's story overview into a partially ordered event graph of narrative beats that adapts to player actions.
PANGeA: Procedural Artificial Narrative Using Generative AI for Turn-Based, Role-Playing Video Games [paper]
Stephanie Buongiorno, Lawrence J. Klinkert, Zixin Zhuang, Tanishq Chawla, Corey ClarkBuilds a system with memory, validation, and a Unity plug-in that keeps LLM-generated RPG content consistent with designer rules despite free-form input.
Word2World: Generating Stories and Worlds through Large Language Models [paper]
Muhammad Umair Nasir, Steven James, Julian TogeliusBuilds Word2World, which prompts LLMs to write a story, extract narrative elements, and place tiles to produce playable game worlds without fine-tuning.
Generating Role-Playing Game Quests With GPT Language Models [paper]
Susanna Värtinen, Perttu Hämäläinen, C. GuckelsbergerFine-tunes GPT-2 on a released dataset of 978 RPG quests, finding about one in five generated quest descriptions acceptable to players.
Ontologically Faithful Generation of Non-Player Character Dialogues [paper]
Nathaniel Weir, Ryan Thomas, Randolph D'Amore, Kellie Hill, Benjamin Van Durme, Harsh JhamtaniIntroduces KNUDGE, a dataset from The Outer Worlds requiring lore-faithful, quest-revealing NPC dialogue trees, with supervised and in-context baselines leaving headroom.
🌟 SceneCraft: Automating Interactive Narrative Scene Generation in Digital Games with Large Language Models [paper]
Vikram Kumaran, Jonathan Rowe, Bradford W. Mott, James C. LesterProposes SceneCraft, an LLM framework that automates NPC interaction scenes to unfold authored plot events in narrative-centered games.
🌟 Language as Reality: A Co-Creative Storytelling Game Experience in 1001 Nights using Generative AI [paper]
Yuqian Sun, Zhouyi Li, Ke Fang, Chang Hee Lee, A. AsadipourPresents 1001 Nights, a game where spoken keywords in co-created LLM tales materialize as in-game items, proposing the notion of AI-native games.
🌟 FIREBALL: A Dataset of Dungeons and Dragons Actual-Play with Structured Game State Information [paper]
Andrew Zhu, Karmanya Aggarwal, Alexander H. Feng, Lara J. Martin, Chris Callison-BurchReleases a dataset of about 25,000 real Discord D&D sessions with true game state, showing state information improves LLM game-turn generation.
Location-Aware Adaptation of Augmented Reality Narratives [paper]
Wan-Wan Li, Changyang Li, Minyoung Kim, Haikun Huang, L. YuProposes an optimization approach that assigns real-world locations to AR story events and synthesizes a navigation graph across story branches.
Personalized Quest and Dialogue Generation in Role-Playing Games: A Knowledge Graph- and Language Model-based Approach [paper]
Trevor Ashby, Braden K Webb, G. Knapp, John Searle, Nancy FuldaProposes a player-centered RPG quest and dialogue generator grounding content in a hand-crafted knowledge base and an LLM, approaching hand-crafted quest quality.
I Cast Detect Thoughts: Learning to Converse and Guide with Intents and Theory-of-Mind in Dungeons and Dragons [paper]
Pei Zhou, Andrew Zhu, Jennifer Hu, J. Pujara, Xiang Ren, Chris Callison-Burch, Yejin Choi, Prithviraj AmmanabroluTrains a Dungeon Master model with RL that rewards guidance whose intent matches theory-of-mind predictions of player actions in D&D.
🌐 World Models
FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling [paper]
Jia-Long Zuo, Haotong Zuo, Shiwei Zhang, Xiang Wang, Chen Li, Nong Sang, Chang-Xin Gao, Xiang BaiFrames novel-to-film generation as building a persistent cinematic world model from prose, then rendering long multi-scene films from it.
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World [paper]
Qing Zong, Yue (Sophie) Guo, Mengxi Yang, Yiwen Guo, Yangqiu SongModels interactive literary worlds as long-horizon co-evolution of characters and world state, with an open-schema framework and benchmark.
ReactiveGWM: Steering NPC in Reactive Game World Models [paper]
Zeqing Wang, Dan Chen, Zhaohu Xing, Zizhao Tong, Yinhan Zhang, Xingyi Yang, Ye-Ying JinDecouples player control from NPC behavior in a game world model, so text prompts can steer how NPCs react to the player.
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling [paper]
Yawen Luo, Xiao-Yu Shi, Junhao Zhuang, Yu-Tian Chen, Quande Liu, Xintao Wang, Pengfei Wan, Tian-Fan XueReformulates multi-shot video generation as causal next-shot prediction, letting users steer an unfolding story in real time via streaming prompts.
🌟 Unbounded: A Generative Infinite Game of Character Life Simulation [paper]
Jialu Li, Yuanzhen Li, Neal Wadhwa, Y. Pritch, David E. Jacobs, Michael Rubinstein, Mohit Bansal, Nataniel RuizBuilds a generative infinite game in which players raise an autonomous character in an LLM-driven, image-generated world with open-ended, emergent mechanics.
AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction [paper]
Junhao Cheng, Yu-Ying Ge, Yi-Xiao Ge, Jing Liao, Shan YingTurns anime film characters into playable agents for open-ended life simulation, predicting multimodal game states to keep the generated world consistent.
🎞️ Screenplays
One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems [paper]
Yu-Fei Shi, Wei-Long Yan, Naixuan Huang, Yucheng Chen, Chenyu Zhang, Tao He, Si Yong Yeo, Ming LiBuilds a hierarchical multi-agent pipeline turning a one-sentence idea into a short drama via debate-based scripting, 3D-grounded first frames, and reviewer loops.
Text-to-Stage: Spatial Layouts from Long-form Narratives [paper]
Jefferson Hernandez, Swarnadeep Saha, Chenxi Whitehouse, Sanjeel Parekh, Calvin Murdock, Yuliang Li, W. O. Brimijoin, V. Ithapu, I. AnanthabhotlaIntroduces the task of inferring stage layouts and movements from narrative text, with a dramaturgy-based evaluation suite and rejection-SFT plus GRPO training.
COMIC: Agentic Sketch Comedy Generation [paper]
Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. SeitzBuilds an automated agent population mimicking studio roles to produce sketch comedy videos, using LLM critics aligned with YouTube viewer preferences.
DramaBench: A Six-Dimensional Evaluation Framework for Drama Script Continuation [paper]
Shijian Ma, Yun-Chien Huang, Yan LinIntroduces a drama script continuation benchmark scoring six dimensions via rules and LLM labeling, evaluating eight LLMs on 1,103 scripts.
Beyond Direct Generation: A Decomposed Approach to Well-Crafted Screenwriting with LLMs [paper]
Hang Lei, Shengyi Zong, Zhaoyan Li, Ziren Zhou, Hao Liu, Liang YuDecouples screenplay writing into outline-to-prose then prose-to-screenplay stages with hybrid data synthesis, winning 75% against strong baselines per professional screenwriters.
CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation [paper]
Mingzhe Zheng, Dingjie Song, Guanyu Zhou, Jun You, Jia-Hao Zhan, Xuran Ma, Xin-Yuan Song, Ser-Nam Lim, Qi-Feng Chen, Harry YangIntroduces a movie script benchmark scoring dialogue coherence, character consistency, and plot reasonableness, plus an instruction-based prompting strategy for better scripts.
🌟 IBSEN: Director-Actor Agent Collaboration for Controllable and Interactive Drama Script Generation [paper]
Senyu Han, Lu Chen, Li-Min Lin, Zhen Xu, Kai YuProposes IBSEN, where a director agent steers actor agents and human players toward plot objectives to generate controllable drama scripts.
🌟 HoLLMwood: Unleashing the Creativity of Large Language Models in Screenwriting via Role Playing [paper]
Jing Chen, Xinyu Zhu, Cheng Yang, Chufan Shi, Ya-Dong Xi, Yuxiang Zhang, Junjie Wang, Jiashu Pu, Rongsheng Zhang, Yu-Jiu Yang, et al.Builds HoLLMwood, a screenwriting framework assigning LLMs Writer, Editor, and role-playing Actor roles to enrich characters and plots in generated screenplays.
🌟 Co-Writing Screenplays and Theatre Scripts with Language Models: An Evaluation by Industry Professionals [paper]
Piotr Wojciech Mirowski, K. Mathewson, Jaylen Pittman, Richard EvansBuilds Dramatron, which hierarchically prompts LLMs to co-write scripts and screenplays, evaluated in a study with 15 theatre and film professionals.
🖼️ Visual2Story
Generating Visual Stories with Grounded and Coreferent Characters [paper]
Danyang Liu, Mirella Lapata, Frank KellerPresents a character-centric visual storytelling model trained on VIST enriched with visual and textual coreference chains, plus metrics for character richness.
🌟 StoryLLaVA: Enhancing Visual Storytelling with Multi-Modal Large Language Models [paper]
Li Yang, Zhihao Xiao, Wen-Xin Huang, Xian ZhongProposes a visual storytelling MLLM trained with a topic-driven narrative optimizer for data refinement and preference-based ranked story sampling for alignment.
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? [paper]
M. Gado, Towhid Taliee, M. Memon, Dmitry Ignatov, R. TimofteAdapts large multimodal models to visual storytelling on VIST and advocates reference-free metrics RoViST and GROOVIST over BLEU-style evaluation.
From Panels to Prose: Generating Literary Narratives from Comics [paper]
Ragav Sachdeva, Andrew ZissermanBuilds a system that converts manga into literary prose for visually impaired readers, introducing the Magiv3 comic-understanding model and annotated panel captions.
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition [paper]
Aditya K Surikuchi, Raquel Fernández, Sandro PezzelleProposes a human-likeness metric over visual grounding, coherence, and repetition, finding a small upgraded TAPM rivals LLaVA, yet good stories need more.
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline [paper]
Dingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang, Tiezheng Ge, Bo Zheng, Qin JinIntroduces synchronized video storytelling, generating clip-aligned narrations of fitting length, with the E-SyncVidStory dataset and a storyline-guided VideoNarrator framework.
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling [paper]
E. Wang, Caren Han, Josiah PoonProposes a visual storytelling framework that builds a social-commonsense plot graph from images and derives storylines via weighted shortest paths with Floyd-Warshall.
🌟 Visual Writing Prompts: Character-Grounded Story Generation with Curated Image Sequences [paper]
Xudong Hong, A. Sayeed, K. Mehra, Vera Demberg, B. SchieleIntroduces a dataset of about 2K curated movie-shot sequences with 12K character-grounded crowdsourced stories, plus a coherence-driven character-based generation baseline.
DiffuVST: Narrating Fictional Scenes with Global-History-Guided Denoising Models [paper]
Shengguang Wu, Mei Yuan, Qi SuProposes a non-autoregressive diffusion model that generates visual story narrations for fictional image sequences with bidirectional history guidance, improving speed and diversity.
Sound of Story: Multi-modal Storytelling with Audio [paper]
Jaeyeon Bae, Seokhoon Jeong, Seokun Kang, Namgi Han, Jae-Yon Lee, Hyounghun Kim, Taehwan KimIntroduces Sound of Story, a dataset of 27K stories pairing image-text sequences with background audio, plus cross-modal retrieval and audio generation benchmarks.
GROOViST: A Metric for Grounding Objects in Visual Storytelling [paper]
Aditya K Surikuchi, Sandro Pezzelle, Raquel FernándezProposes a modular, interpretable metric for visual storytelling that measures how well stories are grounded in image entities, handling temporal misalignment.
Visual Storytelling with Question-Answer Plans [paper]
Danyang Liu, Mirella Lapata, Frank KellerProposes visual storytelling that feeds images as a visual prefix to a pretrained language model and plans with question-answer blueprints.
Attractive Storyteller: Stylized Visual Storytelling with Unpaired Text [paper]
Dingyi Yang, Qin JinIntroduces stylized visual storytelling and a memory-augmented multitask model trained with unpaired style text to generate styled stories from photo streams.
Multimodal Event Transformer for Image-guided Story Ending Generation [paper]
Yucheng Zhou, Guodong LongProposes an event-graph reasoning transformer for image-guided story ending generation, with cross-modal fusion, a multimodal injector, and incoherence detection.
🌄 Story2Visual
🌟 MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling [paper]
Qian Wang, Zi-Qi Huang, Ruoxi Jia, Paul E. Debevec, Ning YuProposes a multi-agent pipeline spanning scripting, shot design, character modeling, keyframes, animation, and audio for long-sequence video storytelling.
Better Call CineCrew: Consistent Ultra-Long Narrative-to-Film Generation [paper]
Jiaben Chen, Si Dong, Qinhong Zhou, Raine Ma, Zhi-Yang Dou, Wojciech Matusik, Chuang GanProposes a multi-agent orchestration layer built on FilmDSL, a film-specific language making shot, continuity, and persona constraints explicit for long script-to-video generation.
Learning Long-form Movie Prior via Large Language Models. [paper]
Jin-Heng Xie, Jia-Jun Feng, M. ShouRepresents movies as text and bounding-box or keypoint tokens and curates Storyboard20K, letting LLMs learn movie priors to sample storyboards.
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution [paper]
Mao-Lin Ran, Xiaoyan Lu, Jia-Qi Liu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan ZhangBuilds a deployed storyboarding system that learns directing rules from expert examples, evolves them via attribution feedback, and releases the PROSE dataset.
AniMaster: From Story Texts to Animated Videos via Cinematic Script Generation and Interactive Authoring [paper]
Ruiqi Yu, De-Kun Qian, Jia-Le Xu, Si-Zhe Cheng, Yize Li, Xiang-yang Wu, Zhiguang Zhou, Wei Chen, Yong WangBuilds an authoring tool that expands brief story texts into cinematic scripts, then animated videos, guided by a three-layer design framework.
MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation [paper]
Mu-Yao Wang, Ze-Ke Xie, Yanhao Chen, Lixin Xiu, Hideki NakayamaProposes an agentic story-to-manga framework decomposing creation into planning, grounding, layout, rendering, composition, and lettering for controllable page generation.
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration [paper]
Bo Gao, Chang Liu, Yu-Yang Miao, Siyuan Ma, S. LimProposes a safety-aware multi-agent framework for end-to-end illustrated storybook generation with page-level text-image calibration and global consistency repair.
CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding [paper]
I. Mondal, Yi-Wen Song, Mihir Parmar, Palash Goyal, J. Boyd-Graber, Tomas Pfister, Yale SongProposes a multi-agent storyboarding framework that plans character, background, and location continuity, plus a new long-range consistency benchmark.
LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization [paper]
Chutian Meng, Fan Ma, Chi Zhang, Jiaxu Miao, Yi Yang, Yue-Ting ZhuangProposes a multi-agent story visualization framework that grounds roles, extracts causal chains, and verifies consistency to model visual logic explicitly.
EmoStory: Emotion-Aware Story Generation [paper]
Jing-Yuan Yang, Rucong Chen, Weibin Luo, Hui HuangIntroduces emotion-aware visual story generation and a two-stage framework combining agent-based planning with region-aware generation for emotional, subject-consistent image sequences.
Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning [paper]
Zhengjian Yao, Yong-Zhi Li, Xinyu Gao, Quan Chen, Peng Jiang, Yan LuCombines an MLLM narrative planner with a memory-bank control module for long consistent visual sequences and releases a 330K-image e-commerce storyboard dataset.
MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration [paper]
Wenzhang Sun, Zhenyu Wang, Zhang-Chi Hu, Chun-Feng Wang, Hao Li, Wei ChenProposes a multi-agent plan-execute-verify-revise loop for long audio-visual stories from short prompts, plus a reference-free evaluation protocol.
🌟 SEED-Story: Multimodal Long Story Generation with Large Language Model [paper]
Shuai Yang, Yu-Ying Ge, Yang Li, Yu-Kang Chen, Yixiao Ge, Shan Ying, Ying-Cong ChenProposes SEED-Story, an MLLM generating interleaved text and consistent images for long stories via multimodal attention sinks, and releases the StoryStream dataset.
Generating Storytelling Images with Rich Chains-of-Reasoning [paper]
Xiujie Song, Qi Jia, Shota Watanabe, Xiao-Yi Pang, Ruijie Chen, Mengyue Wu, Ke ZhuDefines storytelling image generation with chains of visual reasoning clues and proposes an LLM-plus-text-to-image pipeline with dedicated evaluation metrics.
From Outline to Detail: An Hierarchical End-to-end Framework for Coherent and Consistent Visual Novel Generation and Assembly [paper]
Yilin Zhang, Yanyan Wei, Zhao Zhang, Jicong Fan, Haijun Zhang, Shui-Cheng YanProposes an outline-guided pipeline that generates and assembles executable visual novels, using vision-LLM self-correction for cross-modal consistency and script validation.
LLMs Behind the Scenes: Enabling Narrative Scene Illustration [paper]
Melissa Roemmele, John Joon Young Chung, Taewook Kim, Yuqian Sun, Alex Calderwood, Max KreminskiUses LLMs to prompt text-to-image models for narrative scene illustration and releases SceneIllustrations, a dataset of pairwise human quality judgments.
AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation [paper]
Haoyuan Shi, Yunxin Li, Xinyu Chen, Long-Yue Wang, Bao-Tian Hu, Min ZhangProposes AniMaker, a multi-agent animation framework using MCTS-driven multi-candidate clip generation and the AniEval evaluator to produce story-coherent long videos from text.
VinaBench: Benchmark for Faithful and Consistent Visual Narratives [paper]
Silin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler, Mengjie Zhao, Hiromi Wakaki, Yuki Mitsufuji, Syrielle Montariol, Antoine BosselutIntroduces a benchmark annotating commonsense and discourse constraints in visual narratives, with metrics for consistency and text alignment of generated image sequences.
MM-StoryAgent: Immersive Narrated Storybook Video Generation with a Multi-Agent Paradigm across Text, Image and Audio [paper]
Xuenan Xu, Jiahao Mei, Chenliang Li, Yuning Wu, Ming Yan, Shaopeng Lai, Ji Zhang, Mengyue WuProposes a multi-agent framework combining LLMs with image, speech, music, and sound tools to generate narrated storybook videos for children.
VISIAR: Empower MLLM for Visual Story Ideation [paper]
Zhaoyang Xia, Somdeb Sarkhel, Md Mehrab Tanjim, Stefano Petrangeli, Ishita Dasgupta, Yuxiao Chen, Jinxuan Xu, Di Liu, Saayan Mitra, Dimitris N. MetaxasIntroduces visual story ideation, arranging visual assets into storylines, with an MLLM framework using a story graph and a VTravel benchmark.
Multimodal Persona Based Generation of Comic Dialogs [paper]
Harsh Agrawal, A. Mishra, Manish Gupta, M. -Introduces multimodal persona-based comic dialogue generation with a 54K-strip dataset and an architecture that generates next-panel dialogues.
✍️ Text Stories
🗺️ Planning
LLM-Driven MCTS for Conditional Story Generation via Logic-Guided Evidence Tree Optimization [paper]
Hongyan Wu, Zhiliang Tian, Zhen Huang, Nankai Lin, Yi-Ping Song, Zhihua Wen, Menglong Lu, Feng Liu, Dongsheng LiProposes a plug-and-play MCTS planner that builds logic-validated evidence chains for retrieval-based conditional story generation to reduce incoherence and thematic drift.
Planning Beyond Text: Graph-based Reasoning for Complex Narrative Generation [paper]
Hanwen Gu, Chao Guo, Junle Wang, Wen-Da Xie, Yi-Sheng LvProposes PLOTTER, which runs an Evaluate-Plan-Revise cycle on event and character graphs to fix causality and structure before generating full narrative text.
BiT-MCTS: A Theme-based Bidirectional MCTS Approach to Chinese Fiction Generation [paper]
Zhaoyi Li, Xu Zhang, Xiaojun WanGenerates Chinese fiction by writing the climax first, then expanding plot backward and forward with bidirectional MCTS inspired by Freytag's Pyramid.
Lightweight Latent Reasoning for Narrative Tasks [paper]
Alexander Gurung, Esmeralda S. Whitammer, Mirella LapataProposes a lightweight reasoning projector producing continuous latent tokens that RL policies toggle, cutting reasoning length on plot-hole detection and chapter generation.
🌟 Generating Long-form Story Using Dynamic Hierarchical Outlining with Memory-Enhancement [paper]
Qianyue Wang, Jinwu Hu, Zhengpin Li, Yufeng Wang, Daiyuan Li, Yu Hu, Mingkui TanProposes DOME, which fuses planning and writing through dynamic hierarchical outlines and uses a memory module to reduce contradictions in long stories.
🌟 Agents' Room: Narrative Generation through Multi-step Collaboration [paper]
Fantine Huot, Reinald Kim Amplayo, J. Palomaki, Alice Shoshana Jakobovits, Elizabeth Clark, Mirella LapataProposes Agents' Room, which splits fiction writing into subtasks for specialized agents, and releases the Tell Me A Story dataset and evaluation.
StoryWriter: A Multi-Agent Framework for Long Story Generation [paper]
Haotian Xia, Hao Peng, Yunjia Qi, Bin Xu, Juan-Zi Li, Hou Lei, Xiaozhi WangProposes a multi-agent long story framework and uses it to build a 6,000-story dataset for fine-tuning Llama3.1-8B and GLM4-9B.
STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation [paper]
Jiaming Li, Yu-Kun Chen, Ziqiang Liu, Minghuan Tan, Lei Zhang, Yunshui Li, Run Luo, Long-Ze Chen, Jing Luo, A. Argha, et al.Proposes a plot-planning approach using SVO-triplet plot nodes plus interacting storyline and narrative entity knowledge graph modules for coherent story generation.
Ex3: Automatic Novel Writing by Extracting, Excelsior and Expanding [paper]
H. Lei, Jiaming Guo, Guanhua He, Xi-Shan Zhang, Rui Zhang, Shaohui Peng, Shaoli Liu, Tianshi ChenProposes Ex3, which extracts structure from raw novels to build instruction data, fine-tunes an LLM, and expands tree-like into arbitrarily long novels.
SWAG: Storytelling With Action Guidance [paper]
Zeeshan Patel, Karim El-Refai, Jonathan Pei, Tianle LiProposes SWAG, framing story writing as search where an auxiliary LLM picks the next action steering the generator toward engaging stories.
Little Red Riding Hood Goes around the Globe: Crosslingual Story Planning and Generation with Large Language Models [paper]
E. Razumovskaia, Joshua Maynez, Annie Louis, Mirella Lapata, Shashi NarayanIntroduces crosslingual story generation with planning and a dataset, finding three-act plans yield more coherent, interesting, controllable stories across languages.
🌟 DOC: Improving Long Story Coherence With Detailed Outline Control [paper]
Kevin Yang, D. Klein, Nanyun Peng, Yuan-Dong TianImproves long-story plot coherence by generating a detailed hierarchical outline and a controller that keeps drafted passages aligned with outline details.
🧵 Coherence
FossilWriter: Learning Hypergraph World Models with Latent Narratives for Creative Story Generation [paper]
Heng Zhang, Yi-Hao Zhong, Lubin Gan, Zhihe Chen, Tianyi Zhang, Jing Liu, Jin HuangGrows a hypergraph world model whose unresolved elements seed latent narratives, improving plot coherence and reducing long-range factual conflicts in story generation.
Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction [paper]
M. Saifullah, Thomas Kornmaier, Taaha Kazi, Vasu Sharma, A. Kanade, Aanand Kumar YadavProposes a writer-memory system pairing a narratology-typed temporal state graph with hybrid retrieval, outperforming Graphiti/Zep and GraphRAG on multi-hop story questions.
ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control [paper]
Jindong Li, Yang Yang, Zihao Liu, Yutao Yue, Meng-Lin YangProposes a training-free scene-by-scene writer that tracks symbolic story states, checks narrative transitions, and uses uncertainty signals to repair inconsistencies.
Octopus: Entropy-Controlled Science Fiction Literature Generation with Persistent Memory-Context Binding [paper]
Xu Wang, Jiaju Kang, Puyu Han, Zeyu Ai, Luqi GongProposes Octopus, combining entropy regulation via narrative divergence thresholds with hierarchical memory of characters, plots, and scientific rules for long sci-fi generation.
SCORE: Story Coherence and Retrieval Enhancement for AI Narratives [paper]
Qiang Yi, Yang-Fan He, Jian-Hui Wang, Xin-Yuan Song, Shi-Yao Qian, Xin-Hang Yuan, Yi Xin, Yi-Jin Wang, Jingqun Tang, Yuchen Li, et al.Proposes SCORE, which tracks key item states and episode summaries and uses retrieval-augmented generation to detect and fix inconsistencies in LLM-generated stories.
MLD-EA: Check and Complete Narrative Coherence by Introducing Emotions and Actions [paper]
Jin-Ming Zhang, Yun-Fei LongProposes MLD-EA, which uses LLMs with emotion and action cues to detect missing logic in narratives and generate sentences that restore coherence.
FACTTRACK: Time-Aware World State Tracking in Story Outlines [paper]
Zhiheng Lyu, Kevin Yang, Lingpeng Kong, Daniel KleinProposes FACTTRACK, which decomposes events into atomic facts with time-aware validity intervals to track world state and detect contradictions in story outlines.
🌟 RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text [paper]
Wangchunshu Zhou, Y. Jiang, Peng Cui, Tiannan Wang, Zhenxin Xiao, Yifan Hou, Ryan Cotterell, Mrinmaya SachanProposes RecurrentGPT, which simulates LSTM-style recurrence with natural-language long- and short-term memories so LLMs can interactively generate arbitrarily long text.
🧑🤝🧑 Characters
ANIMASK: What the Model Contributes to Role Play in Simulated Story Worlds [paper]
Xiu-Cheng Zhang, Zhuo-Ning Xu, Han-Jun Luo, Yankai Chen, Hanan Salam, Xue LiuReplays story worlds from freeze points with and without personas, finding actor LLMs push characters toward cautious, flatter outcomes than canon.
Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents [paper]
Xushuo Tang, Junhe Zhang, Zi-Han Yang, Yi-Fu Tang, Sichao Li, Longbin Lai, Zheng-Yi YangProposes a three-layer, perspective-bounded memory for book-based role-playing agents that prevents characters using unknown facts, with a 4,386-question knowledge-boundary benchmark.
EvoSpark: Endogenous Interactive Agent Societies for Unified Long-Horizon Narrative Evolution [paper]
Shiyu He, Min-Chi Kuang, Mengxian Wang, Bin Hu, Tingxiang GuProposes a multi-agent framework with stratified narrative memory and role-location-plot alignment to sustain coherent, open-ended long-horizon story evolution.
Deriving Character Logic from Storyline as Codified Decision Trees [paper]
Letian Peng, Kun Zhou, Longfei Yun, Yu-Peng Hou, Jingbo ShangInduces executable, interpretable decision trees of validated scene-conditioned behavior rules from narrative data to ground role-playing agents more reliably.
StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models [paper]
Zehao Chen, Rong Pan, Hao-Ran LiProposes bottom-up long-form story generation in which multi-agent sandbox simulation yields emergent events that form coherent stories exceeding 10,000 words.
🌟 BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation [paper]
Yi-Ting Ran, Xintao Wang, Tian Qiu, Jiaqing Liang, Yanghua Xiao, Deqing YangBuilds BookWorld, which simulates multi-agent societies from established novels' characters and worldviews to generate creative stories faithful to the source books.
Steering Narrative Agents Through a Dynamic Cognitive Framework for Guided Emergent Storytelling [paper]
Chen Yang, M. Gross, R. WampflerProposes a cognitive agent framework where tensions between agents' beliefs and ideal worlds drive actions, steering emergent stories toward authored storylines.
🌟 StoryVerse: Towards Co-authoring Dynamic Plot with LLM-based Character Simulation via Narrative Planning [paper]
Yi Wang, Qian Zhou, David LedoProposes StoryVerse, where authors write abstract acts that LLM narrative planning turns into character actions, balancing authorial intent with emergent game plots.
🎨 Creativity
MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing [paper]
Jian-Xiang Ma, Xiaocui Yang, Da-Ling Wang, Yue-Song Hou, Ming-Fu Zhang, Yi-Chen Gao, Jun-Zhao HuangBuilds a story engine encoding McKee's story theory as atomized rules inside an agent harness, improving WritingBench and consistency across four models.
CraftAlign: Feature-Grounded Evaluation and Revision Guidance for AI Stories [paper]
Yang Yang, Boyun Xu, Shaofeng Liang, Yun Han, Zining Zhong, Songning Lai, Kaishen Yuan, Yutao YuePredicts 304 writing features to score stories against human and AI patterns, turning feature shifts into revision guidance that reduces AI flavor.
StorySpark: Module-wise Evolutionary Search for Story Premise Generation [paper]
Yang Yang, Zining Zhong, Qian Cao, Jindong Li, Boyun Xu, Kaishen Yuan, Meng-Lin Yang, Yutao YueProposes module-wise evolutionary search over premise components like persona, event, and twist, producing more original premises that yield better downstream stories.
PlotTwist: A Creative Plot Generation Framework with Small Language Models [paper]
A. Thorat, Ravi Kolla, Jyotin Goel, Madhav Kataria, N. PedanekarProposes a framework where sub-3B models generate premise-conditioned plots using an aspect reward model, DPO-aligned MoE generator, and cross-family jury evaluation.
LLM Review: Enhancing Creative Writing via Blind Peer Review Feedback [paper]
Weiyue Li, Mingxiao Song, Zhenda Shen, Dachuan Zhao, Yunfan Long, Yi Li, Yongce Li, Ruyi Yang, Meng-Yu WangProposes blind peer review among LLM agents that exchange feedback but revise independently, avoiding homogenization, and introduces the SciFi-100 writing dataset.
Frankentext: Stitching random text fragments into long-form narratives [paper]
Chau Minh Pham, Jenna Russell, Dzung Pham, Mohit IyyerProposes generating long narratives by having LLMs stitch mostly verbatim human-written fragments, improving diversity and originality while often evading AI-text detectors.
Avoidance Decoding for Diverse Multi-Branch Story Generation [paper]
Kyeongman Park, Nakyeong Yang, Kyomin JungProposes a decoding strategy that penalizes concept- and narrative-level similarity to earlier outputs, increasing diversity across multiple story branches from one prompt.
A Character-Centric Creative Story Generation via Imagination [paper]
Kyeongman Park, Minbeom Kim, Kyomin JungProposes character-centric story generation that uses text-to-image imagination of story elements and multi-writer persona selection to deepen characters and creativity.
🌟 Collective Critics for Creative Story Generation [paper]
Minwook Bae, Hyounghun KimProposes CritiCS, where a group of LLM critics collectively revise story plans and text to make long stories more creative and expressive.
🌟 MoPS: Modular Story Premise Synthesis for Open-Ended Automatic Story Generation [paper]
Yan Ma, Yu Qiao, Pengfei LiuProposes MoPS, which composes story premises from modular elements like background and persona, yielding more diverse and original premises for story generation.
🌟 Creating Suspenseful Stories: Iterative Planning with Large Language Models [paper]
Kaige Xie, Mark RiedlProposes a zero-shot iterative prompting planner grounded in cognitive-psychology and narratology theories of suspense to generate suspenseful stories with LLMs.
A Conflict-Embedded Narrative Generation Using Commonsense Reasoning [paper]
Youngrok Song, Gunhee Cho, Hyun-Jee Kim, Youngjune Kim, Byung-Chull Bae, Yun-Gyung CheongProposes a neuro-symbolic framework that embeds conflict in stories by using commonsense defeasible inference to weaken causal links toward protagonist goals.
Returning to the Start: Generating Narratives with Related Endpoints [paper]
A. Brei, Chao Zhao, Snigdha ChaturvediProposes RENarGen, which first generates related opening and closing sentences then infills the middle, producing stories with stronger narrative closure.
Improving Pacing in Long-Form Story Planning [paper]
Yichen Wang, Kevin Yang, Xiaoming Liu, Dan KleinProposes CONCOCT, which trains a concreteness evaluator to guide vaguest-first outline expansion and filtering, yielding more consistent pacing in story outlines.
Affective and Dynamic Beam Search for Story Generation [paper]
Tenghao Huang, Ehsan Qasemi, Bangzheng Li, He Wang, Faeze Brahman, Muhao Chen, Snigdha ChaturvediProposes a decoding method combining bandit-driven dynamic beam sizing and affect-intensity reranking to generate stories with more interesting twists.
GROVE: A Retrieval-augmented Complex Story Generation Framework with A Forest of Evidence [paper]
Zhihua Wen, Zhi-Liang Tian, Wei Wu, Yu-Xin Yang, Yanqi Shi, Zhen Huang, Dongsheng LiProposes GROVE, which retrieves human-written story examples and builds an asking-why forest of evidence to add complex, credible plot details.
Narrative Order Aware Story Generation via Bidirectional Pretraining Model with Optimal Transport Reward [paper]
Zhicong Lu, Li Jin, Guang-Luan Xu, Linmei Hu, Nayu Liu, Xiaoyu Li, Xian Sun, Zequn Zhang, Kaiwen WeiProposes a bidirectional pretrained event model with RL using an optimal transport reward to generate coherent stories with flashbacks.
🎯 Training
Retell, Reward, Repeat: Reinforcement Learning for Narrative Theory-Informed Story Generation [paper]
David Y. Liu, Xanthe Muston, A. Joshi, Sebastian Sequoiah-GraysonShows that reinforcement learning from narrative-theory-informed AI feedback (d-RLAIF) yields more diverse, convention-aligned stories than supervised fine-tuning.
POLARIS: Guiding Small Models to Write Long Stories [paper]
Rishanth Rajendhran, Jenna Russell, Mohit Iyyer, J. WietingProposes a GRPO recipe with LLM-judge rewards and injected human reference stories, letting a 9B model write long stories beyond training length.
Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction [paper]
Ze-Han Li, Yu-Tong Zhu, Si-Yang Wu, Honglin Bao, James A. EvansFinds by comparing OLMo checkpoints that post-training compresses thematic, affective, and stylistic variation in fiction, most for professional literary text.
StoryAlign: Evaluating and Training Reward Models for Story Generation [paper]
Hao-Tian Xia, Hao Peng, Yunjia Qi, Xiaozhi Wang, Bin Xu, Lei Hou, Juan-Zi LiIntroduces StoryRMB, a benchmark exposing weak reward models for story preferences, and StoryReward, trained on 100K preference pairs for best-of-n story selection.
UniCreative: Unifying Long-form Logic and Short-form Sparkle via Reference-Free Reinforcement Learning [paper]
Xiao-Long Wei, Zerun Zhu, Simin Niu, Xingyu Zhang, Peiying Yu, Chang Xiao, Yu-Chen Li, Ji-Cheng Yang, Zhejun Zhao, Chong Meng, et al.Proposes a reference-free RL framework with an adaptive constraint-aware generative reward model and ACPO policy optimization, unifying long-form and short-form creative writing.
DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing [paper]
Qian Cao, Yahui Liu, Wei Bi, Yi Zhao, Rui-jie Song, Xiting Wang, Rui-Ming Tang, Guo-Rui Zhou, Han LiProposes an RL framework for creative writing that branches diverse plans in long chain-of-thought and adds a group-aware diversity reward.
Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling [paper]
Zhaoyan Li, Hang Lei, Yuji Wang, Lan Liu, Hao Liu, Liang YuProposes RL for storytelling with a reasoning generative reward model aligned to human creativity judgments and entropy-based reward shaping for training stability.
From Style to Story: A Curriculum Learning Approach for Imitative Novel Generation [paper]
Xueran Han, Yuhan Liu, Mingzhe Li, Wei Liu, Sen Hu, Rui Yan, Zhiqiang Xu, Xiuying ChenIntroduces imitative novel generation and trains WriterAgent via curriculum learning with hierarchical LoRA modules to mimic an author's style, characters, and plots.
Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning [paper]
Jinlong Liu, Mohammed Bahja, Venelin Kovatchev, Mark LeeTrains an authorship-verification style judge and uses it as a GRPO reward to fine-tune an 8B model for writing like classic authors.
RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing [paper]
Jian-Xing Liao, Tian Zhang, Xiao Feng, Yusong Zhang, Rui Yang, Hao-Rui Wang, Bosi Wen, Ziyi Wang, Run-Zhi ShiProposes RL with a dynamically weighted mix of writing-quality and constraint-verification rewards in GRPO, improving both creative quality and instruction following.
🌟 Learning to Reason for Long-Form Story Generation [paper]
Alexander Gurung, Mirella LapataProposes RL for story reasoning via Next-Chapter Prediction, rewarding plans that raise completion likelihood of real book chapters without labeled data.
🌟 Modifying Large Language Model Post-Training for Diverse Creative Writing [paper]
John Joon Young Chung, Vishakh Padmakumar, Melissa Roemmele, Yuqian Sun, Max KreminskiAdds deviation from other same-prompt samples into DPO and ORPO objectives, increasing creative writing output diversity with minimal quality loss.
LiteraryTaste: A Preference Dataset for Creative Writing Personalization [paper]
John Joon Young Chung, Vishakh Padmakumar, Melissa Roemmele, Yi Wang, Yuqian Sun, Tiffany Wang, S. Almeda, Brett A. Halperin, Yuwen Lu, Max KreminskiReleases reading preferences from 60 people over creative text pairs, finding tastes diverge and stated preferences poorly predict revealed ones.
COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes [paper]
Yunwen Li, Shuangshuang Ying, Xingwei Qu, Xin Li, Sheng Jin, Minghao Liu, Zhoufutu Wen, Tianyu Zheng, Xeron Du, Qiguang Chen, et al.Releases 1,665 Chinese creative writing triplets with reverse-engineered prompts and reasoning traces, finding process supervision helps only when mixed with general data.
🌟 Weaver: Foundation Models for Creative Writing [paper]
Tiannan Wang, Jiamin Chen, Qi Jia, Shuai Wang, Ruoyu Fang, Huilin Wang, Zhaowei Gao, Chunzhao Xie, Chuou Xu, Jihong Dai, et al.Introduces Weaver, a 1.8B-34B LLM family pre-trained and aligned for creative and professional writing, with a routing agent balancing quality and cost.
MirrorStories: Reflecting Diversity through Personalized Narrative Generation with Large Language Models [paper]
Sarfaroz Yunusov, Hamza Sidat, Ali EmamiIntroduces MirrorStories, 1,500 LLM-generated stories personalized to reader identity, and finds they engage readers more than generic human or LLM stories.
📏 Evaluation
🧪 Benchmarks
🌟 LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing [paper]
Daniel Fein, S. Russo, Violet Xiang, Kabir Jolly, Rafael Rafailov, Nick HaberIntroduces a benchmark of 2,480 human-labeled story comparisons and 43,827 training pairs for creative writing evaluation, benchmarking LLM judges and reward models.
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs [paper]
Junjie Li, Xinru Guo, Yuhao Wu, Roy Ka-Wei Lee, Hong-Zhi Li, Yutao XieIntroduces a 2,000-prompt benchmark and automated checker for consistency errors in long story generation, analyzing where and which contradictions LLMs make.
ChangJuan: A Comprehensive Benchmark for Book-Length Chinese Story Evaluation [paper]
Dingyi Yang, Mingshuo Wang, Qin JinIntroduces a benchmark of 300 Chinese novels with human ratings and distilled reader viewpoints, plus CLEM, an 8B evaluator for book-length stories.
Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives? [paper]
Karin de Langis, Püren Öncel, Ryan Peters, Andrew Elfenbein, Laura K. Allen, Andreas Schramm, Dongyeop KangFinds LLM internal representations detect incoherent narratives but their ratings do not, and models notice setting violations more than character-trait violations.
WebNovelBench: Placing LLM Novelists on the Web Novel Distribution [paper]
Leon Lin, Jun Zheng, Haidong WangIntroduces a benchmark of 4,000+ Chinese web novels that scores LLM synopsis-to-story outputs on eight dimensions and ranks them against human-authored percentiles.
What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation [paper]
Dingyi Yang, Qin JinIntroduces LongStoryEval, 600 books averaging 121K tokens with reader reviews, compares long-story evaluation methods, and trains NovelCritique, an 8B summary-based evaluator.
Finding Flawed Fictions: Evaluating Complex Reasoning in Language Models via Plot Hole Detection [paper]
Kabir Ahuja, Melanie Sclar, Yulia TsvetkovIntroduces FlawedFictions, a benchmark built by synthesizing plot holes in human stories, to test LLM narrative reasoning via plot hole detection.
Towards A "Novel" Benchmark: Evaluating Literary Fiction with Large Language Models [paper]
Wenqing Wang, Mingqi Gao, Xinyu Hu, Xiaojun WanProposes a ten-metric macro/meso/micro evaluation framework and bilingual annotated fiction dataset, revealing a high-starting, low-ending pattern in LLM-written novels.
CollabStory: Multi-LLM Collaborative Story Generation and Authorship Analysis [paper]
Saranya Venkatraman, N. Tripto, Dongwon LeeIntroduces CollabStory, a dataset of 32k stories co-written by up to five LLMs, with authorship analysis tasks and baselines for multi-LLM writing.
CS4: Measuring the Creativity of Large Language Models Automatically by Controlling the Number of Story-Writing Constraints [paper]
Anirudh Atmakuru, Jatin Nainani, Rohith Siddhartha Reddy Bheemreddy, Anirudh Lakkaraju, Zonghai Yao, Hamed Zamani, Haw-Shiuan ChangIntroduces CS4, a benchmark that measures LLM story creativity by varying the number of prompt constraints to prevent retelling memorized stories.
StoryWars: A Dataset and Instruction Tuning Baselines for Collaborative Story Understanding and Generation [paper]
Yulun Du, Lydia B. ChiltonIntroduces StoryWars, 40k collaborative stories from 9,400 authors forming 101 understanding and generation tasks, with an instruction-tuned InstructStory baseline.
📐 Metrics
Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling [paper]
Pei-Qi Sui, Yu-Tong Zhu, Tianyi Cheng, Peter West, R. So, Hoyt Long, Ari HoltzmanIntroduces 100-Endings, measuring narrative tension by how often repeated ending predictions fail as a story unfolds, and a pipeline that raises tension.
EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation [paper]
Xinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan, Jia-Lin Liu, Chen-Zhuo Zhao, Zhi-Bo Yang, Bin-Bin Yang, Feng XiaoTrains a pairwise story evaluator on self-synthesized, multi-agent-filtered chain-of-thought data and uses it as a reward model to improve story generation.
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing [paper]
Guillermo Marco, Julio Gonzalo, Víctor Fresno-FernándezFinds conflicting evaluations of AI fiction reflect reader differences, clustering 101 annotators into surface-focused and holistic reader profiles via textual feature preferences.
🌟 Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation [paper]
Cyril Chhun, Fabian M. Suchanek, Chloé ClavelStudies LLMs as automatic story evaluators, finding they beat existing metrics at system-level correlation with humans but struggle to explain their ratings.
CHIRON: Rich Character Representations in Long-Form Narratives [paper]
Alexander Gurung, Mirella LapataProposes character sheet representations built by LLM question-answering and entailment-based fact validation, improving masked-character prediction and measuring character-centricity.
Learning Personalized Alignment for Evaluating Open-ended Text Generation [paper]
Danqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang, Andrew Cohen, Lei Li, Yuandong TianProposes PerSE, a LLaMA-2 based evaluator that infers reader preferences from in-context profiles to give personalized, interpretable scores for open-ended generation.
DeltaScore: Evaluating Story Generation with Differentiating Perturbations [paper]
Zhuohan Xie, Miao Li, Trevor Cohn, Jey Han LauProposes DeltaScore, which evaluates story aspects like fluency and interestingness by measuring likelihood changes under aspect-specific perturbations.
🔍 Analyses
CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories [paper]
A. Brei, Abhisheik Sharma, Nicholas Sanaie, Lu Wang, Snigdha ChaturvediCompares characters in LLM-generated and human-written stories along eight narratological dimensions, examining similarity and variety of character types.
StoryScope: Investigating idiosyncrasies in AI fiction [paper]
Jenna Russell, Rishanth Rajendhran, Chau Minh Pham, Mohit Iyyer, J. WietingFinds that discourse-level narrative features alone separate human from AI fiction and attribute AI stories to specific models, independent of stylistic cues.
LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers [paper]
Pei-Qi SuiFinds across 28 LLMs that model story continuations carry much lower information-theoretic uncertainty than human writing, worsened by instruction tuning.
🌟 Echoes in AI: Quantifying Lack of Plot Diversity in LLM Outputs [paper]
Wei-Jia Xu, Nebojsa Jojic, Sudha Rao, C. Brockett, Bill DolanFinds LLM stories from one prompt reuse plot element combinations far more than human stories, and proposes an automatic narrative-level diversity metric.
Evaluating Creative Short Story Generation in Humans and Large Language Models [paper]
Mete Ismayilzada, Claire E. Stevenson, Lonneke van der PlasCompares stories by 60 LLMs and 60 humans, finding LLMs lag in novelty and surprise though non-experts rate LLM stories more creative.
Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs [paper]
Guillermo Marco, Luz Rello, Julio GonzaloFinds a fine-tuned BART-large outscores average human writers on short fiction in human ratings, contrasting its linguistic traits with GPT-3.5 and GPT-4o.
🌟 Are Large Language Models Capable of Generating Human-Level Narratives? [paper]
Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen, Jonathan May, Nan-Yun PengAnalyzes story arcs, turning points, and affect, finding LLM stories are homogeneously positive and lack tension compared with suspenseful, diverse human narratives.
🌟 Art or Artifice? Large Language Models and the False Promise of Creativity [paper]
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, S. Muresan, Chien-Sheng WuProposes the Torrance Test of Creative Writing, finding LLM stories pass 3-10X fewer expert tests than professional stories, and LLM judges misalign.
Pron vs Prompt: Can Large Language Models already Challenge a World-Class Fiction Author at Creative Text Writing? [paper]
Guillermo Marco, Julio Gonzalo, M.Teresa Mateo-Girona, Ramón SantosStages a contest between novelist Patricio Pron and GPT-4, where expert critics judge the LLM far from a top human fiction author.
Measuring Psychological Depth in Language Models [paper]
Fabrice Y. Harel-Canada, Hanyu Zhou, Sreya Muppalla, Zeynep Yildiz, Miryung Kim, Amit Sahai, Nan-Yun PengIntroduces the Psychological Depth Scale for stories' emotional and empathic impact, automates it with LLM personas, finding GPT-4 rivals top Reddit stories.
A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing [paper]
Carlos G'omez-Rodr'iguez, Paul WilliamsCompares LLMs and humans on an unusual comic epic prompt, finding top commercial LLMs match humans on most criteria except creativity.
More human than human: LLM-generated narratives outperform human-LLM interleaved narratives [paper]
Z. Zhao, Sophie Song, Bridget Duah, J. Macbeth, Scott A. Carter, Monica P. Van, N. Bravo, M. Klenk, Kate Sick, Alexandre L. S. FilipowiczFinds through two roughly 500-participant studies that readers prefer purely LLM-generated stories over human-LLM interleaved stories.
The Next Chapter: A Study of Large Language Models in Storytelling [paper]
Zhuohan Xie, Trevor Cohn, Jey Han LauFinds that prompted LLMs write stories rivaling human authors and beating prior generators, though they sometimes replicate real stories.
🤝 Co-creation
🛠️ Tools
Generating Constructive Feedback on Stories via Reinforcement Learning [paper]
Maja Stahl, Timon Ziegenbein, Henning WachsmuthTrains LLMs with GRPO and a multi-component constructiveness reward to give story-specific feedback, finding actionable suggestions drive constructiveness most.
Fabula: Building a Narrative Storytelling Sidekick with the Writers' Community [paper]
Piotr Mirowski, Benjamin D. Wedin, Reinald Kim Amplayo, Rich Galt, Duncan Williams, Rida Qadri, Jaume Sanchez-Elias, Erin Drake-Kajioka, Sian Gooding, Lucía López-Rivilla, et al.Designs and evaluates a narratology-based fiction writing app with 42 writers, probing auto-evaluators, plan-exposing interfaces, and cultural fit of story structures.
Exploring Creator-Centric Methods for LLM-Assisted Interactive Storytelling [paper]
Yue-Lu Li, Siyi Wu, Lu-Jin Zhang, Zhihan Guo, Wenchuan Lu, David YipDesigns CoNoder, a creator-centered LLM prototype for interactive narratives with node-graph editing, ripple-effect analysis, and simulated reader feedback, informed by creator interviews.
Orchid-Creator: An Authoring Tool Supporting LLM-Driven Interactive Narrative Creation [paper]
Zhen Wu, Serkan Kumyol, Zhengyang Ma, T. BraudBuilds an LLM authoring tool representing interactive narratives as card-based story graphs, found easier for structuring than Twine and AI Dungeon.
Plotania: Exploring Transparency Trade-offs in AI Co-Writing Through Virtual Readers and Transparent Attribution [paper]
Yu-Feng Hu, Jinyi Zhang, Ze-Hua Wang, Chun YuBuilds a co-writing system with virtual reader reactions and AI attribution, finding transparency raises awareness but lowers creative agency and AI usage.
Narrix: Remixing Narrative Strategies from Examples for Story Writing [paper]
Chao Zhang, Shunan Guo, Abe Davis, Eunyee KohBuilds a writing tool that highlights narrative strategies in example stories and lets novices apply them to drafts via strategy-steered generation.
NarrativeLoom: Enhancing Creative Storytelling through Multi-Persona Collaborative Improvisation [paper]
Yuxi Ma, Yongqian Peng, Fengyuan Yang, Siyu Zha, Chi Zhang, Zi-Xia Jia, Zilong Zheng, Yixin ZhuBuilds a multi-persona co-creative storytelling system based on blind variation and selective retention; experts rated co-authored stories more creative.
PlayWrite: A Multimodal System for AI Supported Narrative Co-Authoring Through Play in XR [paper]
Esen K. Tutuncu, Qian Zhou, Frederik Brudy, George W. Fitzmaurice, Fraser AndersonBuilds a mixed-reality system where users author stories by manipulating virtual characters and props, which multi-agent AI turns into rearrangeable narrative beats.
DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection [paper]
Yuying Tang, Xinyi Chen, Haotian Li, Xing Xie, Xiaojuan Ma, Huamin QuBuilds a screenplay refinement system whose AI agent first simulates character experience, then evaluates it to give feedback that deepens screenwriters' reflection.
Vidmento: Creating Video Stories through Context-Aware Expansion with Generative Video [paper]
Catherine Yeh, Anh Truong, Mira Dontcheva, Bryan WangBuilds a tool that fills narrative gaps in video stories by generating context-aware clips that blend stylistically and narratively with captured footage.
DiaryPlay: AI-Assisted Creation of Interactive Story Vignettes for Everyday Storytelling [paper]
Jiangnan Xu, Haeseul Cha, Gosu Choi, Gyu-cheol Lee, Y. Yoon, Zucheul Lee, Konstantinos Papangelis, D. Kim, Juho KimBuilds an AI authoring system turning text stories into interactive vignettes, using LLM-controlled divergence to keep NPC behavior within the intended story.
Elsewise: Authoring Open-ended Interactive Narrative with Possibility Space Visualization [paper]
Yi Wang, John Joon Young Chung, Melissa Roemmele, Yuqian Sun, Tiffany Wang, S. Almeda, Brett A. Halperin, Yuwen Lu, Max KreminskiBuilds an authoring tool that visualizes bundled storylines of LLM-driven interactive narratives, helping authors anticipate player-experienced stories in a 12-user study.
Toward Personalizable AI Node Graph Creative Writing Support: Insights on Preferences for Generative AI Features and Information Presentation Across Story Writing Processes [paper]
Hua-Xuan Qin, Guangzhi Zhu, Mingming Fan, Pan HuiStudies a FigJam plugin combining node-graph story structure, LLM audience impersonation, and image/audio generation for personalized story writing and moral reflection.
WhatELSE: Shaping Narrative Spaces at Configurable Level of Abstraction for AI-bridged Interactive Storytelling [paper]
Zhuoran Lu, Qian Zhou, Yi WangBuilds an authoring system deriving narrative possibility spaces from example stories, letting authors bound them and unfold them into game events.
Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols [paper]
John Joon Young Chung, Melissa Roemmele, Max KreminskiBuilds a storytelling system where users steer LLM story text by moving character symbols like toys, via a shared motion-text semantic space.
🌟 CharacterMeet: Supporting Creative Writers' Entire Story Character Construction Processes Through Conversation with LLM-Powered Chatbot Avatars [paper]
Hua-Xuan Qin, Shan Jin, Ze Gao, Mingming Fan, Pan HuiBuilds a system letting writers develop characters by conversing with customizable chatbot avatars; a 14-writer study shows it supports iterative character construction.
Ai.llude: Encouraging Rewriting AI-Generated Text to Support Creative Expression [paper]
David Zhou, S. StermanFinds from 27 writing sessions that deliberately imperfect intermediate AI suggestions encourage writers to rewrite, supporting creative ownership and reflection.
🌟 ID.8: Co-Creating Visual Stories with Generative AI [paper]
Victor Antony, Chien-Ming HuangIntroduces ID.8, an open-source system for co-creating visual stories with generative AI, with a user study highlighting enjoyment and remaining gaps.
Fiction-Writing Mode: An Effective Control for Human-Machine Collaborative Writing [paper]
Wenjie Zhong, Jason Naradowsky, Hiroya Takamura, Ichiro Kobayashi, Yusuke MiyaoAnnotates narrative paragraphs with writing-mode labels and fine-tunes LLMs conditioned on these modes, finding authors prefer mode-controlled suggestions in collaborative fiction writing.
👥 User Studies
The Garden of Forking Prompts: How Users Explore Narrative Space in Story Generation [paper]
Advait Deshmukh, N. Benedict, Melanie Walsh, Maria AntoniakAnalyzes how users iteratively revise story prompts in wild chatbot logs, releasing WildStories and WildEdits and an edit-type framework for benchmarking.
AI Fiction in the Wild [paper]
Neel Gupta, Maria Antoniak, Melanie WalshAnalyzes 500,000 ChatGPT conversations, finding over a third involve fiction generation, dominated by power users favoring fanfiction, erotica, and repetition.
Proactive AI as a Catalyst for Creativity? Balancing Human Agency and AI Contribution in Collaborative Story Writing [paper]
Yiwen Yin, Ming-Ze Wu, R. Huang, X. Tong, Jun Zhou, Chun Yu, Yuanchun ShiWizard-of-Oz study of intrusive versus non-intrusive proactive AI suggestions in story outlining, revealing a creativity-agency trade-off moderated by how inspiring suggestions are.
Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback [paper]
Hannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella LapataIntroduces a task and 1,300 deliberately corrupted stories to evaluate LLM writing feedback, finding models often miss the biggest writing issue.
Understanding Screenwriters' Practices, Attitudes, and Future Expectations in Human-AI Co-Creation [paper]
Yuying Tang, Haotian Li, Minghe Lan, Xiao-Juan Ma, Huamin QuInterviews 23 screenwriters on how they integrate AI across workflow stages and categorizes expected AI roles as actor, audience, expert, and executor.
'It was 80% me, 20% AI': Seeking Authenticity in Co-Writing with Large Language Models [paper]
Angel Hsing-Chi Hwang, Q. Liao, Su Lin Blodgett, Alexandra Olteanu, Adam TrischlerInterviews 19 professional writers and surveys readers on authenticity in AI co-writing, finding personalization should support writer growth beyond text production.
🌟 Social Dynamics of AI Support in Creative Writing [paper]
Katy Ilonka Gero, Tao Long, Lydia B. ChiltonInterviews 20 creative writers to identify what help they want, how they perceive supporters, and values shaping AI-versus-human support choices.
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers [paper]
Tuhin Chakrabarty, Vishakh Padmakumar, Faeze Brahman, S. MuresanStudies 30 writers using an LLM interface based on the cognitive process model, finding LLMs most helpful for translating and reviewing.
📚 Surveys
Narrative Theory-Driven LLM Methods for Automatic Story Generation and Understanding: A Survey [paper]
David Y. Liu, A. Joshi, Paul DawsonSurveys LLM story generation and understanding through narratology, finding generation lags understanding and recommending theory-based metrics over a single quality benchmark.
🌟 A Survey on LLMs for Story Generation [paper]
Maria Teleki, Vedangi Bengali, Xiangjue Dong, Sai Janjur, Haoran Liu, Tian Liu, Cong Wang, Ting-Yiu Liu, Yin Zhang, Frank Shipman, et al.Surveys LLM story generation, organizing work into autonomous generation versus author assistance and comparing methods, datasets, story types, and evaluations.
🌟 What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation [paper]
Dingyi Yang, Qin JinSurveys story evaluation across text-to-text, visual-to-text, and text-to-visual tasks, proposing a taxonomy of human criteria, benchmarks, and automatic metrics.
Open-world Story Generation with Structured Knowledge Enhancement: A Comprehensive Survey [paper]
Yuxin Wang, Jieru Lin, Zhiwei Yu, Wei Hu, Börje F. KarlssonSurveys structured knowledge-enhanced story generation, offering a taxonomy of how knowledge is injected to improve coherence and grounding, plus future directions.
🧰 Public Resources
📦 Datasets
- WritingPrompts: about 300K human-written stories paired with Reddit writing prompts; the most widely used dataset for open-ended story generation.
- VIST: photo sequences paired with human-written stories; the standard dataset for Visual2Story.
- PG-19: full-length books from Project Gutenberg, a common source of long-form fiction.
🏆 Leaderboards
- EQ-Bench Creative Writing: LLM-judged short creative writing leaderboard, updated as new models are released.
- EQ-Bench Longform Creative Writing: multi-chapter story writing, measuring how quality holds up over length.
- LLM Creative Story-Writing Benchmark: tests how well models weave ten required elements into a short story, graded by a panel of LLMs.
- LMArena Creative Writing: crowdsourced human preferences on creative writing prompts.
🏛️ Venues & Workshops
- ICIDS: International Conference on Interactive Digital Storytelling.
- AIIDE: AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment.
- WNU: Workshop on Narrative Understanding, co-located with ACL conferences.
- In2Writing: Workshop on Intelligent and Interactive Writing Assistants.
- Wordplay: When Language Meets Games workshop.
🔗 Related Lists
- Awesome LLM Role-Playing with Persona: role-playing language agents and persona research.
🤝 Contributing
We welcome paper recommendations and corrections. Please read CONTRIBUTING.md for what the list includes and how to format an entry, then open a paper recommendation issue or a pull request.
📝 Citation
If you find this list useful, please consider citing it:
@misc{ma2023awesomestorygeneration,
title = {Awesome-Story-Generation: A Curated List of Papers on Story Generation in the Era of Large Language Models},
author = {Ma, Yingpeng and Ma, Yan},
year = {2023},
howpublished = {\url{https://github.com/yingpengma/Awesome-Story-Generation}}
}