Awesome GUI Agent Paper List
A curated list of 722 research papers on GUI agents — models, frameworks, benchmarks, datasets, and more — spanning topics like GUI grounding, planning, memory, safety, and reinforcement learning.
🌐 Read this list on the web
The website is the recommended way to read this list:
→ osu-nlp-group.github.io/GUI-Agents-Paper-List
What it adds over the raw markdown:
- Full-text search across titles, authors, institutions, TLDRs, and keywords
- Multi-axis filtering — environment, keyword (AND/OR), author, institution, year, venue — with shareable URLs
- Per-paper detail pages with the full TLDR, all keywords, and related papers
- One-click BibTeX copy — LaTeX-paste-ready (
@miscfor arXiv-only,@inproceedings/@articlefor venue papers). Auto-generated, so please verify before citing. - Interactive stats — quarterly publication trend by environment, top keywords / institutions / authors / venues
- Warm-paper light theme + dark theme, keyboard shortcuts (
/,j/k,Esc,?), no tracking
The structured store papers.yaml (and adjacent.yaml) is the source of truth — everything on the website and the README is generated from it.


Browse by Environment
🌐 Web (278) · 🖥️ Desktop (178) · 📱 Mobile (221) · 🖼️ General GUI (153)
Browse by Keyword
benchmark (223) · dataset (109) · framework (73) · reinforcement learning (71) · model (63)
GUI grounding (59) · safety (44) · security (37) · OSWorld (30) · long-horizon tasks (24)
WebArena (23) · world model (21) · memory (20) · training-free (19) · reward model (19)
AndroidWorld (16) · GRPO (15) · prompt injection (15) · planning (13) · survey (11)
Browse by Author
Wei Liu (23) · Jian Luan (23) · Pengzhi Gao (17) · Zhuosheng Zhang (14) · Graham Neubig (14)
Yu Su (14) · Huan Sun (14) · Zhengxi Lu (13) · Shuyan Zhou (12) · Mike Zheng Shou (12)
Fei Tang (11) · Tao Yu (11) · Boyuan Zheng (11) · Yongliang Shen (10) · Jie Tang (10)
Tianbao Xie (10) · Qiushi Sun (10) · Yuanchun Li (10) · Kevin Qinghong Lin (10) · Yuxiang Chai (10)
Contributing
We welcome contributions from the community!
- Missing a paper? Open an issue with the paper title, link, and any relevant details — we'll add it.
- Want to add papers yourself? Edit
papers.yaml, runbash scripts/update_repo.sh, then submit the regenerated diff. See CLAUDE.md for the YAML schema and local update workflow. - Spotted an error? Feel free to open an issue or PR to correct any paper metadata (authors, dates, institutions, etc.).
Recent Papers (from most recent to oldest)
This README shows the 500 most recent papers. See
papers.yamlfor the full structured source — including BibTeX, OpenReview / publisher / homepage / code / dataset links, and thebibtex_confirmedflag. For non-canonical adjacent papers seeadjacent.yaml.
-
Securing Computer-Use Agents Against Branch Steering Attacks
- Giulio Zingrillo, Hanna Foerster, Ilia Shumailov, Yiren Zhao, Robert Mullins
- 🏛️ Institutions: ETH, University of Cambridge, AI Sequrity Company, Imperial College London
- 📅 Date: October 02, 2026
- 📑 Publisher: NeurIPS 2026 Workshop on Agents in the Wild
- 💻 Env: [Web]
- 🔑 Key: [STEER-Bench], [COBRA], [benchmark], [security], [indirect prompt injection], [branch steering]
- 📖 TLDR: Studies branch steering attacks, in which untrusted page content pushes a computer-use agent down a hazardous but pre-approved branch of a Dual-LLM plan without injecting explicit instructions, and introduces STEER-Bench (101 tasks across 9 domains) to measure them. Proposes COBRA, which pairs trusted branching plans with ahead-of-time capability constraints on the parameters and destinations each branch may use; it reduces attack success on STEER-Bench to 0% while retaining 97% benign utility.
-
- Jiangang Han
- 🏛️ Institutions: Independent Researcher
- 📅 Date: October 02, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [WebFovea], [WebRetriever Challenge], [vision-based agent], [harness], [failure analysis]
- 📖 TLDR: Technical report of WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 on live websites. Many failures it observed occur in the harness between the model and the browser (action parsing, action execution, result reporting, and observation) rather than in model reasoning; hardening each stage raised the official hidden-set score from 31.0 to 57.0 with the same model.
-
- Yulong Ming, Jie Xu, Zihan Wu, Xiaohua Jia
- 🏛️ Institutions: CityU
- 📅 Date: October 02, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [PACE], [efficiency], [workflow compilation], [token cost]
- 📖 TLDR: Studies when it pays to compile a GUI procedure that an agent executes repeatedly into a program, given uncertain compilation cost and unknown future reuse. Proposes PACE, which combines a measurement protocol for payback counts with an online compilation algorithm that bounds total cost to at most 1+ε times that of running every task with the agent.
-
Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories
- Chengguang Gan, Zimeng He, Yoshihiro Tsujii, Ken-ichiro Kobayashi, Hiroki Itoh, Kotaro Funakoshi
- 🏛️ Institutions: Techtouch, Institute of Science Tokyo
- 📅 Date: October 01, 2026
- 📑 Publisher: NeurIPS 2026 Workshop on Who Verifies the Agents? Toward Reliable Agent Development
- 💻 Env: [Web]
- 🔑 Key: [WebArena-Lite], [evaluation], [human review], [trajectory evaluation], [failure analysis], [memory]
- 📖 TLDR: Audits all 165 WebArena-Lite tasks under six evaluation conditions with human review of outcomes and trajectories. Human review recovers successes the automatic evaluator missed, a failure analysis of 102 trajectories identifies recurring error patterns, and a memory-and-analysis mechanism and guide text improve success.
-
AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
- Cheng Yang, Yifan Wu, Yutao Huang, Zhaohua Zhang, Beiduo Chen, Muxi Chen, Chenchen Zhao, Hexuan Deng, Haolin Yang, Geyuan Zhu, Sa Zhu, Jianhuan Zhuo, Qiuyong Xiao, Jianhao Ruan, Yiran Peng, Jiayi Zhang, Tian Ye, Xinlei Yu, Tianwen Jiang, Jihong Zhang, Yuyu Luo
- 🏛️ Institutions: Tencent Hunyuan
- 📅 Date: October 01, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [AutoGUIWorld], [model], [world model], [data generation], [trajectory synthesis], [OSWorld], [ScienceBoard]
- 📖 TLDR: AutoGUIWorld synthesizes GUI interaction trajectories without running the corresponding software: a planner specifies atomic actions and their intended visual consequences, and an image generator iteratively edits the current screenshot to produce the next observation. The 79,266 step-level samples across Ubuntu, Windows, macOS, and Chrome raise Qwen3.5-35B-A3B from 33.0% to 40.8% on OSWorld and from 14.0% to 32.2% on ScienceBoard.
-
DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents
- A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys, Peter W. J. Staar
- 🏛️ Institutions: ETH, IBM Research, Microsoft
- 📅 Date: October 01, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [DeskForge], [dataset], [model], [GUI grounding], [environment synthesis], [long-horizon tasks]
- 📖 TLDR: DeskForge is a controllable desktop environment that composes real applications under varied states, window layouts, appearances, and resolutions, and fuses screenshots, accessibility trees, and window geometry into dense element annotations. The resulting DeskForge-1M corpus (1.2M annotated observations) improves four fine-tuned vision-language models on five GUI grounding benchmarks and on the long-horizon WebArena-Infinity and OpenApps tasks.
-
GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
- Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang, Min Zhang, Shipei Zeng, Zhongxiang Dai
- 🏛️ Institutions: CUHK-Shenzhen, Tianjin University, HIT-Shenzhen, East China Normal University, Shenzhen Research Institute of Big Data
- 📅 Date: October 01, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [GUI-HARVEST], [harness optimization], [self-improvement], [OSWorld], [WindowsAgentArena]
- 📖 TLDR: GUI-HARVEST automatically optimizes the executable harness around a frozen GUI model by aligning actions with before-and-after screenshots, treating repeated runs of a task as one evidence unit, and mapping recurring failure patterns to bounded source-code edits. It yields held-out gains on OSWorld-Verified across six backbones and transfers a frozen harness to WindowsAgentArena.
-
Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
- Beining Wu, Zihao Ding, Jun Huang
- 🏛️ Institutions: South Dakota State University
- 📅 Date: October 01, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [component routing], [self-improvement], [memory], [experience]
- 📖 TLDR: Studies where a self-improving GUI agent should put its experience, in the weights or in the prompt context, by splitting trajectories into locators, procedures, state facts, and lessons. Locators and lessons are better placed in the weights and procedures and state facts in the context; a rule based on recurrence and state-conditionality routes components accordingly and beats every whole-trajectory baseline.
-
Action Conditioned Bisimulation For GUI Agent Memory
- Hongbo Zhang, Liuyang Song, Quanquan Li, Daqian Yang, Yan Wen, Zhengtao Yao
- 🏛️ Institutions: PKU, East China Normal University, USC
- 📅 Date: September 30, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [bisimulation], [memory], [state abstraction], [MiniWoB], [training-free]
- 📖 TLDR: Defines when a web agent's memory should treat two pages as the same state using an action-conditioned bisimulation over the empirical state graph a frozen agent builds as it acts: two states merge only when their shared actions lead to agreeing outcomes. Nothing is trained, and replacing an outcome-value memory's merge rule with it raises success on MiniWoB++.
-
ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
- Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen
- 🏛️ Institutions: ZJU, Ant Group
- 📅 Date: September 30, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [ComputerSD], [self-distillation], [online reinforcement learning], [GRPO], [OSWorld]
- 📖 TLDR: ComputerSD is an online self-distillation method that turns real-time feedback from executed GUI transitions into token-level guidance for computer-use agents. A fine-tuned GUI analyzer supplies guidance and a step-level value score, and training jointly optimizes on-policy self-distillation and trajectory-level GRPO, outperforming outcome-only GRPO on OSWorld-Verified.
-
Learning Reliable GUI Agents under Imperfect Priors
- Bo Han, Qianyi Wang, Shuai Liu, Xiong Zifan, Changqiao Wu, Yuanfa Li, Pengzhi Gao, Wei Liu, Jian Luan, Heng Qu, Yunpeng Song, Zhongmin Cai
- 🏛️ Institutions: XJTU, Xiaomi
- 📅 Date: September 30, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [imperfect priors], [exploration], [data synthesis], [robustness], [UI state-transition graph]
- 📖 TLDR: Argues that GUI agents should act correctly under imperfect, drifting priors rather than pursue perfect app knowledge. Structured exploration builds a UI state-transition graph and synthesizes task-trajectory pairs without human annotation, and a noise-aware training strategy injects five types of realistic errors so the agent learns to assess prior reliability before acting.
-
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
- Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang, Bo Zhou, Haixin Wang, Yufan Du, Shi Bo, Ruihan Lin, Mengqi Yuan, Dunjie Lu, Steven Dillmann, Yiming Shi, Tina Su, Amy Xin, Minghao Liu, Xi Wang, Xu Huang, Ge Zhang, Pengyu Nie, Zhen Yang, Jie Tang, Juanzi Li, Weihao Xuan, Tianyu Liu
- 🏛️ Institutions: Tsinghua, UCLA, RIKEN AIP, University of Waterloo, CMU, Yale University, Northwestern University, ZJU, UC Berkeley, University of Illinois Chicago, Boston University, HKU, Stanford, New York University, TokenWave.AI, University of Tokyo
- 📅 Date: September 30, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [OSWorld-Science], [benchmark], [scientific software], [OSWorld]
- 📖 TLDR: OSWorld-Science is a benchmark of 146 expert-proposed tasks for computer-use agents on scientific software, covering workflows such as molecular drawing, pathology image analysis, statistical computing, and physical simulation. Execution-based evaluators inspect application states and generated artifacts with partial credit, and 12 vision-language models still struggle even with a strong harness.
-
Talk2Agent: Benchmarking Voice Interfaces for Text Agents
- Terumi Chiba, Guangzhi Sun, Zheqi Yuan, Chao Zhang
- 🏛️ Institutions: Tsinghua, University of Cambridge
- 📅 Date: September 30, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [Talk2Agent], [benchmark], [voice interface], [speech recognition], [OSWorld], [WildClawBench]
- 📖 TLDR: Talk2Agent benchmarks how well voice interfaces convey human-spoken instructions to computer-use agents, using spoken versions of WildClawBench and OSWorld tasks. It proposes an execution-free, task-conditioned evaluation that measures how much task-relevant information survives transcription, correlating with downstream completion better than WER/CER.
-
cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
- Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh
- 🏛️ Institutions: CMU
- 📅 Date: September 30, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [cua-speedrun], [benchmark], [efficiency], [evaluation], [latency]
- 📖 TLDR: cua-speedrun provides standardized virtual-machine infrastructure, a common agent interface, and task sets for measuring the speed and cost of computer-use agents across four benchmarks. It shows how reasoning effort, agent harness, and environment latency affect speed, and that most benchmark task sets can be reduced without losing statistical power.
-
Absorbed in Inertia: Activation Analysis for Computer-Use Agents
- Giulio Segalini, Zhi Wen Soi, Jérémie Decouchant, Lydia Chen
- 🏛️ Institutions: University of Neuchâtel, Delft University of Technology
- 📅 Date: September 29, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [inertia], [activation analysis], [failure analysis], [R3], [training-free]
- 📖 TLDR: Shows that computer-use agents can exhibit inertia, repeating fruitless actions while recognizing that they are ineffective, and that inertia corresponds to an absorbing region of the underlying model's activation space. Proposes R3 (Reset, Reroute, Restore), which temporarily resets the context trajectory and then restores it, lowering measured inertia by 17-55% across models.
-
AdaptArena: Evaluating Test-Time Personalization of Web Agents
- Dongchan Shin, Xing Han Lù, Jiaqi Deng, Jay Gala, Tomás Vergara Browne, Jaewon Moon, Fengyuan Liu, Alexandre Drouin, Siva Reddy, Alexandre Lacoste
- 🏛️ Institutions: Mila, McGill University, ServiceNow Research, Laval University
- 📅 Date: September 29, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [AdaptArena], [benchmark], [personalization], [preference inference]
- 📖 TLDR: AdaptArena is a benchmark of 480 tasks for test-time personalization of web agents, where the agent must infer a latent user preference from a relevant historical trajectory. Oracle agents given the true preference succeed on 82.92% of tasks while evaluated agents reach at most 15.62%, and correct preference inference does not guarantee downstream execution success.
-
EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?
- Hongcheng Gao, Hailong Qu, Yu Lei, Henghui Sun, Haoyang Li, Yipeng Wei, Naihao Xue, Xiaohan Yu, Zhuo Tao, Yihe Zang, Yajiao Wang, Jingyi Tang, Yi Li, Jingjing Zhou, Jie Luo, Bohan Zeng, Chengyu Shen, Hao Jiang, Chong Chen, Bowen Qu, Olive Huang, Zeqiang Wang
- 🏛️ Institutions: Tsinghua, Zhiman Inc., Chongqing University, University of Chinese Academy of Sciences, Shandong University, BUPT, Fudan, Henan Polytechnic University, XJTU, PKU, ZJU
- 📅 Date: September 29, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [EngiWorld], [benchmark], [professional software], [engineering], [CAD]
- 📖 TLDR: EngiWorld is a benchmark of 1,301 expert-curated tasks across six engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces. Domain verifiers check the geometric validity, physical feasibility, and rule compliance of artifacts; the strongest of seven frontier models reaches an EngiScore of 44.3.
-
HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents
- Tongbo Chen, Junbo Niu, Zhengxi Lu, Niu Lian, Fei Tang, Yuchen Yan, Yike Hong, Yong Du, Yizhou Liu, Bofan Chen, Yongliang Shen
- 🏛️ Institutions: ZJU, PKU, Tsinghua
- 📅 Date: September 29, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [HybridCUA], [dataset], [model], [CLI], [reinforcement learning], [verifiable rewards], [OSWorld], [WindowsAgentArena]
- 📖 TLDR: HybridCUA trains computer-use agents to combine GUI interaction with the command line, using a data pipeline that produces GUI-only, CLI-only, and interleaved trajectories (HybridCUA-8K), followed by supervised fine-tuning and reinforcement learning with CLI-aware rewards. HybridCUA-9B reaches 53.6% on OSWorld, 14.8 points above its base model.
-
MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows
- Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, Lu Chen
- 🏛️ Institutions: SJTU, Suzhou Laboratory, Shanghai AI Laboratory, BIGAI, Shanghai Innovation Institute
- 📅 Date: September 29, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [MatToolBench], [benchmark], [scientific software], [materials science], [OriginPro]
- 📖 TLDR: MatToolBench is a real-environment benchmark of 204 tasks across 10 professional materials science tools, covering GUI operation, OriginPro scripting, and database queries in a Windows 11 VM, with expert-defined sub-criteria for partial credit. The best model reaches 25% on GUI tasks, and failures stem from missing domain-specific operational knowledge rather than visual grounding alone.
-
Neuro-Symbolic Computer Use: Learning Reusable Policies for Reliable and Efficient Execution
- Hyewon Suh, Thanh Minh Nguyen, Chih-Lun Lee, Darrow Hartman, Lizhao Liu, Xin Eric Wang, Ang Li, Jiachen Yang
- 🏛️ Institutions: Georgia Tech, Stanford, Simular Research
- 📅 Date: September 29, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [neuro-symbolic computer use], [reusable policies], [efficiency], [OSWorld], [ScienceBoard]
- 📖 TLDR: Executes recurring computer workflows with a learned neuro-symbolic policy instead of re-planning every run: executable code fixes the stable decisions (ordering, variables, loops, branches) and neural models handle observation-dependent ones such as grounding and state checks. On OSWorld-Verified and ScienceBoard the policies reach the highest Pass^3 among compared methods while cutting per-run cost by 15-217x.
-
PrecogUI: Proactive GUI Agents via Pre-cognitive Simulation and Experience Retrieval
- Bin Kang, Jiarui Ouyang, Li Jiang, Bin Chen, Zhuotao Tian
- 🏛️ Institutions: University of Chinese Academy of Sciences, HKUST, Harbin Institute of Technology, CUHK, Shenzhen Loop Area Institute
- 📅 Date: September 29, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [PrecogUI], [benchmark], [InterfereBench], [world model], [memory], [long-horizon tasks]
- 📖 TLDR: PrecogUI moves GUI agents from reactive to proactive decision-making with an experience pool of anomaly and success patterns, a simulator that forecasts the next symbolic UI layout for a candidate action, and a controller that fuses both with closed-loop error correction. It also builds InterfereBench for long-horizon tasks with strong disturbances, where PrecogUI surpasses prior methods.
-
SCA: Spatial Credit Assignment for Reinforcement Learning of GUI Agents
- Shengtian Yang, Ziyu Xiong, Kaibing Yang, Guangfeng Cai, Yewen Li, Peng Jiang, Gai Kun, Qingpeng Cai, Lei Feng
- 🏛️ Institutions: Southeast University, Kuaishou Technology
- 📅 Date: September 29, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [SCA], [reinforcement learning], [GRPO], [GUI grounding], [credit assignment]
- 📖 TLDR: Spatial Credit Assignment (SCA) refines group-relative reinforcement learning for GUI agents using the screen coordinates of sampled clicks. It predicts each response's reward from the others in mixed groups, and when every sampled click fails it ranks them by distance to the target, improving grounding and offline action prediction.
-
- Ruozhao Yang, Mingfei Cheng, Xiaofei Xie
- 🏛️ Institutions: Singapore Management University
- 📅 Date: September 28, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [Veer], [dark patterns], [safety], [runtime defense], [TrickyArena], [WebDecept]
- 📖 TLDR: Treats task-relevant web state as a runtime control target for defending web agents against deceptive interfaces, since a task-valid action can still realize an unauthorized consequence. Veer constructs and verifies a prospective intervention trajectory toward a safe state before the action executes, and it achieves the highest safe task completion on TrickyArena and WebDecept.
-
GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions
- Shaoqing Zhang, Kehai Chen, Xuefeng Bai, Zhuosheng Zhang, Pengfei Zhang, Yang Xiang, Min Zhang
- 🏛️ Institutions: HIT-Shenzhen, Pengcheng Laboratory, SJTU
- 📅 Date: September 28, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile], [Web]
- 🔑 Key: [GUITAR], [failure analysis], [state transition graph], [offline evaluation], [AndroidControl], [Mind2Web]
- 📖 TLDR: GUITAR is a state-centric diagnostic framework that maps visually diverse screens to shared functional states in a State Transition Graph and analyzes GUI-agent failures over states and transitions. On AndroidControl and Mind2Web it finds that 60.4% of failures occur in 20% of states, and bottleneck-targeted guidance improves success rate.
-
LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles
- Bingo Zhang, Haochuan Lu, Zongjie Li, Genjian Li, Ari Yu Zhang, Chaozheng Wang
- 🏛️ Institutions: Vera Praxis, Tencent, HKUST, Independent Researcher, CUHK
- 📅 Date: September 28, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [LongPuzzleBench], [benchmark], [long-horizon tasks], [puzzle games], [planning]
- 📖 TLDR: LongPuzzleBench evaluates GUI agents on 114 levels of six puzzle games played through native GUI actions, where a legal move can make the puzzle unsolvable several moves later. Seven of ten general-purpose agents solve nothing harder than Medium, and diagnostics show agents judge a move by its visible progress rather than the future options it leaves.
-
One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents
- Ziqiang Wang, Li Gu, Zhixiang Chi, Linlian Jiang, Zihuan Jiang, Linqiang Guo, Siobhan Reid, Zhi Liu, Yang Wang
- 🏛️ Institutions: Concordia University, Mila, University of Toronto, Shanghai University
- 📅 Date: September 28, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web], [Mobile]
- 🔑 Key: [SOLO], [test-time adaptation], [self-distillation], [continual learning], [WebArena], [VisualWebArena], [MobileWorld]
- 📖 TLDR: Defines fully test-time adaptation for GUI agents, where a deployed agent gets one attempt per task and no ground truth, and proposes SOLO. A judge selects episodes it deems successful, a proposer-verifier pair relabels failed prefixes, and a small adapter is updated by top-K self-distillation, improving UI-TARS-7B and Qwen3-VL-8B by three to six points on recurring web and mobile task streams.
-
PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents
- Yangqin Jiang, Lingrui Xu, Chao Huang
- 🏛️ Institutions: HKU
- 📅 Date: September 28, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [PhoneCLI], [compilation], [efficiency], [AndroidLab], [AndroidWorld], [training-free]
- 📖 TLDR: PhoneCLI compiles a mobile app's GUI navigation into callable commands by exploring the app offline into an annotated screen map; online, the agent verifies and replays a command deterministically and falls back to the VLM for open-ended interaction. It needs no app-internal API or training and improves success while cutting steps and tokens on AndroidLab.
-
SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety
- Jianxing Chen, Xiao Yu, Shipra Agrawal, Zhou Yu
- 🏛️ Institutions: Columbia
- 📅 Date: September 28, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [SCOUT], [safety], [verification], [agent-as-a-judge], [AutoElicit-Bench], [OS-Blind]
- 📖 TLDR: SCOUT is a two-stage agentic safety verifier for computer-use agents: a rubric generator reasons over the task and trajectory to define safe and complete execution, and a probing agent interacts with the post-execution environment to collect grounded evidence. It outperforms LLM-as-a-judge verifiers on AutoElicit-Bench and OS-Blind, and test-time reflection lowers the unsafe execution rate from 30.2% to 17.2%.
-
Same Tasks, Different Apps: Why Mobile GUI Agents Fail to Generalize?
- Tien Tran, Namho Koh, Daiki E. Matsunaga, Ayush Jain, Kee Eung Kim
- 🏛️ Institutions: KAIST, CMU
- 📅 Date: September 28, 2026
- 📑 Publisher: Findings of EMNLP 2026
- 💻 Env: [Mobile]
- 🔑 Key: [AnyAppBench], [benchmark], [cross-application generalization], [failure analysis], [Android]
- 📖 TLDR: AnyAppBench is a live Android benchmark with 100 task templates and 520 task-application pairs over 52 applications that keeps the user goal fixed across apps. Across 13 agents, success on the original application does not transfer reliably to new applications with the same goal, and sub-goal decomposition does not close the gap.
-
- Ziqi Zhang, Shaohui Li, Bing Li
- 🏛️ Institutions: Beijing Key Laboratory of Multimodal Super-Intelligent Security, People AI
- 📅 Date: September 28, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [WebRetriever Challenge], [technical report], [visual localization], [context management]
- 📖 TLDR: Technical report of the winning system in the WebRetriever Challenge. It uses semantic webpage information for routine browser operations and invokes vision only when structured representations fall short, with grid-assisted visual localization, hierarchical context management, and fault-aware execution across eight concurrent browser workers; it achieved a 59% pass rate on the official Protocol 3.
-
WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents
- Anton Emelyanov, Maria Tikhonova, Zaven Martirosian, Sergei Averkiev, Alena Fenogenova
- 🏛️ Institutions: DAIMLD, HSE University, NUST MISIS
- 📅 Date: September 28, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [WebPageBench], [benchmark], [event-level verification], [UI variation], [evaluation]
- 📖 TLDR: WebPageBench verifies each web-agent task from the interface's own event log across six instrumented mock sites, with no judge model, and supports controlled UI variation that re-renders one control while keeping the prompt and success conditions fixed. On its 152-task leaderboard, the gap between what agents declare finished and what the log confirms reaches 41 points.
-
Probe to Act: Elevating Browser-Use Agent via Active Visual Probing
- Keliang Li, Heng Wang, Chen Hu, Daxin Jiang, Hong Chang, Shiguang Shan
- 🏛️ Institutions: Institute of Computing Technology, CAS, University of Chinese Academy of Sciences, StepFun
- 📅 Date: September 27, 2026
- 📑 Publisher: Findings of EMNLP 2026
- 💻 Env: [Web]
- 🔑 Key: [P2A], [Set-of-Marks], [memory], [VisualWebArena], [long-horizon tasks]
- 📖 TLDR: Probe to Act (P2A) lets a browser-use agent issue lightweight probes at decision time that translate DOM handles into pixel evidence and map screen regions back to DOM candidates, keeping only probed or committed observations as memory. It works as a prompting strategy for proprietary models and can be distilled into open-weight models, improving results on three browser-use benchmarks.
-
Relevance Does Not Imply Applicability: Experience Activation for Personal GUI Agents
- Fuyao Zhang, Xuan Wang, Zherui Li, Jiaming Zhang, Longtao Huang, Wei Yang Bryan Lim
- 🏛️ Institutions: NTU, Alibaba Group
- 📅 Date: September 27, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [ExpActivator], [personalization], [memory], [training-free], [proactive agent]
- 📖 TLDR: Argues that relevance does not imply applicability for personal GUI agents: retrieved history mainly helps the first step of an episode. ExpActivator is a training-free framework that matches each screen to historical states in the backbone's latent space and activates a recurring intent only when time and scenario support it, improving within-trajectory step success on four backbones.
-
AgentTell: Behavioural Side-Channel Leakage in Browser-Use Agents
- Asif Shahriar, Md Nafiu Rahman, Sadif Ahmed, Farig Sadeque, Md Rizwan Parvez
- 🏛️ Institutions: BRAC University, Qatar Computing Research Institute
- 📅 Date: September 26, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [AgentTell], [benchmark], [privacy], [side-channel leakage], [browser-use agent]
- 📖 TLDR: AgentTell studies behavioural side-channel leakage, where a browser-use agent's actions reveal a private fact learned on a prior website despite an instruction not to disclose it. Across 9,760 sessions on six backbones, agents reveal the secret through their actions in 61.1% of sessions.
-
CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
- Prince Zizhuang Wang, Chenhao Liang, Zelong Xu, Aojie Yuan, Xiaolin Zhou, Haiyue Zhang, Yue Zhao, Xiyang Hu, Shuli Jiang
- 🏛️ Institutions: CMU, USC, University of Wisconsin-Madison, Arizona State University, AWS Agentic AI
- 📅 Date: September 26, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [CUA-SWE], [benchmark], [software engineering], [environment], [visual feedback]
- 📖 TLDR: CUA-SWE is a benchmark, environment, and evaluation pipeline for software engineering with computer use, in which agents modify code, run commands, interact with the running application, and inspect visual feedback within one task. It spans four software engineering domains with deterministic task-specific tests, and includes tasks where required information is available only through the application's interface.
-
CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning
- Xin Yan, Zhengbo Jiao, Jiaqi Liu, Zhenglin Wan, SiYuan Ma, Xuliang Yu, Tianyi Jiang, Chubin Zhang, Pengfei Zhou, Wangbo Zhao, Xingrui Yu, Bo An, Yang You, Ivor Tsang
- 🏛️ Institutions: A*STAR, HKUST, Beijing Normal University, NUS, NTU, ZJU, PKU
- 📅 Date: September 26, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web], [Desktop]
- 🔑 Key: [CUA-Sandbox], [training environment], [reinforcement learning], [efficiency]
- 📖 TLDR: CUA-Sandbox makes computer-use reinforcement learning cheaper by sharing initialized application runtimes across concurrent rollouts while keeping each trajectory's mutable state in a private capsule with transactional resets and branches. It matches or improves task success relative to Docker with up to 6.20x rollout throughput and 9.2x lower per-environment memory.
-
DashAct: A Progressive Diagnostic Benchmark for GUI Agents in Interactive Dashboard Analysis
- Chuhan Zhang, Qi Xie, Ziyue Wang, Jianing Yin, Yunfan Zhou, Dazhen Deng, Yingcai Wu
- 🏛️ Institutions: ZJU
- 📅 Date: September 26, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [DashAct], [benchmark], [dashboard analysis], [diagnostic evaluation], [visual grounding]
- 📖 TLDR: DashAct is a diagnostic benchmark of 357 human-verified trajectories for GUI agents on interactive dashboards. A progressive cascade evaluates end-to-end execution, then restores verified context for next-action prediction, then adds target semantics and a local view for grounding, revealing bottlenecks hidden by end-to-end scores.
-
- Yangzhe Peng, Xiaoyang Wang, Yiyang Zhao, Lijun Wu, Kun He
- 🏛️ Institutions: HUST, Shanghai AI Laboratory, ZJU, Shanghai Innovation Institute
- 📅 Date: September 26, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [SurfPRM], [dataset], [process reward model], [reward model], [data synthesis], [WebArena-Lite], [WebPRMBench]
- 📖 TLDR: SurfPRM synthesizes process preference data for comparative web process reward models by building an Interaction Element Graph from demonstrations and using it to propose grounded contrastive negative actions, raising the share of grounded minimal contrastive pairs from 24.19% to 74.60%. PRMs trained on the resulting data improve on WebPRMBench and guide step-level search on WebArena-Lite.
-
PastForward: Faster On-Device GUI Agents via Computational Experience Reuse
- Taehwan Park, Changmin Lee, Hayeon Lee, Taesik Gong
- 🏛️ Institutions: UNIST, Meta
- 📅 Date: September 26, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [PastForward], [efficiency], [on-device agents], [speculative decoding], [AndroidWorld]
- 📖 TLDR: PastForward speeds up on-device GUI agents by reusing computation from earlier task executions: it retrieves prior outputs as multi-token proposals verified in one forward pass, and begins next-step inference from predicted screens while the current action runs. It cuts action-step latency by 1.63-2.36x on AndroidWorld workloads while maintaining task success.
-
REBASE: Device-Cloud Experience Coherence for GUI Agents Across App Updates
- Beining Wu, Jun Huang, Yanxiao Zhao
- 🏛️ Institutions: South Dakota State University, Virginia Commonwealth University
- 📅 Date: September 26, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [REBASE], [app updates], [device-cloud collaboration], [experience reuse], [robustness]
- 📖 TLDR: REBASE keeps the device and cloud copies of a mobile GUI agent's recorded experience coherent across app updates: the cloud replays experience on the new version, the device verifies each step, and failures trigger escalating evidence, starting from an accessibility-subtree diff, until a version-keyed patch is derived. On two Jetson devices it restores the success rate of stale experience to that of fresh experience within one episode.
-
The GUI Is Not the State: Diagnosing State Aliasing in GUI World Models
- Dongsheng Liu, Chao Jin, Wenkui Yang, Hejin Wang, Junwei Yang, Zeren Zhang, Ziwei Chen, Huaibo Huang, Jie Cao, Ran He
- 🏛️ Institutions: University of Chinese Academy of Sciences, CASIA, Huawei Noah's Ark Lab
- 📅 Date: September 26, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [StateAliasBench], [benchmark], [world model], [state aliasing], [AndroidWorld]
- 📖 TLDR: Identifies state aliasing in GUI world models, where the visible interface omits transition-relevant environment state so identical observations can lead to different valid futures. Introduces StateAliasBench to isolate it and a lightweight predictive-state recovery that augments frozen world models, restoring state-sensitive prediction and improving agents on AndroidWorld.
-
- Zhenhao Zhang, Zhaoyu Fan, Haohan Ying, Jingwen Hu, Hancen Fan, Junhao Zhou, Zitian Chen, Linchao Zhu
- 🏛️ Institutions: Columbia, ZJU, University of Rochester, UIUC
- 📅 Date: September 25, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [CaptchaArena], [CAPTCHA], [dataset], [model], [reinforcement learning]
- 📖 TLDR: CaptchaArena is a training dataset of 50K execution-verified interactive CAPTCHA puzzles across 20 types and 5 interaction modes, with screenshot-action trajectories, step-by-step reasoning annotations, and pixel-mask annotations. Supervised fine-tuning followed by reinforcement learning on it yields CaptchaAgent, a single 9B policy for all 20 types.
-
From Tapping to Hopping: Augmenting Mobile GUI Agents with App-Native Deeplinks
- Yuchen Sun, Chenglin Cai, Gongjie Zhang, Tianyu Xia, Quyu Kong, Panrong Tong, Zhengwen Zeng, Long Chen, Steven Hoi, Chongyang Zhang, Yue Wang
- 🏛️ Institutions: SJTU, Tongyi Lab
- 📅 Date: September 25, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [GUI-Hopper], [deeplink], [hybrid interaction]
- 📖 TLDR: GUI-Hopper lets a mobile GUI agent jump to a target screen through an app-native deeplink and fall back to taps and swipes for everything else. Deeplinks are found by static analysis, validated on real devices, and stored with descriptions of their landing screens as a verified catalog; using this catalog improves task success in commercial apps on real devices.
-
ScreenHaystack: Finding Blind Zones in GUI Grounding
- Chenyue Li, Xiaoxiao Sun, Yubo Deng, Qinlin Zhao, Serena Yeung-Levy, Yuhui Zhang
- 🏛️ Institutions: Stanford
- 📅 Date: September 25, 2026
- 📑 Publisher: EMNLP 2026
- 💻 Env: [General GUI]
- 🔑 Key: [ScreenHaystack], [benchmark], [GUI grounding], [ScreenSpot-Pro], [blind zones]
- 📖 TLDR: ScreenHaystack is a needle-in-a-haystack benchmark that relocates controlled target icons across high-resolution GUI backgrounds, revealing blind zones where grounding accuracy drops sharply for models including Qwen3-VL, UI-TARS, GTA, and UI-Venus. The blind zones transfer to ScreenSpot-Pro, and augmenting training data in them improves accuracy.
-
Jev-Mobile: Jev as an Executor for Mobile GUI Agents
- Linghua Zhang
- 🏛️ Institutions: Rice University
- 📅 Date: September 24, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [Jev-Mobile], [Jev], [accessibility tree], [efficiency]
- 📖 TLDR: Jev-Mobile splits a mobile GUI agent into low-frequency VLM planning and high-frequency execution: the VLM sets local goals, the accessibility tree defines the executable action space, and Jev, a fast typed decision model, selects actions inside it. On the full AndroidWorld suite it reaches 79% task success against 84% for a step-wise VLM baseline, while cutting execution time by 32.7% and model API cost by 73.4% on successful trajectories.
-
- Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasović
- 🏛️ Institutions: University of Utah
- 📅 Date: September 24, 2026
- 📑 Publisher: Findings of EMNLP 2026
- 💻 Env: [Web]
- 🔑 Key: [KNOWS], [benchmark], [long-horizon tasks], [artifact generation]
- 📖 TLDR: KNOWS is a benchmark of open-ended, long-horizon, browser-based tasks in which an agent retrieves information, synthesizes it, and produces an artifact such as a document, presentation, or spreadsheet, scored by evaluators that combine deterministic checks with LLM judgments. The best frontier computer-use agent or browser harness fully succeeds on fewer than 3% of tasks, and failures on visual steps make the produced artifacts unusable.
-
CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
- Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang
- 🏛️ Institutions: CMU, MSR
- 📅 Date: September 23, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [CAVEAT], [CAVEAT-Harness], [benchmark], [incentive robustness]
- 📖 TLDR: CAVEAT is a benchmark of nine marketplace environments and eight steering mechanisms that tests whether computer-use agents keep the user's objective when the environment itself has a stake in the outcome. Agents buy the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering is enabled, and CAVEAT-Harness, which targets the diagnosed failure modes, raises user-optimal purchasing by 55.0%.
-
- Chengguang Gan, Yunhao Liang, QingHao Zhang, Shiwen Ni
- 🏛️ Institutions: Independent Researcher, University of Chinese Academy of Sciences, Pusan National University, Shenzhen University of Advanced Technology
- 📅 Date: September 23, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [WebMRE], [benchmark], [WebArena], [offline evaluation], [mutual reinforcement]
- 📖 TLDR: WebMRE is an offline benchmark of 541 tasks and 5,293 steps derived from successful WebArena trajectories that scores a web agent checkpoint deterministically without an environment, with each step pairing a guide sentence with a grounded action. Decoding the guide causally improves action accuracy (forcing the gold guide lifts it from .422 to .684), and the mutual reinforcement effect grows from the 4B to the 9B model.
-
Invisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents
- Zeyu Zhang, Dingyi Rong, Zijian Chen, Zicheng Zhang, Xiongkuo Min, Guangtao Zhai
- 🏛️ Institutions: SJTU, Shanghai AI Laboratory
- 📅 Date: September 23, 2026
- 📑 Publisher: ACM MM 2026
- 💻 Env: [Web]
- 🔑 Key: [MVCAP], [MVCAP-Bench], [benchmark], [CAPTCHA]
- 📖 TLDR: Motion Vision CAPTCHA (MVCAP) hides target semantics in motion-defined foreground structures that can only be recovered by temporal segregation from a dynamic background, at three levels from coherent to biological motion. On MVCAP-Bench, a browser-based benchmark of 600 live instances, humans reach 99.6% while the best GUI agent reaches 16.8%, close to six-way chance.
-
Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents
- Yan Zhang, Daiqing Wu, Huawen Shen, Liang Li, Gang Cao, Zhi Gong, Wei Dai, Xiaode Zhang, Can Ma, Yu Zhou
- 🏛️ Institutions: Institute of Information Engineering, Tencent, Tsinghua, University of Chinese Academy of Sciences, Nankai University
- 📅 Date: September 23, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [GUI-SD-v2], [self-distillation], [model]
- 📖 TLDR: GUI-SD-v2 extends on-policy self-distillation from GUI grounding to multi-turn GUI agents in two stages: it first strengthens the self-teacher's ability to follow privileged guidance, then selectively distills step-specific reasoning and memory guidance. It outperforms existing on-policy self-distillation baselines and the evaluated state-of-the-art methods on AndroidWorld and MobileWorld in Pass@1 and Pass@3.
-
Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction
- Yan Zhang, Pei Fu, Daiqing Wu, Huawen Shen, Ruoceng Zhang, Shaojie Zhang, Jiahui Yang, Yu Zhou, Can Ma, Zhenbo Luo, Jian Luan
- 🏛️ Institutions: Institute of Information Engineering, Xiaomi, University of Chinese Academy of Sciences, Nankai University
- 📅 Date: September 22, 2026
- 📑 Publisher: EMNLP 2026
- 💻 Env: [Mobile], [Web]
- 🔑 Key: [MaP], [masked trajectory prediction], [multi-task], [model]
- 📖 TLDR: MaP (Masked Trajectory Prediction) casts multi-turn GUI interaction as a trajectory and defines training objectives by masking and predicting its components, unifying step-wise decision-making, state-action alignment, and long-horizon planning under one objective. A role-aware adapter routes tokens to specialised representation spaces, and MaP outperforms direct mixture training on five GUI navigation benchmarks.
-
How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation
- Yunxiang Li, Xixin Wu, Helen Meng
- 🏛️ Institutions: CUHK
- 📅 Date: September 21, 2026
- 📑 Publisher: EMNLP 2026
- 💻 Env: [General GUI]
- 🔑 Key: [PACE], [confidence estimation], [GUI grounding], [ScreenSpot-Pro]
- 📖 TLDR: GUI agents emit click coordinates as digit tokens, and Place-Aware Coordinate Entropy (PACE) weights each digit's Shannon entropy by its place value to estimate per-click confidence in a single forward pass. On ScreenSpot-Pro and ScreenSpot-v2 it wins AUROC and selective accuracy on all primary comparisons, matching or outperforming K-sample baselines at a fraction of the cost.
-
OSWorld-Pro: Process-based Evaluation for Computer Use Agents
- Zhilin Wang, Shaokun Zhang, Yifan Zhang, Hao Zhang, Jin Xu, Binfeng Xu, Jian Hu, Yunheng Zou, Karan Sapra, Andrew Tao, Jan Kautz, Yi Dong
- 🏛️ Institutions: NVIDIA
- 📅 Date: September 21, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [OSWorld-Pro], [benchmark], [process evaluation], [OSWorld]
- 📖 TLDR: OSWorld-Pro is a set of over 300 tasks with over 2,800 subgoals, grounded in more than 67,000 human annotations, that evaluates computer-use agents on the process rather than only the end state, using human-aligned LLM judges. The top performer, Claude Opus 5, reaches 75.7% against 83.4% on OSWorld, and the analysis surfaces process failures such as subgoal-irrelevant actions and click mistakes.
-
MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory Learning
- Changyue Jiang, Jiayi Wang, Xin Wen, Jiarun Dai, Geng Hong, Xudong Pan
- 🏛️ Institutions: Fudan, Shanghai Innovation Institute
- 📅 Date: September 19, 2026
- 📑 Publisher: USENIX Security 2026
- 💻 Env: [Mobile]
- 🔑 Key: [MATE], [MATEBench], [benchmark], [security], [model]
- 📖 TLDR: MATE is a policy-conditioned auditor that reads a mobile agent's trajectory together with a natural-language security policy, decides whether the trajectory violates it, and explains why, so policies can change without retraining. It is trained on more than 140K synthesized trajectories built from hundreds of apps, is evaluated on the released MATEBench, and audits AutoGLM and Mobile-Agent trajectories on real devices with over 95% accuracy.
-
MintAct: A Unified Visual Agent for Digital Environments
- Mingfei Gao, Rui Tian, Haiming Gang, Bohan Zhai, Le Zhang, Yuanzheng Gong, Di Feng, Ege Özsoy, Kaixin Ma, Vishwesh Kirthivasan, Oğuzhan Fatih Kar, Roman Bachmann, Anders Boesen Lindbo Larsen, Afshin Dehghan
- 🏛️ Institutions: Apple
- 📅 Date: September 18, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile], [Desktop], [Web]
- 🔑 Key: [model], [MintAct], [GUI grounding], [reinforcement learning], [unified agent]
- 📖 TLDR: MintAct is a family of 2B/4B/8B vision-language models that unifies GUI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use in a single agent. Backed by an asynchronous RL infrastructure that hosts hundreds of concurrent environment instances, it matches per-domain specialists across all these capabilities and reaches 48.9 on OSWorld-Verified.
-
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
- Shuai Bai, Jiayong Deng, Yikun Fu, Chang Gao, Xuhao Hu, Mianqiu Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Keliang Li, Ning Li, Wanli Li, Dayiheng Liu, Dunjie Lu, Changwei Luo, Que Shen, Zheyuan Wang, Zijian Wang, Jie Wu, Gao Wu, Zhihui Xie, Rui Xie, Haiyang Xu, An Yang, Jiakang Yuan, Yanming Zhang, Jiajun Zhang, Xi Zhang, Zhenru Zhang, Zhuo Zhen, Mingkang Zhu, Bowen Zhou
- 🏛️ Institutions: Alibaba Group
- 📅 Date: September 18, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop], [Mobile], [Web]
- 🔑 Key: [benchmark], [dataset], [RecreationWorld], [RecreationBench], [hybrid computer use], [execution-grounded reward]
- 📖 TLDR: RecreationWorld is a five-platform environment suite (Ubuntu, macOS, Windows, Android, Web) for hybrid computer-use agents that interleave GUI interaction with coding, using a running reference application as an oracle for hidden behavioral tests. Models trained on its generated trajectories improve across five out-of-distribution coding and hybrid benchmarks, and on the held-out 250-task RecreationBench the leading agent reaches 58.1% overall while passing all programmatic tests on only 2.8% of tasks.
-
A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents
- Haya Halimeh, Sascha Kaltenpoth, Kevin Bösch, Oliver Müller
- 🏛️ Institutions: Paderborn University
- 📅 Date: September 17, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [safety], [digital nudging], [dual-process theory], [choice architecture]
- 📖 TLDR: This study tests whether LLM-based GUI agents are susceptible to digital nudges, running 21,600 simulations of 3,600 agents across six frontier models in a randomized online-shopping experiment. Agents proved vulnerable to both automatic and reflective nudges, and extended reasoning did not confer robustness: it reduced susceptibility to default nudges while heightening susceptibility to social-influence nudges.
-
EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data
- Yinzhu Quan, Zefang Liu
- 🏛️ Institutions: Georgia Tech, Capital One
- 📅 Date: September 17, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [EconSkills], [skill library], [skill transfer], [retrieval], [EconWebArena]
- 📖 TLDR: EconSkills distills verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data, each recording scope, navigation steps, verification checks, and recovery steps with instance values replaced by placeholders. Matched skills improve success over no-skill prompting and abstraction beats replaying raw trajectories, but at library scale retrieval is only competitive with the no-skill baseline because approximate matches on uncovered tasks offset the gains.
-
Affora: A Design System for Agent-Friendly Interfaces
- Jin Gao
- 🏛️ Institutions: Independent Researcher
- 📅 Date: September 16, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [Affora], [design system], [agent-friendly interface], [interface affordance]
- 📖 TLDR: Affora is a design system for making software interfaces legible to computer-use agents without constraining their visual design for human users, grounded in three controlled studies of component implementations, visual variation, and interaction-design principles. It reports that agent performance depends on the interaction meaning exposed through the interface representation, and that substantial visual variation remains possible when that meaning is preserved.
-
RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection
- Liyang Fan, Xinping Bi, Yitai Li, Shuaimin Li, Hui Li, Min Yang
- 🏛️ Institutions: Shenzhen University, Shenzhen University of Advanced Technology, Shenzhen Institutes of Advanced Technology, CAS
- 📅 Date: September 16, 2026
- 📑 Publisher: ACM MM 2026
- 💻 Env: [General GUI]
- 🔑 Key: [RankGround], [GroundRanker], [GUI grounding], [reranking], [crop selection]
- 📖 TLDR: RankGround achieves high-resolution GUI grounding with a single VLM call per query by first using GroundRanker, a lightweight multimodal reranker, to pick the most promising crop from a dense candidate set. Ranking supervision is constructed from existing grounding datasets under a strict containment criterion and trained with a pointwise-then-listwise curriculum, yielding 1.4x faster inference and 5.5% average accuracy gains over the second-best method across backbones and screen scales.
-
The Verifiable Action Card: Trustworthy Human-in-the-Loop Control for Secure Autonomous Agents
- Hasnain Irshad, Anam Mughees, Neelam Mughees, Abdullah Mughees, Imtiaz Ali Soomro
- 🏛️ Institutions: Sir Syed CASE Institute of Technology, University of Engineering and Technology Lahore, National Textile University, KFUPM
- 📅 Date: September 16, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [security], [prompt injection], [human-in-the-loop], [agentic browser]
- 📖 TLDR: The Verifiable Action Card reconstructs approval information from the ground-truth pending browser action and trusted intent provenance, renders it out-of-band in trusted browser chrome, and binds the approval to the exact action re-verified at dispatch. On a 24-scenario benchmark covering confused-deputy attacks, dialog forging, indirect prompt injection, and provenance evasion, attack success falls from 68-100% without the mechanism to 0% on every evaluated model, with 78% legitimate-task completion and no false blocks.
-
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
- Kratika Bhagtani, Kusha Sridhar, Maziyar Baran Pouyan, Yuying Zhao, Eugene Siow
- 🏛️ Institutions: Accenture
- 📅 Date: September 15, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [benchmark], [ERPBench], [enterprise software], [state-grounded evaluation]
- 📖 TLDR: ERPBench evaluates screenshot-only computer-use agents on a live, reproducible ERP system and scores each task against ground-truth values in the underlying database rather than on-screen appearance. Across six closed and open-source agents it shows that strong general GUI performance does not transfer to enterprise reliability: some agents save a form in up to 85% of runs but write the correct value in as few as 3%.
-
EchoPath: Execution-Level Replayable Memory for GUI Agents
- Yao Zhao, Aditya Shanmugham, Swastik Roy, Yanxun Xu
- 🏛️ Institutions: JHU, Amazon AGI
- 📅 Date: September 15, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [framework], [EchoPath], [memory], [trajectory replay], [enterprise automation]
- 📖 TLDR: EchoPath converts artifact-validated GUI trajectories into standardized, parameter-controlled callable memories that a host agent invokes only when they can be deterministically replayed in the current runtime. An image-based target-reaiming algorithm treats stored coordinates as visual evidence and corrects them against the current screen before execution, reducing median token cost by more than 90% and median execution time by about 60% on real computer-use tasks.
-
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
- Bofan Chen, Boxuan Zhang, Fei Tang, Zhengxi Lu, Yong Du, Tongbo Chen, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
- 🏛️ Institutions: ZJU, UESTC
- 📅 Date: September 15, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile], [Desktop]
- 🔑 Key: [framework], [EvoSkill-GUI], [skill evolution], [failure recovery], [training-free adaptation]
- 📖 TLDR: EvoSkill-GUI makes GUI agent skills revisable at deployment time instead of fixed beforehand, representing each skill as a multi-file package holding retrieval metadata, executable plans, backup localization, failure-recovery rules, and failure cases. A reflect-revise-reuse loop lets the executor patch skills mid-rollout while an isolated critic diagnoses failed trajectories, improving multiple base models without any training by up to 16.2% on MobileWorld, 6.0% on AndroidWorld, and 10.5% on OSWorld.
-
When Agents See Differently: Exposing UI Desynchronization Threats in Mobile Agents
- Heng Li, Fulin Zhao, Zhe Geng, Zhiyuan Yao, Wei Yuan, Xiapu Luo
- 🏛️ Institutions: PolyU, Huazhong University of Science and Technology
- 📅 Date: September 15, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [security], [UI desynchronization], [repackaging attack], [human oversight]
- 📖 TLDR: This paper identifies human-agent UI desynchronization, where one mobile UI state presents materially different information to a human viewer, bound by occlusion and luminance contrast, than to an agent consuming screenshots and accessibility metadata. An automated framework builds deployable repackaged APKs that exploit the gap without runtime adaptation, reaching 77.9% and 66.9% average misleading rates across five mobile-agent frameworks and three backbones on 546 tasks, while a 186-participant study finds the perturbations hard for people to notice.
-
AnchorGUI: Asymmetric Memory for Dual-Scale Learning in GUI Navigation
- Shengjie Jin, Zelong Sun, Hengbo Xu, Yanbiao Ma, Zhiwu Lu
- 🏛️ Institutions: Renmin University of China
- 📅 Date: September 14, 2026
- 📑 Publisher: ECCV 2026
- 💻 Env: [Mobile]
- 🔑 Key: [AnchorGUI], [memory], [cross-trial distillation], [prediction error], [AndroidWorld]
- 📖 TLDR: AnchorGUI addresses an informational asymmetry in GUI navigation, where expected transitions compress into lightweight text but unexpected outcomes need preserved screenshots as causal evidence, through a Cognitive State Anchor that turns trajectories into explicit prediction-error signals. Its asymmetric memory drives both in-episode correction and cross-trial credit assignment, reaching 57.3% on AndroidWorld with 2.4x fewer tokens per step and 69.2% after cross-trial distillation.
-
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
- Zihan Dong, Yuanzhe Liu, Zhiyuan Ma, Qishi Zhan, Dehan Kong, Guohao Li, Kaixin Li
- 🏛️ Institutions: Georgia Tech, North Carolina State University, Marquette University, Saros Lab, CAMEL-AI.org, Eigent.AI, NUS
- 📅 Date: September 14, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [benchmark], [CADWorld], [computer-aided design], [long-horizon task], [FreeCAD]
- 📖 TLDR: CADWorld is a 200-task benchmark for long-horizon computer use in FreeCAD, spanning 11 mechanical-CAD workflow categories from sketching and part modeling to CAM, FEM, and technical drawing. Agents act through screenshots and GUI actions while success is decided by executable checks over saved FreeCAD artifacts, and the strongest of seven agents reaches 17.5% against an 87.0% expert reference pass rate.
-
Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding
- Yizhou Liu, Fei Tang, Yuchen Yan, Zhengxi Lu, Songqin Nong, Tao Jiang, Wenhao Xu, Wenqi Zhang, Weiming Lu, Jun Xiao, Yongliang Shen
- 🏛️ Institutions: ZJU, Ant Group
- 📅 Date: September 14, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [GUI grounding], [test-time adaptation], [negative learning], [CANL], [ScreenSpot-Pro]
- 📖 TLDR: This work introduces a label-free test-time training paradigm for GUI grounding built on two observations: coordinate-token confidence indicates correctness better than full-sequence confidence, and in sparse GUI coordinate spaces negative samples give more reliable signal than potentially noisy pseudo-positives. The resulting Confidence-Anchored Negative Learning reaches 92.1% on ScreenSpot-V2 and 33.8% on ScreenSpot-Pro, an 8.9-point absolute gain over the base model.
-
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
- Sibo Zhu, Shicheng Fan, Xinyue Wang, Wenyi Wu, Kun Zhou, Biwei Huang
- 🏛️ Institutions: Aether AI, UC San Diego, University of Illinois Chicago
- 📅 Date: September 14, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [framework], [RSIAgent], [recursive self-improvement], [memory construction], [OSWorld]
- 📖 TLDR: RSIAgent is a training-free multi-agent framework in which curriculum, actor, and verifier agents autonomously explore a new environment and accumulate reusable causal knowledge linking actions, conditions, and consequences into a memory that is then frozen for downstream reuse. Using a broad-then-deep exploration strategy, it lifts strong open-source models on OSWorld-v2 and Agent's Last Exam enough to outperform frontier closed-source models.
-
CoArena: Evaluating Computer-Use and Multi-Agent Systems in Real Time
- Nitish Kovuru, Prateek Jannu
- 🏛️ Institutions: Coasty Research Lab
- 📅 Date: September 13, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [benchmark], [CoArena], [live evaluation], [Bradley-Terry model], [pairwise comparison]
- 📖 TLDR: CoArena proposes a real-time evaluation platform for computer-use agents in which real users submit tasks, two systems execute them concurrently in identical sandboxed desktops, and a public leaderboard is refit from blind pairwise judgments. The paper's contribution is a formal account of real-time evaluation as five measurable properties plus the full Bradley-Terry rating methodology; its five-system worked example with 211 votes is labeled illustrative rather than a measurement of a deployed system.
-
PriMobiBench: Characterizing Visual Privacy Leakage in VLM-Driven Mobile GUI Agents
- Qihang Cen, Tianshuo Cong, Da Song, Xinlei He, Jiaxing Song, Ke Xu, Qi Li
- 🏛️ Institutions: Tsinghua, Shandong University, Wuhan University
- 📅 Date: September 12, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [benchmark], [dataset], [privacy], [PriMobiBench], [MobiLeak], [user profiling]
- 📖 TLDR: PriMobiBench is a benchmark for visual privacy leakage in screenshot-driven mobile GUI agents, paired with MobiLeak, a dataset of execution traces from 16 apps covering 25 privacy attributes across 2,960 embedded privacy instances. VLMs extract sensitive on-screen information with up to 82.5% success and infer user profiles from aggregated visual evidence about 70% of the time; masking privacy-sensitive but task-irrelevant UI elements before cloud processing cuts profiling success by up to 58% at roughly 8% performance loss.
-
AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
- Xinyun Cao, Adriana Szekeres, Fazle Elahi Faisal
- 🏛️ Institutions: UMich, MSR
- 📅 Date: September 11, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [framework], [AutoTailor], [tool synthesis], [capability selection], [WebArena]
- 📖 TLDR: AutoTailor builds and maintains a compact set of MCP APIs derived from web trajectories, filtering 1,283 unrefined APIs down to 87 offline by granularity, redundancy, and usage likelihood, then to 33 online through dynamic reselection driven by observed coverage gaps. On 106 WebArena Postmill tasks the resulting set with ReAct fallback reaches 90.6% correctness versus 87.5% for ReAct alone, while cutting average request-token cost by 57.8% and latency by 29.4%.
-
BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents
- Tong Ye, Kunyang Han, Guozhi Wang, Longqiang Luo, Zhifeng Ding, Yongxiang Zhang, Xiaolei Shen, Yuxuan Zhang, Zhuping Zhang, Tao Xu, Yue Pan, Yucheng Zhao, Yupei Hu, Yuanjiang Ouyang, Danfeng Shen, Runqi Lin, Hongda Cai, Zhaoxiong Wang, Mengjia Yan, Yingjie Zhong, Chen Zhou, Zeyu Zhang, Xuwen Zhu, Penggang Shi, Mingcheng Luo, Ziyang Wu, Min Jin, Mingfu Shen, Zairong Xu, Fan Zhang, Hao Wang, Liang Liu, Zhulin Xie, Lijun Yao, Xiao Liang, Liangmin Wen, Liqiang Feng, Feilong Wu, Min Hu, Min Chen, Guanjing Xiong, Xiaohu Ruan, Xiaoxin Chen
- 🏛️ Institutions: vivo AI Lab
- 📅 Date: September 11, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [model], [real-device training], [self-improvement], [AndroidWorld], [BlueLM-GUI]
- 📖 TLDR: BlueLM-GUI is a 35B-A3B native mobile GUI agent trained through continual pre-training, SFT, and agentic RL run on hundreds of real phones instead of sandboxes, with a pipeline that salvages failed trajectories into supervision. It reaches 87.4 on MobileGUI-VBench and 84.9 on AndroidWorld.
-
Token Efficient Task Execution via Application Behavior Modeling for Web Agents
- Alexandru Ianta, Eleni Stroulia
- 🏛️ Institutions: University of Alberta
- 📅 Date: September 11, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [OdoBot], [token efficiency], [application behavior model], [demonstration learning]
- 📖 TLDR: OdoBot is a web-agent architecture that builds a behavioral model of the target application from successful task demonstrations and uses it to cut the token cost of execution. On 45 tasks in the Canvas learning management system it consumes 44% and 80% fewer tokens than Agent-E and WebVoyager respectively while also exceeding WebVoyager's task success rate.
-
VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets
- Yu Bai, Yukai Miao, Dawei Wang, Li Chen, Yanyu Ren, Yuqian Shi, Dan Li, Ying Xiong, Chengqiu Tan, Run Zhou, Li Li
- 🏛️ Institutions: Zhongguancun Laboratory, Tsinghua, Network Management Center, China Mobile
- 📅 Date: September 11, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web], [General GUI]
- 🔑 Key: [benchmark], [verbal reinforcement learning], [trial budget], [exploration-exploitation], [VRL-Bench]
- 📖 TLDR: VRL-Bench is a harness for fairly comparing verbal-reinforcement-learning and trial-and-error-memory methods on computer control under a fixed trial budget, evaluated on MiniWoB and WebShop. It finds reflection can reduce success under tight budgets and proposes VEX², an exploration-exploitation scheduler.
-
- Xiaolei Li, Jialun Cao, Zhijian Hou, Yuzhi Zhao, Yepang Liu, Shing-Chi Cheung
- 🏛️ Institutions: HKUST, SUSTech, CityU
- 📅 Date: September 09, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [framework], [GUI testing], [GraphDroid], [asynchronous exploration], [bug detection]
- 📖 TLDR: GraphDroid is an intent-driven Android GUI testing framework that uses a cluster-based memory over visited states to identify uncovered functionality, generates test intents asynchronously so exploration never blocks, and delegates simple intents to a lightweight heuristic while reserving the LLM for complex ones. Across 41 real apps it achieves up to 36.4% higher code coverage at under one eighth the cost of the best pure-LLM baseline, exposing 19 bugs of which seven were previously unknown.
-
JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition
- Zixiang Chen, Yuheng Lu, Zihao Cheng, Zeming Liu, Jizeng Bai, Ziye Huang, Zhiyin Lin, Zihan Li, Yuhang Guo, Yunhong Wang, Haifeng Wang
- 🏛️ Institutions: Beihang University, Beijing Institute of Technology, Baidu Inc.
- 📅 Date: September 09, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile], [Desktop]
- 🔑 Key: [benchmark], [cross-device], [task composition], [JarvisGUI]
- 📖 TLDR: JarvisGUI is a dynamic benchmark for GUI agents on cross-device workflows spanning Android, Windows, and Ubuntu, formulating tasks as typed input-output transformations so multi-step cross-device workflows can be composed and evaluated automatically.
-
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
- Zhangxuan Gu, Haoxing Chen, Qi Qin, Yi Xin, Kai Gan, Lin Liu, Long Cui, Xiaomei Wang, Beitong Zhou, Yunzhu Zhang, Zhengwen Zeng, Changlong Gao, Weizhi Chen, Rongchao Zhang, Haoyuan Wu, Shuheng Shen, Changhua Meng, Weiqiang Wang, Jianguo Li, Zhenzhong Lan
- 🏛️ Institutions: Inclusion AI, Venus Team, Westlake University
- 📅 Date: September 09, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile], [Desktop], [Web]
- 🔑 Key: [model], [LLaDA-UI], [diffusion language model], [block-wise decoding], [mixture of experts]
- 📖 TLDR: LLaDA-UI is a 16.7B-parameter MoE block-wise diffusion vision-language GUI agent, built by aligning a native-resolution vision encoder with the LLaDA2.0-mini-base diffusion backbone and then fine-tuning on mobile, desktop, web, and grounding data. It substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks, establishing block-wise diffusion as a practical generative paradigm for GUI agents.
-
TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents
- Yuhao Wang, Mu Qiao, Xindong Zhang, Yunzhi Zhuge, Lei Zhang, Huchuan Lu
- 🏛️ Institutions: Dalian University of Technology, OPPO Research Institute, PolyU
- 📅 Date: September 09, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [efficiency], [visual token pruning], [admission control], [TRACE]
- 📖 TLDR: TRACE is a training-free visual-token pruning method for GUI agents that treats pruning as an irreversible admission decision, combining a layout-derived interaction prior with instruction relevance and feature novelty to keep operable regions under tight token budgets.
-
APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents
- Jintian Feng, Long Chen, Xiao Yu, Jiayi Dai, Chenglong Liu, Haoru Wang, Zizhen Xue, Yuxuan Shi, Ziyang Wang, Yichen Gong
- 🏛️ Institutions: Central China Normal University, Agentic Labs, Acrab AI
- 📅 Date: September 07, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [benchmark], [reproducibility], [simulated apps], [APPSim-Bench]
- 📖 TLDR: APPSim-Bench provides 557 tasks across 17 high-frequency Chinese and English apps, rebuilt as controllable simulated apps that preserve interaction logic while removing stochasticity from ads, recommendations, and accounts, and evaluates 19 GUI agents on them.
-
FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?
- Jingpu Yang, Fengxian Ji, Jinri Guo, Tianhao Li, Qian Jiang, Fan Zhang, Min Peng, Qianqian Xie, Preslav Nakov, Zhuohan Xie
- 🏛️ Institutions: Wuhan University, MBZUAI, Northeastern University, Zhongguancun Academy
- 📅 Date: September 07, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop], [General GUI]
- 🔑 Key: [benchmark], [financial agents], [autonomous benchmark construction], [FinCUABuild]
- 📖 TLDR: FinCUABuildBench evaluates whether agents can autonomously construct computer-use-agent evaluation tasks for financial scenarios, via 576 construction requests over 24 financial workflows with a task-qualification mechanism for the resulting benchmarks.
-
AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents
- Zhihao Liu, Hongyu Sun, Zhiyuan Fu, Xiaonan Duan, Jice Wang, Shangru Zhao, Weizhi Meng, Wuxin Yang, Yangfan Zhou, Yuqing Zhang
- 🏛️ Institutions: Hainan University, University of Chinese Academy of Sciences, Lancaster University
- 📅 Date: September 06, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop], [General GUI]
- 🔑 Key: [safety], [security], [adversarial patch], [prompt injection], [AgentHijack]
- 📖 TLDR: AgentHijack is an end-to-end evaluation of image-triggered command injection against computer-use agents: adversarial patches deployed on real pages are tested across five GUI-agent/VLM backends over 600 online cases, reaching 84.5% trigger success rate and 20.3% end-to-end attack success rate.
-
Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool Actions
- Juyong Lee, Woogyeol Jin, Kimin Lee
- 🏛️ Institutions: KAIST
- 📅 Date: September 06, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [framework], [tool actions], [self-generation], [DroidTool]
- 📖 TLDR: DroidTool augments Android GUI agents with a hybrid GUI-plus-tool action space by auto-generating Python tools over app state through an agentic propose/implement/test/repair workflow with relational cross-tool tests, improving both proficiency and efficiency.
-
Selective Knowledge Control for Continual GUI Agent Learning over Application Streams
- Zirui Shang, Xin Shu, Yang Liu, Zhi Gao, Xinxiao Wu, Lifeng Fan
- 🏛️ Institutions: Beijing Institute of Technology, BIGAI, Wuhan University, Shenzhen MSU-BIT University
- 📅 Date: September 06, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [continual learning], [catastrophic forgetting], [neuron-level gradient control]
- 📖 TLDR: This work proposes activation-conditioned selective knowledge control, a neuron-level gradient surgery method that protects highly-activated MLP neurons so GUI agents can learn new applications over a continual stream without interfering with knowledge from earlier ones.
-
CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents
- Haoting Shi, Wenhao Wang, Weicheng Fang, Yaozhong Liang, Tian Jin, Pengxiang Zhao, Guangyi Liu, Siheng Chen, Yanfeng Wang
- 🏛️ Institutions: SJTU, ZJU
- 📅 Date: September 04, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [framework], [environment], [hybrid GUI and CLI], [CUA-Universe]
- 📖 TLDR: CUA-Universe is an environment-to-data pipeline that turns real desktop software into hybrid GUI+CLI environments over shared application state, targeting agents that coordinate visual inspection with command-line throughput.
-
ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults
- Weide Zhan, Qumu Shaqu, Yuanqing Liu, Peng Zhang, Jiahao Liu, Kam Him Lam, Ning Gu, Zhan Hu, Tun Lu
- 🏛️ Institutions: Fudan
- 📅 Date: September 04, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [benchmark], [accessibility], [older adults], [ElderBench]
- 📖 TLDR: ElderBench collects 249 naturally elicited smartphone tasks from older adults across 20 applications, characterizing how elderly instructions (indirect speech, referential ambiguity, under-specification) diverge linguistically from existing GUI benchmark instructions.
-
From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents
- Longtao Hu, Xiao Liang, Linchao Zhu
- 🏛️ Institutions: University of Electronic Science and Technology of China, ZJU
- 📅 Date: September 04, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop], [General GUI]
- 🔑 Key: [skill library], [online evolution], [procedural memory]
- 📖 TLDR: This work converts computer-use-agent trajectories and evaluator feedback into a persistent, versioned library of reusable procedures through online skill evolution, running each iteration against a frozen library snapshot to isolate the library's incremental value.
-
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
- Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng, Zheng Wu, Yansi Li, Chuanbiao Song, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
- 🏛️ Institutions: SJTU, Ant Group
- 📅 Date: September 03, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [conflict detection], [termination], [feasibility verification], [CONFLICTGUI]
- 📖 TLDR: CONFLICTGUI benchmarks instruction-internal and instruction-vs-GUI-context conflicts, exposing execution-biased overcompliance in GUI agents; CONFLICTGUARD is an inference-time feasibility-verification framework that aligns termination decisions with action generation.
-
Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails
- Suyoung Lee, Myungsub Choi
- 🏛️ Institutions: Tynapse
- 📅 Date: September 02, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [safety], [security], [prompt injection], [guardrail evaluation], [Mind2Web-Injection]
- 📖 TLDR: Mind2Web-Injection provides 9,954 instruction-screenshot pairs with pixel-exact evidence boxes and image-side counterfactuals, measuring whether a web-agent guardrail VLM actually localizes the on-screen injected text it claims to detect, rather than only judging accuracy of the verdict.
-
Discriminative World Models for Web Agents
- Kelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang, Leon Oks, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig
- 🏛️ Institutions: UC Berkeley, MIT-IBM Watson AI Lab, Cal Poly San Luis Obispo, Xero
- 📅 Date: September 02, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [world model], [discriminative training], [state prediction]
- 📖 TLDR: This work replaces supervised next-state prediction with predicted-state matching, so a web world model's predicted HTML/AXTree state must discriminate the true resulting state from states reached by alternative actions, aligning the world model with the downstream ranker or process reward model.
-
Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization
- Bizhe Bai, Jiakang Yuan, Hongming Wu, Xinyue Wang, Jie Ren, Siyao Chen, Yuchen Ya, Fan Bai, Pai Peng, Huafeng Qin, Tao Chen
- 🏛️ Institutions: Fudan, Shanghai Innovation Institute, Chongqing Technology and Business University
- 📅 Date: September 02, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [survey], [efficiency], [observation], [memory], [runtime optimization]
- 📖 TLDR: This survey organizes GUI-agent efficiency research along observation, context/memory, action, and planner/system axes, arguing that real-world deployment depends on efficiency as much as on task success rate.
-
Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision
- Sitong Pan, Yipeng Shen, Yilin Lu, Caiwen Ding, Lu Cheng, Qianwen Wang
- 🏛️ Institutions: University of Minnesota, Purdue University, Pennsylvania State University
- 📅 Date: September 02, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [runtime monitoring], [risk prediction], [observable trajectories]
- 📖 TLDR: This work does prefix-level risk prediction for web agents using only observable trajectory signals with no access to token logits, combining macro cross-step features with micro intention/action/state-change consistency features, labeled by the first uncorrected critical error.
-
OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations
- Yixiong Xiao, Lang An, Hucheng Yang, Pinxue Ma, Yongquan Chen, Jingjia Cao, Yusai Zhao, Ting Wang, Ting Liu, Siqi Bao, Jingbo Zhou, Hua Wu
- 🏛️ Institutions: Baidu Inc., Ningxia Electric Power Engineering Co., Ltd.
- 📅 Date: September 02, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [framework], [standard operating procedures], [human demonstrations], [OmegaUse-SOP]
- 📖 TLDR: OmegaUse-SOP is a human-in-the-loop system that turns demonstrations of professional computer use into standard operating procedures for GUI agents, targeting domain-specific workflows with implicit knowledge and software-specific conventions.
-
WiP: Characterizing and Defending Against Mobile-Agent-Driven MFA Automation
- Yimeng Liu, Hua Huang
- 🏛️ Institutions: University of California, Merced
- 📅 Date: September 02, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [safety], [security], [multi-factor authentication], [risk signal]
- 📖 TLDR: This work-in-progress paper identifies "factor collapse" -- a mobile GUI agent completing all authentication factors inside one autonomous environment, completing 10/10 authorized MFA workflows -- and proposes a motion-based Android risk signal for separating human from agent-driven logins.
-
- Satwik Ram Kodandaram, Monalika Padma Reddy, Xiaojun Bi, Jiawei Zhou, I. V. Ramakrishnan, Vikas Ashok
- 🏛️ Institutions: Stony Brook University, Old Dominion University
- 📅 Date: September 01, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [accessibility], [blind users], [failure taxonomy], [OLLA]
- 📖 TLDR: A three-week diary study with 8 blind users of OLLA, a screen-reader-accessible computer-use-agent prototype, collects 1,258 commands across 12 applications; GPT-5 reaches 52.5% success, with a failure taxonomy spanning grounding, planning, constraint tracking, and termination.
-
SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction
- Bowei He, Xiaokun Zhang, Meng Ding, Xue Liu
- 🏛️ Institutions: MBZUAI, McGill University, CityU, UMass Boston
- 📅 Date: August 31, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [self-improvement], [skill abstraction], [hierarchical skills], [SCAFFOLD]
- 📖 TLDR: SCAFFOLD induces parametric, executable skills from successful trajectories under a multi-instance abstraction constraint and maintains a recursively composed hierarchy where higher-level skills invoke lower-level ones, for visual web agents.
-
SIR: Self-improving Red-teaming for Compute Use Agents
- Chen Xiong, Zhiyuan He, Pin-Yu Chen, Stjepan Picek, Tsung-Yi Ho
- 🏛️ Institutions: CUHK, IBM Research, University of Zagreb, Radboud University
- 📅 Date: August 31, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop], [General GUI]
- 🔑 Key: [safety], [security], [red-teaming], [prompt injection], [SIR]
- 📖 TLDR: SIR is a black-box indirect-prompt-injection attack that composes stealthy injections from a reusable principle library and self-improves over rounds, arguing that fixed hand-written injection benchmarks underestimate adaptive adversaries against computer-use agents.
-
When and What to Teach: Budget-Aware Online Adaptation for Web Agents
- Jianwei Zhang, Sihan Cao, Pengcheng Zheng, Ya Wen, Pei Ke, Kuien Liu, Shen Gao, Wei Dong, Yang Yang, Chaoning Zhang
- 🏛️ Institutions: University of Electronic Science and Technology of China, Institute of Software, Chinese Academy of Sciences, Xi'an University of Architecture and Technology
- 📅 Date: August 31, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [online adaptation], [budget-aware teaching], [distillation]
- 📖 TLDR: Score-Guided Online Teaching with Budgeted Trajectory Trimming does post-deployment online teaching of a lightweight local web agent by a stronger teacher, avoiding budget spent on unresolvable episodes and redundant turns.
-
ActReal: System-Level Mobile Agents Challenge Mobile Automation Detection
- Mingshuo Wang, Hanqing Guo, Huining Li, Yuliang Fu, Jing Xu, Chenhan Xu
- 🏛️ Institutions: North Carolina State University, Indiana University
- 📅 Date: August 30, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [security], [automation detection], [sensor spoofing], [ActReal]
- 📖 TLDR: ActReal shows a privileged system-level mobile agent can jointly synthesize time-aligned touch and six-axis IMU signals, defeating automation detectors that rely on touch trajectories, timing, and touch-IMU physical coupling.
-
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
- Lin Fu, Zheyuan Yang, Tianhui Zhang, Jinbiao Wei, Guo Gan, Boxu Liu, Yilun Zhao, Yu Rong
- 🏛️ Institutions: Yale University, ZJU, China University of Geosciences, DAMO Academy, Alibaba Group, Tongji University, UC San Diego
- 📅 Date: August 30, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [benchmark], [world model], [contextual consistency], [GUI-CC]
- 📖 TLDR: GUI-CC evaluates GUI world models as multi-step environments rather than one-step next-screen predictors, via an offline reference-action track over real mobile trajectories and an online agent-loop track that stresses contextual consistency over long horizons.
-
Learning Simple Test-Time Environments for LLM Web Agents
- Junxuan Li, Zijun Liu, Ziyi Huang, Peng Li, Yuzhou Liu, Ming Yan, Yang Liu
- 🏛️ Institutions: Jilin University, Tsinghua, Institute for AI Industry Research, Tsinghua, Beijing Jiaotong University, Tongyi Lab, Alibaba Group
- 📅 Date: August 29, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [test-time adaptation], [environment decomposition], [label-free]
- 📖 TLDR: Test-Time Environment Decomposition is a label-free method with trial steps that let a web agent decompose a complex environment observation into sub-modules and adapt its behavior during inference, without additional training.
-
CURA: Certified Runtime Alarms for Computer-Use Agents
- Divake Kumar, Sina Tayebati, Devashri Naik, Amanda Sofie Rios, Nilesh Ahuja, Omesh Tickoo, Ranganath Krishnan, Amit Ranjan Trivedi
- 🏛️ Institutions: University of Illinois Chicago, Intel Labs, Capital One AI Labs
- 📅 Date: August 28, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [runtime monitoring], [false-success detection], [sequential testing], [CURA]
- 📖 TLDR: Documenting that 90% of an OSWorld pipeline's failures end with a false success claim, CURA proposes an external monitor reading only harness-visible telemetry, with no access to model internals or extra LLM calls, that turns a running trajectory into a sequential statistical test.
-
- Jiahe Ying, Wendong Bu, Kaihang Pan, Bingchen Miao, Siyu Chen, Wen Wang, Xueming Jiang, Juncheng Li, Siliang Tang
- 🏛️ Institutions: Fudan, ZJU
- 📅 Date: August 28, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [framework], [reward shaping], [cycle-consistency], [Iron]
- 📖 TLDR: Iron is an annotation-efficient GUI-agent training framework using a stepwise cycle-consistent reward to align low-level actions with high-level intents, plus retrospective reuse of failed trajectories.
-
ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions
- Rui Xie, Lu Chen
- 🏛️ Institutions: X-LANCE Lab, Shanghai Jiao Tong University, BIGAI
- 📅 Date: August 27, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [framework], [structured state], [semantic actions], [ASIL]
- 📖 TLDR: ASIL argues screenshot-and-click interaction is state-incomplete and brittle, and introduces an agent-native interaction layer exposing software via structured JSON observations and code-executable semantic actions, instantiated across 15 applications with a 300-task benchmark.
-
- Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhen Lin, Zhenyu Ma, Changhua Meng, Han Peng, Xinyu Qiu, Shuheng Shen, Zhongyi Shui, Weiqiang Wang, Ming Wen, Zhuoer Xu, Hang Yan, Kaiwen Yang, Ruilin Yao, Nanjun Yu, Zhengwen Zeng, Lianrui Zhang, Yunzhu Zhang, Zhe Zhao, Beitong Zhou
- 🏛️ Institutions: Ant Group
- 📅 Date: August 27, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [model], [foundation agent], [reward verification], [UI-Venus-2]
- 📖 TLDR: UI-Venus-2 is a general-purpose foundation GUI agent for mobile, web, and desktop under a unified closed-loop reasoning-action framework, scaling environments (170+ multilingual mobile apps plus native desktop OSes), task construction, and reward verification.
-
WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning
- Yu Han, Tianwen Qian
- 🏛️ Institutions: East China Normal University
- 📅 Date: August 27, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [reinforcement learning], [world model], [reasoning], [WM-R1]
- 📖 TLDR: WM-R1 trains mobile GUI agents with a world model substituting for the real Android environment during RL rollouts, embedding the world model into the reasoning process so the agent predicts the consequences of candidate actions before acting.
-
LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI Agents
- Weiming Li, Helen Paik, Yulei Sui
- 🏛️ Institutions: University of New South Wales
- 📅 Date: August 26, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [local deployment], [control state], [small models], [LocalLSTC]
- 📖 TLDR: Replacing GPT-5 with Qwen3.5-9B drops average OSWorld SR-100 from 60.9% to 37.7% with control failures in 91.6% of failed trajectories; LocalLSTC is a training-free architecture that makes persistent control state explicit to stabilize locally deployed GUI agents.
-
ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
- Joy Chen, Alejandro Castillejo Munoz, Pierluca D'Oro, Yuxuan Sun, Chloe Evans, Joseph Tighe
- 🏛️ Institutions: FAIR at Meta
- 📅 Date: August 25, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile], [Desktop]
- 🔑 Key: [benchmark], [safety], [trustworthiness], [disambiguation], [ADeptS-Bench]
- 📖 TLDR: ADeptS-Bench is a dual-stream trustworthiness benchmark for computer-use agents on mobile and desktop: a Safety stream with paired benign/malicious tasks and threats embedded in the visual interface, and a Disambiguation stream testing whether agents seek clarification; no model exceeds 80% success while staying under 30% attack success.
-
- Guo Gan, Yilun Zhao, Cong Chen, Jinbiao Wei, Tingyu Song, Zheyuan Yang, Lin Fu, Hong Zhou
- 🏛️ Institutions: ZJU, Yale University, University of Chinese Academy of Sciences, Tongji University
- 📅 Date: August 25, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [AnTrap], [benchmark], [reinforcement learning], [safety], [security], [GRPO], [AndroidWorld]
- 📖 TLDR: AnTrap is an AndroidWorld-based benchmark that injects controllable runtime anomalies, organized into four categories with ten subcategories, into Android GUI-agent trajectories. Testing 16 leading GUI agents reveals substantial robustness degradation under these anomalies, and adversarial RL training can address many single-step traps but struggles with deep contextual failures such as state deadlock.
-
BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes
- Fei Tang, Huawen Shen, Zhiqiong Lu, Zhengxi Lu, Pengyuan Lyu, Chengquan Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
- 🏛️ Institutions: ZJU, Tencent
- 📅 Date: August 25, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [data synthesis], [parallel sandboxes], [BrowserForge]
- 📖 TLDR: BrowserForge generates pixel-based web-interaction training data at scale by driving many browser sandboxes in parallel over the open web, rather than staying bound to predefined site lists or tutorial sources.
-
Reflection with Action-Induced Visual Differences for Desktop GUI Agents
- Yijie Ma, Chaoyue Niu, Fan Wu, Guihai Chen
- 🏛️ Institutions: SJTU
- 📅 Date: August 25, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [reflection], [outcome verification], [Planner-Operator-Reflector]
- 📖 TLDR: Evidence-First Reflection is a two-stage reflector for the Planner-Operator-Reflector agent loop that decouples detection of action-induced visual differences from outcome verification, targeting dense desktop GUIs where state changes are subtle.
-
Task-Adaptive Rubrics for GUI Reward Modeling
- Tao Xiong, Xavier Hu, Wenkai Wang, Qinzhuo Wu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang
- 🏛️ Institutions: ZJU, MiLM Plus, Xiaomi
- 📅 Date: August 25, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [reward model], [task-adaptive rubric], [outcome verification], [reinforcement learning], [AdaptRubric]
- 📖 TLDR: AdaptRubric makes GUI reward verification task-adaptive by routing each instruction to a task family for coarse rubric retrieval and then generating instance-level checks for concrete values, scopes, and constraints. It improves offline reward F1 by 3.6 points over the matched baseline average and raises downstream task success by 4.23 points in online reinforcement learning.
-
WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents
- Lin-Fa Lee, Yi-Yu Chang, Kuo-Hui Yeh
- 🏛️ Institutions: National Yang Ming Chiao Tung University
- 📅 Date: August 25, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [safety], [security], [WebMCP], [trust boundary], [prompt injection]
- 📖 TLDR: WebMCP-Phalanx is a dual-layer runtime for the W3C WebMCP proposal that binds each page-exposed tool to its registering principal with cryptographic capability credentials, addressing subject-attribution spoofing, uncontrolled tool lifecycles, and semantic prompt injection.
-
GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis
- Long Zhang, Yuhan Chen, Chaoran Zhang, Wanxia Cao, Kun Huang, Pengzhi Gao, Wei Liu, Jian Luan, Chenliang Li, Lixin Zou
- 🏛️ Institutions: Wuhan University, Xiaomi
- 📅 Date: August 24, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [reinforcement learning], [reward model], [data synthesis], [goal-state anchors], [GSAR]
- 📖 TLDR: GSAR addresses the limited diversity of GUI-agent training environments and the unreliability of scalable evaluators with self-evolving task synthesis and goal-state anchors. The anchors identify task-relevant UI elements in successful states to provide accurate rewards, achieving over 90% offline trajectory-verification accuracy and improving agents on AndroidWorld and a new benchmark.
-
CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajectories
- Zheyuan Deng, Binghang Lu, Hanqi Feng, Shirley Huang, Dianzhuo Wang, Yuanda Xu, Zhiwei Zhang, Yige Sun, Changhong Mou, Runyu Zhang, Yuexing Hao, Barnabas Poczos, Xiaomin Li
- 🏛️ Institutions: Unknown
- 📅 Date: August 23, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [procedural memory], [self-evolution], [multi-model contrastive learning]
- 📖 TLDR: CONTRAMEM learns self-evolving procedural memory for autonomous computer-use agents by contrasting trajectories produced by multiple different backbone models on the same tasks, extracting procedures that generalize across models rather than overfitting to one model's idiosyncrasies.
-
CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents
- Jiaxuan Luo, Zhanfeng Liao, Jiayao Teng, Yuan Wang
- 🏛️ Institutions: Unknown
- 📅 Date: August 23, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [memory], [screenshot caching], [long-horizon tasks], [CausalCache]
- 📖 TLDR: CausalCache is a budgeted screenshot-fidelity restoration method for long-horizon GUI trajectories, conditionally restoring high-fidelity visual context only where it causally affects the agent's next decision instead of caching every screenshot at full resolution.
-
Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents
- Chengyang Gu, Le Zhang, Jingbo Zhou, Yize Chen, Yu Shi, Siqi Bao, Zheng-Fan Wu, Hua Wu, Hui Xiong
- 🏛️ Institutions: HKUST(GZ), Baidu Inc., University of Alberta
- 📅 Date: August 22, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [reinforcement learning], [reward gradient], [trajectory length], [GRPO]
- 📖 TLDR: This work identifies a reward-gradient misalignment in GRPO-style GUI-agent RL caused by trajectory length and proposes length-aware contrastive learning that credits partial progress instead of collapsing every trajectory to a binary success/failure signal.
-
GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
- Md Abrar Jahin, Md Rizwan Parvez
- 🏛️ Institutions: USC, Qatar Computing Research Institute
- 📅 Date: August 22, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [GUI grounding], [spatial reasoning], [diagnostic benchmark], [GUI-Primitives]
- 📖 TLDR: GUI-Primitives is a 994-item contrastive spatial-relation benchmark that isolates spatial-reasoning failures in vision-language GUI grounding from other sources of grounding error.
-
Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and Web
- Qijia Chen, Giulio Jacucci
- 🏛️ Institutions: University of Helsinki
- 📅 Date: August 22, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile], [Web]
- 🔑 Key: [GUI grounding], [lexical coupling], [embedding analysis]
- 📖 TLDR: This work shows that sentence-embedding similarity used for GUI-element grounding is confounded by lexical coupling -- recovery of the visible on-screen label -- rather than genuine semantic understanding, across both mobile and web interfaces.
-
Spine-Branch Coordination for Multi-agent Computer Use
- Mian Zhang, Manasi Sharma, Sheng Zhang, Minglai Yang, Kejian Shi, Ying Liu, Zhiyu Zoey Chen, Daniel Yue Zhang
- 🏛️ Institutions: Scale AI, JHU, University of Texas at Dallas
- 📅 Date: August 22, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [multi-agent], [coordination], [parallel virtual machines]
- 📖 TLDR: This work proposes spine-branch coordination for multi-agent computer use, splitting a long-horizon task between a persistent coordinating spine agent and short-lived branch agents that each execute in a parallel virtual machine.
-
Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning
- Chenghao Zhang, Canran Xiao, SaiSai Hu, Dan Roth
- 🏛️ Institutions: University of Pennsylvania, Sun Yat-sen University, Pace University
- 📅 Date: August 22, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [synthetic environments], [data verification], [training data]
- 📖 TLDR: This work argues that synthetic web environments used to train agents must themselves be verified for trustworthiness, and proposes a pipeline that checks generated environments and trajectories for correctness before they are used as training data.
-
- Pengshuai Yang, Zijing Gao, Xue Yu, Benhui Zhuang, Bo Yuan, Junlan Feng
- 🏛️ Institutions: Jiutian Research, China Mobile
- 📅 Date: August 21, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [automated evaluation], [trajectory evaluation], [step-level reasoning]
- 📖 TLDR: This work automates mobile-agent trajectory evaluation by reasoning about the consequence of each step individually and aggregating step-level judgments into a trajectory-level verdict, rather than scoring only the final outcome.
-
Benchmarking General Mobile Assistants in Challenging Real-World Scenarios
- Yiqi Zhu, Feiyu Gao, Jiaxing Fan, Jiahui Zeng, Minggang Wu, Chenliang Li, Haiyang Xu, Peng Li, Ming Yan, Yang Liu
- 🏛️ Institutions: Tsinghua, Alibaba Group
- 📅 Date: August 21, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [benchmark], [real-world scenarios], [GMA]
- 📖 TLDR: GMA benchmarks general mobile assistants in challenging real-world scenarios, positioned against prior AndroidWorld/MobileWorld-style benchmarks by emphasizing scenario realism over curated task sets.
-
Inducing Task Models from Computer-Use Traces
- Yucheng Jiang, Zora Zhiruo Wang, Ruishi Chen, Diyi Yang
- 🏛️ Institutions: Stanford, CMU
- 📅 Date: August 20, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]
- 🔑 Key: [task model induction], [screenshot traces], [symbolic representation]
- 📖 TLDR: This work induces symbolic task models from raw computer-use traces -- screenshots plus mouse and keyboard events -- turning unstructured demonstration recordings into reusable, inspectable procedures.
-
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
- Tianchen Guan, Xinlei Lin, Royce Cheng-Yue, Xiangjun Wang, Shuyan Zhou
- 🏛️ Institutions: Duke University, Amazon AGI SF Lab
- 📅 Date: August 18, 2026
- 📑 Publisher: COLM 2026
- 💻 Env: [Web]
- 🔑 Key: [benchmark], [component-level evaluation], [failure diagnosis], [ComponentBench]
- 📖 TLDR: ComponentBench targets the under-instrumented middle layer between long-horizon workflow benchmarks and atomic GUI-grounding tests, with a library-agnostic ontology of 97 canonical UI components instantiated as 2,910 programmatically verified, component-centered tasks on modern web UIs.
-
- Sujin Chen, Lijun Li, Tianyi Du, Jing Shao
- 🏛️ Institutions: Shanghai AI Laboratory
- 📅 Date: August 18, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [safety], [security], [environmental injection], [MobileWorldSafety]
- 📖 TLDR: MobileWorldSafety benchmarks Android GUI-agent safety against environmental injection -- indirect prompt injection and adversarial instructions delivered through everyday mobile channels such as notifications, pop-ups, and shared content.
-
CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications
- Linqiang Guo, Li Gu, Zihuan Jiang, Zhixiang Chi, Siobhan Reid, Ziqiang Wang, Yuanhao Yu, Wei Liu, Yang Wang, Tse-Hsun (Peter) Chen
- 🏛️ Institutions: Concordia University, Mila, University of Toronto, McMaster University
- 📅 Date: August 12, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [CoAdapt-GUI], [test-time adaptation], [workflow context], [GRPO], [AndroidWorld-Generalization]
- 📖 TLDR: Mobile GUI agents stay brittle on applications absent from source training, so CoAdapt-GUI performs test-time adaptation that jointly updates a structured workflow context and the policy from the agent's own target-app rollouts and rewards. The workflow context retains transferable procedures, failure modes, and verification rules while excluding app-bound source details, and policy adaptation uses a task-context-matched group-relative objective over a LoRA adapter on a frozen VLM. It reaches 45.0% on AndroidWorld-Generalization against 37.5% for the reported Policy-Only TTA baseline, and lifts AndroidWorld Plus from 38.6% to 52.9%.
-
Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching
- Yuke Li, Xuehan Hou
- 🏛️ Institutions: Unknown
- 📅 Date: August 10, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [GUI grounding], [layout-aware matching], [coordinate hallucination]
- 📖 TLDR: A regression-free GUI grounding pipeline lets a frozen MLLM turn an instruction into a structured visual description, then matches that description against layout-prior candidates. The paper reports gains on ScreenSpot-Pro and Mind2Web while avoiding coordinate-regression fine-tuning.
-
Software Engineering for and with GUI Agent
- Shengcheng Yu, Yuchen Ling, Junyang Xing, Quan Zhou, Chunrong Fang, Zhenyu Chen
- 🏛️ Institutions: Unknown
- 📅 Date: August 10, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [survey], [software engineering], [reliability], [lifecycle]
- 📖 TLDR: This survey reviews 336 GUI-agent papers through a software-engineering lens, covering architectures, evaluation, lifecycle concerns, and deployment gaps. It identifies recurring weaknesses in recovery, safety enforcement, auditability, testing, and maintenance.
-
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
- Weikai Xu, Yunren Feng, Haoxiang Lei, Kun Huang, Yuxuan Liu, Kang Zhao, Xiaolin Hu, Shuo Shang, Bo An
- 🏛️ Institutions: Unknown
- 📅 Date: August 06, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [world model], [executable HTML], [mobile agents], [AppDeltaWorld]
- 📖 TLDR: AppDeltaWorld models mobile GUI transitions as constrained executable HTML updates rather than unconstrained next-screen generation. It uses retrieved app structures and predicted next-screen content to support a mobile-agent training environment and test-time adaptation.
-
- Jiaming Wei, Zekun Wu, Adriano Koshiyama, Maria Perez-Ortiz
- 🏛️ Institutions: Unknown
- 📅 Date: August 06, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [observation routing], [multimodal observation], [WebArena], [VisualWebArena]
- 📖 TLDR: This paper measures when web agents should use text, pixels, or both. Repeated runs show that apparent oracle routing gains are heavily affected by execution noise; learned routers do not consistently beat a fixed mode, while unsolved-task routing can still reduce cost without lowering success.
-
StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection
- Zhuoxin Zhan, Akbar Rafiey, Avery Ma, Leila Pishdad, Layla El Asri
- 🏛️ Institutions: Unknown
- 📅 Date: August 06, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [benchmark], [safety], [indirect prompt injection], [StepJack]
- 📖 TLDR: StepJack studies indirect prompt injection chains in which a harmful goal is split into seemingly innocuous steps across linked pages. Its 480-example benchmark evaluates whether computer-use agents follow these multi-step attacks and releases the attack-generation pipeline.
-
StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
- Linqiang Guo, Wei Liu, Li Gu, Yang Wang, Tse-Hsun (Peter) Chen
- 🏛️ Institutions: Unknown
- 📅 Date: August 06, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [reflection], [UI transitions], [structured prediction], [StepReflect]
- 📖 TLDR: StepReflect recasts after-action reflection for mobile GUI agents as structured transition prediction over paired visual evidence. Its staged training recipe produces a local reflection model that improves several evaluated agent configurations while reducing dependence on repeated frontier-model calls.
-
The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents
- Weiwei Li, Junzhuo Liu, Tong Chu, Hengfu Yu, Wen Li
- 🏛️ Institutions: Unknown
- 📅 Date: August 06, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Mobile]
- 🔑 Key: [hindsight distillation], [mobile agents], [privileged information], [GHD]
- 📖 TLDR: Gated Hindsight Distillation uses the next screenshot as training-only privileged information to rescore an on-policy student's actions. It distills the signal only when that view recovers a missed demonstrated action, improving mobile-agent task success over the compared GRPO baseline.
-
- Longtao Guo, Zelin Zhang, Kaifeng Huang, Yang Shi
- 🏛️ Institutions: Unknown
- 📅 Date: August 05, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Web]
- 🔑 Key: [security], [indirect prompt injection], [authentication], [LoginTrap]
- 📖 TLDR: LoginTrap is a black-box, task-agnostic attack that uses page-specific indirect instructions to make a controlled login flow appear necessary. The paper evaluates end-to-end credential-oriented attack success across web-agent configurations and defenses.
-
GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
- Zichuan Fu, Shirong Wang, Wenlin Zhang, Guojing Li, Yimin Deng, Jingtong Gao, Junjia Qi, Hanyu Yan, Yefeng Zheng, Xiaopeng Li, Wanyu Wang, Xian Wu, Xiangyu Zhao
- 🏛️ Institutions: CityU, Tencent Jarvis Lab, Westlake University
- 📅 Date: August 04, 2026
- 📑 Publisher: arXiv
- 💻 Env: [General GUI]
- 🔑 Key: [GUI-Lens], [GUI grounding], [coarse-to-fine cropping], [coordinate priming], [visual verification]
- 📖 TLDR: GUI-Lens turns GUI grounding into an iterative visual-search process: OCR and detected components provide coordinate references, a general-purpose VLM selects progressively enlarged crops, and separate verification gates reject bad crop or click proposals. Across four grounding benchmarks and three VLM backends, the paper reports gains of up to 24.9 percentage points, with GPT-5.5 reaching 87.9% on ScreenSpot-Pro.
-
- Siqi Fan, Minghao Li, Xiaoqian Ma, Wenhui Tan, Xiusheng Huang, Juntong Wu, Liujie Zhang, Shuo Shang, Weihang Chen
- 🏛️ Institutions: Unknown
- 📅 Date: August 04, 2026
- 📑 Publisher: arXiv
- 💻 Env: [Desktop]