Jingjing Liu, Ziye Huang, Zihao Cheng, Zeming Liu, Jiahong Wu, Yuhang Guo, Kehai Chen, Yunhong Wang
While Graphical User Interface (GUI) agents have shown promising performance in automated device interaction, they primarily depend on static parametric knowledge from pre-training or instruction tuning. This reliance fundamentally limits their ability to handle long-tailed tasks that require explicit procedural knowledge absent from model parameters, often forcing agents to resort to inefficient and brittle trial-and-error exploration. To mitigate this limitation, we introduce \textbf{Proactive Document-Guided Action} for GUI agents in dynamic, open-web environments, a novel paradigm that mirrors human problem-solving by enabling agents to autonomously search for relevant documentation to resolve long-tailed tasks. To evaluate agents' capability in this paradigm, we propose \textbf{DocOS}, a benchmark designed to assess document-guided problem solving in fully interactive environments. DocOS requires agents to autonomously navigate a web browser, locate relevant online documentation, comprehend procedural instructions, and faithfully ground them into executable GUI actions. Extensive experiments reveal that progress is strictly constrained by dual bottlenecks: agents struggle to reliably locate relevant information during proactive search and frequently fail to faithfully ground retrieved instructions into precise actions, pointing toward document-guided interaction as a crucial pathway for enabling self-evolving GUI agents in dynamic environments.
DocOS introduces the concept of Proactive Document-Guided Action for GUI agents — a paradigm where agents must autonomously search the web for documentation to solve long-tailed tasks they lack parametric knowledge for. This mirrors how human users actually operate: when encountering unfamiliar software functionality, they search for official documentation, read it, and then execute. The paper formalizes this into a two-phase pipeline (Proactive Knowledge Retrieval + Document-Grounded Execution) and constructs a benchmark of 817 tasks across 20 applications in Docker-based interactive desktop environments.
The key insight is that existing GUI agents rely on static parametric knowledge and fail on long-tailed, application-specific tasks. By requiring agents to autonomously retrieve and apply documentation, DocOS tests a capability essential for real-world deployment but previously unexamined in benchmarks.
The paradigm of document-guided GUI agents addresses a genuine practical limitation. In enterprise settings, software frequently updates, and agents that can consult documentation dynamically would be far more robust than those relying on static training data. This has clear applications in:
The benchmark itself fills a clear gap in the evaluation landscape (Table 1 effectively demonstrates this). However, the current results suggest the field is far from solving this problem, which positions DocOS more as a long-term challenge benchmark than an immediately actionable contribution.
This paper is highly timely. The GUI agent space has exploded in 2024-2025 with models like CogAgent, UI-TARS, and OpenCUA, and benchmarks like OSWorld and WebArena. The observation that these agents fail on long-tailed tasks is increasingly recognized, and DocOS directly addresses this gap. The connection to retrieval-augmented generation (RAG) in the GUI agent setting is natural and underexplored.
The paper appears at ICML 2026, which places it at a moment when the community is moving beyond basic GUI grounding toward more autonomous and adaptive agent behaviors.
The AutoGen framework experiment (Table 9) is a welcome addition but shows minimal difference from the base model, suggesting that agentic scaffolding alone doesn't resolve the fundamental capability gaps. The paper would benefit from exploring whether chain-of-thought prompting, better document chunking, or multi-turn retrieval strategies could improve performance.
The dataset contribution is the paper's strongest lasting impact — even if current models perform poorly, DocOS provides a meaningful evaluation target as GUI agents improve.
Generated May 19, 2026
PANDO demonstrates stronger scientific impact through concrete, quantifiable improvements: 58.3% success rate on VisualWebArena (beating prior SOTA), 58-61% token reduction, and introduces novel efficiency metrics. It addresses a practical and timely problem (computational cost of AI agents) with a comprehensive framework validated through rigorous ablations on 910 tasks. Paper 2 (DocOS) introduces an interesting paradigm (document-guided agents) and benchmark, but primarily reveals limitations ('dual bottlenecks') without solving them, making it more diagnostic than solution-oriented. PANDO's combination of methodological innovation, strong empirical results, and practical efficiency gains gives it broader and more immediate impact.
Paper 2 offers a novel theoretical contribution by establishing a formal mathematical correspondence between Transformer blocks and k-means clustering, providing a principled interpretation of Chain-of-Thought reasoning on graphs. This theoretical insight has broader implications across multiple fields (NLP, graph learning, interpretability) and offers a unifying framework. Paper 1, while addressing a practical problem in GUI agents with a useful benchmark, is more incremental and application-specific, primarily identifying bottlenecks rather than proposing fundamental solutions. Paper 2's theoretical depth and cross-domain applicability give it higher potential impact.
Paper 2 addresses a critical bottleneck in medical AI—trustworthiness and explainability—by combining LLMs with neuro-symbolic reasoning and fuzzy logic. Its application in clinical decision-making offers profound real-world impact and addresses urgent societal needs. While Paper 1 introduces an innovative benchmark for GUI agents, Paper 2's methodological rigor in formal logic and its potential to safely integrate AI into healthcare systems give it a broader and more significant scientific and societal impact.
Paper 1 introduces a novel paradigm (proactive document-guided action) for GUI agents that addresses a fundamental limitation—reliance on static parametric knowledge—with broader applicability across diverse GUI automation tasks. The benchmark (DocOS) fills an important gap in evaluating agents in dynamic, open-web environments, which is highly relevant given the rapid growth of autonomous agent research. Paper 2, while technically sophisticated with its asymmetric-view training and cognitive user simulator, addresses a narrower problem (proactive task-oriented dialogue/sales). Paper 1's contribution has wider potential impact across the agent research community and more diverse real-world applications.
Paper 2 is more novel in framing “proactive document-guided actions” as a distinct capability and contributes a benchmark (DocOS) that can standardize evaluation and drive follow-on work. Its applications span web automation, enterprise tooling, accessibility, and general agentic RAG, giving broader cross-field impact and timeliness as GUI agents rapidly evolve. Paper 1 offers valuable, rigorous negative/diagnostic findings about LLM optimization limits in hardware-aware code, but its impact is narrower (compiler/kernel optimization) and mainly characterizes failure modes rather than enabling a new scalable research direction or widely reusable artifact.
Paper 2 offers a rigorous methodological advancement by combining interpretable AI with statistical safety guarantees (conformal risk control) for physical systems. Its direct application to safety-critical, real-world infrastructure (wastewater treatment) addressing energy efficiency and greenhouse gas emissions gives it profound real-world impact. While Paper 1 introduces a useful benchmark for GUI agents, Paper 2 spans AI, control theory, and environmental engineering with validated real-world testing, suggesting broader and more immediate scientific and societal impact.
Paper 2 introduces a novel, practical paradigm (proactive document-guided action) with a concrete benchmark (DocOS) for GUI agents, addressing a clear limitation in current agentic systems. It has broader applicability across the rapidly growing GUI/web agent community, identifies specific bottlenecks that can drive future research, and is highly timely given the surge in LLM-based agents. Paper 1, while intellectually elegant in formalizing trust calibration as preferential Bayesian optimization, is more incremental—connecting existing frameworks (GP classification, PBO) to a specific application without novel algorithmic contributions or empirical validation.
Paper 2 introduces a more novel and cross-disciplinary framework (GVG) that bridges neuroscience, computer vision, and MLLMs by using generative visual grounding to translate EEG signals into proxy images. This addresses a fundamental limitation in brain-computer interfaces with broad implications for clinical neuroscience and brain foundation models. Paper 1, while addressing a practical limitation of GUI agents, is more incremental—introducing a benchmark for document-guided actions in a narrower application domain. Paper 2's methodological innovation (trimodal alignment, EEG-to-image generation) and potential impact across neuroscience and AI give it higher scientific impact.
Paper 1 introduces a fundamental methodological innovation by replacing computationally expensive explicit Chain-of-Thought reasoning with latent think tokens for multimodal representations. This addresses a critical bottleneck in deploying reasoning-heavy models, offering broad applicability across foundation models and representation learning. Paper 2 presents a valuable benchmark for GUI agents, but its impact is relatively confined to the agentic workflow subfield, making Paper 1's architectural advancements more likely to achieve widespread scientific and practical impact.
Paper 1 addresses a fundamental bottleneck in LLM alignment and complex reasoning (token-level credit assignment in RLVR). Its proposed algorithmic solution, AMR-SD, targets core methodological challenges in training state-of-the-art reasoning models. While Paper 2 introduces a valuable benchmark for GUI agents, advancements in foundational reasoning capabilities and RL training paradigms typically exert a more profound, widespread impact across the broader AI ecosystem.