Back to Rankings

DocOS: Towards Proactive Document-Guided Actions in GUI Agents

Jingjing Liu, Ziye Huang, Zihao Cheng, Zeming Liu, Jiahong Wu, Yuhang Guo, Kehai Chen, Yunhong Wang

May 18, 2026arXiv:2605.18048v1
cs.AI
Share
Scorecard· 5/16
5.5/10 impact

Abstract

While Graphical User Interface (GUI) agents have shown promising performance in automated device interaction, they primarily depend on static parametric knowledge from pre-training or instruction tuning. This reliance fundamentally limits their ability to handle long-tailed tasks that require explicit procedural knowledge absent from model parameters, often forcing agents to resort to inefficient and brittle trial-and-error exploration. To mitigate this limitation, we introduce \textbf{Proactive Document-Guided Action} for GUI agents in dynamic, open-web environments, a novel paradigm that mirrors human problem-solving by enabling agents to autonomously search for relevant documentation to resolve long-tailed tasks. To evaluate agents' capability in this paradigm, we propose \textbf{DocOS}, a benchmark designed to assess document-guided problem solving in fully interactive environments. DocOS requires agents to autonomously navigate a web browser, locate relevant online documentation, comprehend procedural instructions, and faithfully ground them into executable GUI actions. Extensive experiments reveal that progress is strictly constrained by dual bottlenecks: agents struggle to reliably locate relevant information during proactive search and frequently fail to faithfully ground retrieved instructions into precise actions, pointing toward document-guided interaction as a crucial pathway for enabling self-evolving GUI agents in dynamic environments.

AI Impact Assessments

(1 model)

Scientific Impact Assessment: DocOS — Towards Proactive Document-Guided Actions in GUI Agents

1. Core Contribution

DocOS introduces the concept of Proactive Document-Guided Action for GUI agents — a paradigm where agents must autonomously search the web for documentation to solve long-tailed tasks they lack parametric knowledge for. This mirrors how human users actually operate: when encountering unfamiliar software functionality, they search for official documentation, read it, and then execute. The paper formalizes this into a two-phase pipeline (Proactive Knowledge Retrieval + Document-Grounded Execution) and constructs a benchmark of 817 tasks across 20 applications in Docker-based interactive desktop environments.

The key insight is that existing GUI agents rely on static parametric knowledge and fail on long-tailed, application-specific tasks. By requiring agents to autonomously retrieve and apply documentation, DocOS tests a capability essential for real-world deployment but previously unexamined in benchmarks.

2. Methodological Rigor

Strengths in design:

  • The POMDP formalization is clean and the decomposition into retrieval and execution phases is well-motivated.
  • The benchmark construction pipeline (task construction → document collection → task filtering) is systematic, with 92.5% quality validation on a 20% sample.
  • The evaluation uses both URL-based metrics (TUI, HPP) and task completion rate (TCR), enabling analysis at different granularities.
  • The two complementary evaluation settings (w/ document vs. w/o document, and oracle document) enable diagnosis of bottlenecks.
  • Weaknesses:

  • The improvement from document guidance is surprisingly small — Table 4 shows only 2-8% relative improvement in TCR. This raises questions about whether the paradigm is truly effective, or whether the agents simply cannot utilize documents well. The paper acknowledges this as a "dual bottleneck" but the marginal gains weaken the motivating argument.
  • The baseline set is limited to 7B-8B scale open-source models, with only Qwen3-VL-32B added in the appendix. No proprietary models (GPT-4o, Claude, Gemini) are tested, which limits understanding of whether these bottlenecks are fundamental or scale-dependent.
  • The TCR numbers are extremely low across the board (best: ~17% for UI-TARS-1.5-7B), making it difficult to draw nuanced conclusions about relative performance.
  • The semantic retrieval evaluation (Appendix D) reports a "similarity score of 0.653 with URL-based evaluation results" but doesn't clearly explain what this correlation means or how it was computed.
  • The human quality validation at 92.5% on 20% of data is reasonable but not exhaustive.
  • 3. Potential Impact

    The paradigm of document-guided GUI agents addresses a genuine practical limitation. In enterprise settings, software frequently updates, and agents that can consult documentation dynamically would be far more robust than those relying on static training data. This has clear applications in:

  • IT automation and helpdesk workflows
  • Software testing across version updates
  • Accessibility tools for users with disabilities
  • General-purpose computer-use agents
  • The benchmark itself fills a clear gap in the evaluation landscape (Table 1 effectively demonstrates this). However, the current results suggest the field is far from solving this problem, which positions DocOS more as a long-term challenge benchmark than an immediately actionable contribution.

    4. Timeliness & Relevance

    This paper is highly timely. The GUI agent space has exploded in 2024-2025 with models like CogAgent, UI-TARS, and OpenCUA, and benchmarks like OSWorld and WebArena. The observation that these agents fail on long-tailed tasks is increasingly recognized, and DocOS directly addresses this gap. The connection to retrieval-augmented generation (RAG) in the GUI agent setting is natural and underexplored.

    The paper appears at ICML 2026, which places it at a moment when the community is moving beyond basic GUI grounding toward more autonomous and adaptive agent behaviors.

    5. Strengths & Limitations

    Key Strengths:

  • Novel and well-motivated paradigm: The proactive document search concept is intuitive, practically important, and clearly differentiated from prior work.
  • Comprehensive benchmark design: 20 applications, 817 tasks, three difficulty levels, Docker-based execution — this is a substantial engineering effort.
  • Diagnostic evaluation framework: The two-phase evaluation with oracle/non-oracle settings enables clear bottleneck identification.
  • Detailed error taxonomy: The qualitative error analysis (imprecise localization, non-official reference, execute before retrieval, action grounding failure, context misidentification) provides actionable insights.
  • Notable Weaknesses:

  • Small performance deltas: The 2-8% relative improvement from documents (Table 4) and 4-11% from oracle documents (Table 5) suggest that current agents barely benefit from the paradigm, undermining its practical value at present.
  • Limited baseline diversity: No closed-source frontier models tested; scale effects underexplored.
  • Writing quality issues: Multiple typos and grammatical errors throughout ("konwledge," "oracel," "limitaions," "througy," "retireval"), suggesting rushed preparation.
  • Shallow analysis of why documents don't help more: The paper identifies bottlenecks but doesn't deeply investigate solutions or why the gaps are so severe.
  • Reproducibility concerns: While Docker environments and code are promised, the reliance on live web documentation introduces temporal instability — official documentation pages change over time, potentially invalidating ground-truth URLs and content.
  • Step-based difficulty categorization (Easy/Medium/Hard by number of steps) is simplistic and may not capture true cognitive difficulty.
  • 6. Additional Observations

    The AutoGen framework experiment (Table 9) is a welcome addition but shows minimal difference from the base model, suggesting that agentic scaffolding alone doesn't resolve the fundamental capability gaps. The paper would benefit from exploring whether chain-of-thought prompting, better document chunking, or multi-turn retrieval strategies could improve performance.

    The dataset contribution is the paper's strongest lasting impact — even if current models perform poorly, DocOS provides a meaningful evaluation target as GUI agents improve.

    Rating:5.5/ 10
    Significance 6.5Rigor 5Novelty 6.5Clarity 5

    Generated May 19, 2026

    Comparison History (25)

    Lostvs. PANDO: Efficient Multimodal AI Agents via Online Skill Distillation

    PANDO demonstrates stronger scientific impact through concrete, quantifiable improvements: 58.3% success rate on VisualWebArena (beating prior SOTA), 58-61% token reduction, and introduces novel efficiency metrics. It addresses a practical and timely problem (computational cost of AI agents) with a comprehensive framework validated through rigorous ablations on 910 tasks. Paper 2 (DocOS) introduces an interesting paradigm (document-guided agents) and benchmark, but primarily reveals limitations ('dual bottlenecks') without solving them, making it more diagnostic than solution-oriented. PANDO's combination of methodological innovation, strong empirical results, and practical efficiency gains gives it broader and more immediate impact.

    claude-opus-4-6·May 26, 2026
    Lostvs. Clustering as Reasoning: A $k$-Means Interpretation of Chain-of-Thought Graph Learning

    Paper 2 offers a novel theoretical contribution by establishing a formal mathematical correspondence between Transformer blocks and k-means clustering, providing a principled interpretation of Chain-of-Thought reasoning on graphs. This theoretical insight has broader implications across multiple fields (NLP, graph learning, interpretability) and offers a unifying framework. Paper 1, while addressing a practical problem in GUI agents with a useful benchmark, is more incremental and application-specific, primarily identifying bottlenecks rather than proposing fundamental solutions. Paper 2's theoretical depth and cross-domain applicability give it higher potential impact.

    claude-opus-4-6·May 26, 2026
    Lostvs. Uncertainty Reasoning with Large Language Models for Explainable Disease Diagnosis

    Paper 2 addresses a critical bottleneck in medical AI—trustworthiness and explainability—by combining LLMs with neuro-symbolic reasoning and fuzzy logic. Its application in clinical decision-making offers profound real-world impact and addresses urgent societal needs. While Paper 1 introduces an innovative benchmark for GUI agents, Paper 2's methodological rigor in formal logic and its potential to safely integrate AI into healthcare systems give it a broader and more significant scientific and societal impact.

    gemini-3.1-pro-preview·May 26, 2026
    Wonvs. Unlocking Proactivity in Task-Oriented Dialogue

    Paper 1 introduces a novel paradigm (proactive document-guided action) for GUI agents that addresses a fundamental limitation—reliance on static parametric knowledge—with broader applicability across diverse GUI automation tasks. The benchmark (DocOS) fills an important gap in evaluating agents in dynamic, open-web environments, which is highly relevant given the rapid growth of autonomous agent research. Paper 2, while technically sophisticated with its asymmetric-view training and cognitive user simulator, addresses a narrower problem (proactive task-oriented dialogue/sales). Paper 1's contribution has wider potential impact across the agent research community and more diverse real-world applications.

    claude-opus-4-6·May 22, 2026
    Wonvs. Prior Knowledge or Search? A Study of LLM Agents in Hardware-Aware Code Optimization

    Paper 2 is more novel in framing “proactive document-guided actions” as a distinct capability and contributes a benchmark (DocOS) that can standardize evaluation and drive follow-on work. Its applications span web automation, enterprise tooling, accessibility, and general agentic RAG, giving broader cross-field impact and timeliness as GUI agents rapidly evolve. Paper 1 offers valuable, rigorous negative/diagnostic findings about LLM optimization limits in hardware-aware code, but its impact is narrower (compiler/kernel optimization) and mainly characterizes failure modes rather than enabling a new scalable research direction or widely reusable artifact.

    gpt-5.2·May 20, 2026
    Lostvs. Explainable Wastewater Digital Twins: Adaptive Context-Conditioned Structured Simulators with Self-Falsifying Decision Support

    Paper 2 offers a rigorous methodological advancement by combining interpretable AI with statistical safety guarantees (conformal risk control) for physical systems. Its direct application to safety-critical, real-world infrastructure (wastewater treatment) addressing energy efficiency and greenhouse gas emissions gives it profound real-world impact. While Paper 1 introduces a useful benchmark for GUI agents, Paper 2 spans AI, control theory, and environmental engineering with validated real-world testing, suggesting broader and more immediate scientific and societal impact.

    gemini-3.1-pro-preview·May 20, 2026
    Wonvs. Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use

    Paper 2 introduces a novel, practical paradigm (proactive document-guided action) with a concrete benchmark (DocOS) for GUI agents, addressing a clear limitation in current agentic systems. It has broader applicability across the rapidly growing GUI/web agent community, identifies specific bottlenecks that can drive future research, and is highly timely given the surge in LLM-based agents. Paper 1, while intellectually elegant in formalizing trust calibration as preferential Bayesian optimization, is more incremental—connecting existing frameworks (GP classification, PBO) to a specific application without novel algorithmic contributions or empirical validation.

    claude-opus-4-6·May 20, 2026
    Lostvs. Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs

    Paper 2 introduces a more novel and cross-disciplinary framework (GVG) that bridges neuroscience, computer vision, and MLLMs by using generative visual grounding to translate EEG signals into proxy images. This addresses a fundamental limitation in brain-computer interfaces with broad implications for clinical neuroscience and brain foundation models. Paper 1, while addressing a practical limitation of GUI agents, is more incremental—introducing a benchmark for document-guided actions in a narrower application domain. Paper 2's methodological innovation (trimodal alignment, EEG-to-image generation) and potential impact across neuroscience and AI give it higher scientific impact.

    claude-opus-4-6·May 19, 2026
    Lostvs. TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens

    Paper 1 introduces a fundamental methodological innovation by replacing computationally expensive explicit Chain-of-Thought reasoning with latent think tokens for multimodal representations. This addresses a critical bottleneck in deploying reasoning-heavy models, offering broad applicability across foundation models and representation learning. Paper 2 presents a valuable benchmark for GUI agents, but its impact is relatively confined to the agentic workflow subfield, making Paper 1's architectural advancements more likely to achieve widespread scientific and practical impact.

    gemini-3.1-pro-preview·May 19, 2026
    Lostvs. AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment

    Paper 1 addresses a fundamental bottleneck in LLM alignment and complex reasoning (token-level credit assignment in RLVR). Its proposed algorithmic solution, AMR-SD, targets core methodological challenges in training state-of-the-art reasoning models. While Paper 2 introduces a valuable benchmark for GUI agents, advancements in foundational reasoning capabilities and RL training paradigms typically exert a more profound, widespread impact across the broader AI ecosystem.

    gemini-3.1-pro-preview·May 19, 2026