Back to Rankings
Checking arXiv for updates...

KWBench: Measuring Unprompted Problem Recognition in Knowledge Work

Ankit Maloo

Apr 17, 2026arXiv:2604.15760v1
cs.AIcs.GT
Share
Scorecard· 5/16
6.5/10 impact

Abstract

We introduce the first version of KWBench (Knowledge Work Bench), a benchmark for unprompted problem recognition in large language models: can an LLM identify a professional scenario before attempting to solve it. Existing frontier benchmarks have saturated, and most knowledge-work evaluations to date reduce to extraction or task completion against a specification. KWBench targets the step before that: recognizing the governing structure of the situation from raw inputs alone. The benchmark contains 223 tasks sourced from practitioners across acquisitions, contract negotiations, clinical pharmacy, organizational politics, fraud analysis, and incentive design. Each task encodes a formal game-theoretic pattern (principal-agent conflict, signaling, mechanism design failure, strategic omission, coalitional dynamics, strategic interdependence) and carries structured ground truth recording the expert reading of the situation and the anticipated failure modes. Models receive raw data and a task prompt with no indication of problem type. Scoring is a three-tier rubric gated by a mandatory conjunctive check. Mandatory criteria encode the predicted wrong paths. We evaluate 16 models. The best model passes on 27.9% of tasks. The top two models agree on only 31.7% of their passes. Among the top 8, 44 tasks are solved by exactly one model; routing across the top 8 covers 50.7% of the benchmark, nearly double the best single model. Conditional on passing, quality scores converge (approx 83% across models); unconditional scores do not. Same models articulate the relevant game-theoretic concept correctly when asked, then fail to apply it unprompted. We release KWBench to shift how frontier models are evaluated on knowledge work, scoring them on whether they recognize the right problem from the situation alone, not only on how well they execute once the problem has been framed for them.

AI Impact Assessments

(3 models)

Scientific Impact Assessment: KWBench

Core Contribution

KWBench introduces a genuinely novel evaluation axis: unprompted problem recognition — whether LLMs can identify the correct framing of a professional scenario before executing on it. The paper argues convincingly that existing benchmarks test execution given a correctly specified problem, while real knowledge work requires first recognizing *what* the problem is. The benchmark contains 223 tasks spanning acquisitions, contract negotiations, clinical pharmacy, organizational politics, fraud analysis, and incentive design, each encoding a game-theoretic pattern (signaling, principal-agent, mechanism design failure, etc.) that must be recognized from raw inputs alone.

The conceptual contribution is sharp and well-articulated: the distinction between perfect-information problems (where benchmarks saturate) and imperfect-information games (where knowledge workers actually operate) provides a clean theoretical motivation. The "chess vs. poker" framing is effective and the mapping of six game-theoretic patterns to professional scenarios is intellectually coherent.

Methodological Rigor

Strengths in design: The "don't instruct, measure" principle is methodologically sound. The mandatory conjunctive gate — score zero if any core criterion fails — is well-justified by the argument that domain expertise is conjunctive (one missed liability clause compromises the entire contract review). The three-stage rubric construction (metadata specification → multi-model generation → human synthesis) provides some quality control.

Significant weaknesses: The paper acknowledges but does not resolve several critical methodological gaps:

1. No human baseline. This is the most damaging omission. Without knowing how human experts perform, we cannot interpret the 27.9% pass rate. Is this benchmark impossibly hard, poorly calibrated, or genuinely measuring a meaningful gap? The authors claim practitioner validation of "realism and difficulty calibration," but structured consultations are not the same as having experts take the test.

2. Single-judge evaluation. All scoring relies on Gemini 3 Flash as judge. While the binary, verifiable nature of criteria helps, the paper provides no inter-rater reliability data, no multi-judge comparison, and no analysis of judge failure modes. Given the centrality of the mandatory gate, even a small systematic bias in judging could dramatically affect results.

3. No recognition ablation. The paper's central claim — that models *possess* game-theoretic knowledge but fail to *apply* it unprompted — is stated but not formally tested. Running the same tasks with explicit game-theoretic hints would be a straightforward and critical experiment.

4. Rubric construction concerns. The rubrics are generated by LLMs and synthesized by a single author. The "practitioner validation" is described vaguely ("structured consultations") without details on how many practitioners reviewed how many tasks, what their credentials were, or what the disagreement rate was.

5. Best-of-3 evaluation protocol introduces selection bias. Taking the best run inflates reported performance, though the authors note modest variance (1-3 percentage points).

Potential Impact

The paper targets a real and important gap. As LLMs are increasingly deployed in advisory roles — drafting memos, reviewing contracts, triaging decisions — the failure mode of "polished analysis of the wrong problem" is genuinely dangerous. The examples are vivid and compelling: a PIP that would fail in court, an acquisition offer misread as a valuation exercise, a deal celebrated when the buying process hasn't started.

Practical implications are significant:

  • The finding that no single model dominates (Jaccard overlap of 31.7% between top two models) has direct implications for system architecture, suggesting ensemble/routing approaches.
  • The "cooperative default" analysis — that RLHF training may systematically suppress adversarial reasoning — is a valuable hypothesis for the alignment community.
  • The 107 unsolved tasks provide a concrete capability target.
  • Limitations on impact: The benchmark's professional focus means it primarily benefits enterprise AI deployment rather than broader ML research. The task count (223) is modest, and the domain skew toward Western corporate norms (acknowledged by the authors) limits generalizability.

    Timeliness & Relevance

    This is highly timely. Frontier benchmarks are saturating (MMLU, HumanEval), and the field is actively searching for meaningful evaluation axes. The deployment of LLMs in knowledge work is accelerating faster than our ability to measure their fitness for it. The specific failure mode KWBench targets — confident, well-structured output that answers the wrong question — is arguably the most dangerous failure mode in professional AI deployment, precisely because it evades casual quality checks.

    Strengths

    1. Novel and well-defined evaluation axis. The recognition-execution distinction is crisp, empirically supported (decoupled scores, knowledge-application gap), and practically important.

    2. Rich task design. The detailed walkthrough examples (Appendix B, C) demonstrate genuine depth. The PIP example alone is a masterclass in what "testing for pitfalls, not correct answers" means.

    3. Surprising empirical findings. The disjoint recognition profiles (low Jaccard overlap), the coverage analysis showing every top-8 model contributes unique passes, and the convergence of conditional scores are genuinely informative results.

    4. Structured expert annotations (5,800 items across 223 tasks) are a separable, reusable contribution.

    5. Transparent about limitations. The paper clearly states what it does not do.

    Limitations

    1. Absence of human baselines undermines interpretability of all absolute numbers.

    2. Single author, single judge creates concentration risk in both construction and evaluation.

    3. Small benchmark size (223 tasks) with only 85 in the core game-theoretic category.

    4. Potential subjectivity in "correct" framing. While the paper argues criteria are objective (verifiable traps), reasonable experts might disagree on whether a given scenario truly requires adversarial framing.

    5. Reproducibility concerns. The reliance on practitioner knowledge that is "rarely articulated" makes independent validation difficult.

    6. 38 tasks adapted from existing benchmarks without clear analysis of how they compare to the 185 original tasks.

    Overall Assessment

    KWBench makes a compelling conceptual contribution by identifying and operationalizing a previously unmeasured capability axis. The empirical findings are interesting and the task design is sophisticated. However, the methodological gaps — particularly the absence of human baselines, recognition ablation, and multi-judge validation — prevent the paper from fully substantiating its claims. The paper is stronger as a provocation and framework than as a definitive measurement instrument. It opens an important research direction but needs significant validation work to become a trusted evaluation standard.

    Rating:6.5/ 10
    Significance 7.5Rigor 5Novelty 8Clarity 8.5

    Generated Apr 20, 2026

    Comparison History (37)

    Wonvs. LLMs can persuade only psychologically susceptible humans on societal issues, via trust in AI and emotional appeals, amid logical fallacies

    Paper 2 introduces a novel benchmark targeting a critical, underexplored gap in LLM capabilities—unprompted problem recognition. By moving beyond saturated task-completion evaluations, it sets a new target for AI reasoning and agentic systems in real-world knowledge work. Benchmarks that expose fundamental model limitations typically drive significant, field-wide progress and garner high citations, giving it a higher potential for broad scientific impact than the behavioral findings in Paper 1.

    gemini-3-pro-preview·May 5, 2026
    Wonvs. Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys

    KWBench introduces a novel evaluation paradigm—unprompted problem recognition—that addresses a fundamental gap in LLM benchmarking. Its finding that models can articulate game-theoretic concepts but fail to apply them unprompted reveals a critical limitation with broad implications for AI deployment in knowledge work. The benchmark spans multiple professional domains, offers a public resource for the community, and addresses benchmark saturation, a timely concern. Paper 1, while methodologically rigorous and practically useful, represents an incremental optimization within an established framework (Prediction-Powered Inference) with narrower scope.

    claude-opus-4-6·May 5, 2026
    Wonvs. When AI reviews science: Can we trust the referee?

    Paper 1 introduces a fundamentally novel evaluation paradigm for LLMs (unprompted problem recognition), addressing a critical gap in current AI benchmarks. While Paper 2 tackles an important and timely issue (AI in peer review), Paper 1's benchmark has broader implications across multiple disciplines of knowledge work and cognitive AI evaluation. By shifting the focus from task execution to problem identification, KWBench has the potential to guide the next generation of frontier model development, yielding a wider and more foundational scientific impact.

    gemini-3-pro-preview·May 5, 2026
    Lostvs. Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

    Agent-World addresses a fundamental bottleneck in agent training—lack of realistic environments and continuous learning mechanisms—with a scalable, self-evolving framework validated across 23 benchmarks. Its contributions to environment synthesis, multi-environment RL, and self-evolving training have broad applicability across the rapidly growing AI agent ecosystem. While KWBench introduces a valuable and novel evaluation paradigm (unprompted problem recognition), it is primarily a benchmark contribution with a narrower scope. Agent-World's methodological contributions and demonstrated scaling laws offer more transformative potential for advancing general agent intelligence.

    claude-opus-4-6·Apr 21, 2026
    Lostvs. From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?

    Paper 2 offers a more general, theory-driven reframing of LLMs-as-annotators as latent opinion estimators, deriving conditions/regimes where LLMs can statistically outperform humans and where they cannot—insights broadly applicable across HCI, NLP evaluation, computational social science, and survey/measurement. This has clear real-world implications for scaling subjective annotation and estimating subgroup perspectives, and is timely given widespread LLM annotation use. Paper 1 is novel as a benchmark targeting unprompted problem recognition in knowledge work, but its impact is narrower (benchmark-centric, 223 tasks) and more domain-specific, with less immediate cross-field theoretical leverage.

    gpt-5.2·Apr 21, 2026
    Wonvs. ASMR-Bench: Auditing for Sabotage in ML Research

    KWBench introduces a fundamentally novel evaluation paradigm—unprompted problem recognition—that addresses a significant gap in LLM benchmarking. Its findings (best model at 27.9%, low inter-model agreement, the recognition-application gap) reveal deep limitations in current frontier models with broad implications across knowledge work domains. The benchmark spans multiple professional fields and introduces rigorous game-theoretic grounding. ASMR-Bench addresses an important but narrower AI safety concern (sabotage detection in ML codebases) with a smaller benchmark (9 codebases). KWBench's broader applicability, novel conceptual contribution, and richer empirical findings suggest higher impact.

    claude-opus-4-6·Apr 20, 2026
    Wonvs. Grounding Clinical AI Competency in Human Cognition Through the Clinical World Model and Skill-Mix Framework

    Paper 1 introduces an empirical benchmark addressing a critical, unsolved capability in LLMs (unprompted problem recognition), applying broadly across multiple domains of knowledge work. Its quantitative evaluation of frontier models reveals a significant performance gap, which is highly likely to drive immediate follow-up research and model optimization. While Paper 2 offers a valuable theoretical framework for clinical AI, conceptual models typically have a slower, more domain-restricted impact compared to actionable, broadly applicable AI benchmarks.

    gemini-3-pro-preview·Apr 20, 2026
    Wonvs. Do Agent Rules Shape or Distort? Guardrails Beat Guidance in Coding Agents

    Paper 2 introduces a novel benchmarking paradigm (unprompted problem recognition) that addresses a critical gap in LLM evaluation across diverse knowledge-work domains. While Paper 1 offers valuable, immediate insights for coding agents, Paper 2 has a broader scientific impact by fundamentally shifting how we evaluate and develop frontier models for complex, real-world reasoning and autonomous problem framing.

    gemini-3-pro-preview·Apr 20, 2026
    Wonvs. Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents

    Paper 2 introduces a novel benchmark (KWBench) addressing a critical gap in LLM evaluation (unprompted problem recognition) where frontier models currently struggle. Challenging new benchmarks that reveal fundamental model limitations typically drive immediate follow-up research, model development, and high citation rates, leading to higher measurable scientific impact compared to theoretical or survey frameworks like the one proposed in Paper 1.

    gemini-3-pro-preview·Apr 20, 2026
    Wonvs. Agent-Aided Design for Dynamic CAD Models

    Paper 1 introduces a paradigm shift in LLM evaluation, focusing on unprompted problem recognition rather than mere execution. Because existing AI benchmarks are rapidly saturating, a rigorous benchmark evaluating situational awareness addresses a critical bottleneck in AI development. While Paper 2 offers a valuable advance in AI-aided CAD design, its impact is largely confined to mechanical engineering and manufacturing. Paper 1's findings regarding LLM failure modes in structural reasoning have broad implications across AI alignment, cognitive science, and diverse knowledge-work domains like law, medicine, and finance, giving it a much wider scientific footprint.

    gemini-3-pro-preview·Apr 20, 2026