Back to Rankings

Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf

Yu-Yu Yang, Ti-Rong Wu, Hung Guei, Hsing-Yu Chen, I-Chen Wu

Sep 11, 2026arXiv:2609.12446v1
cs.AIcs.CL
Share
Scorecard· 16/16
5.5/10 impact

A well-executed, timely diagnostic benchmark with a clear conceptual reframing and strong reproducibility, but bounded impact due to a single narrow environment and receiver-side-only scope.

Abstract

Social-deduction games such as Werewolf are increasingly used to evaluate LLM agents, but existing evaluations often rely on final game outcomes. We propose a belief-shift evaluation benchmark in Werewolf for analyzing communication skills through belief updating. Using LLM-played games, we annotate suspicion and accusation messages and measure how an observing village-side model's beliefs change after each message. We evaluate 40 open-weight LLM configurations on 1,224 annotated messages. Our results show that larger models better distinguish true wolves from villagers based on game history, but accusations still strongly influence their beliefs. Models become more suspicious of the accused target and less suspicious of the accuser, especially when the accuser is trusted, even if the accuser is wolf-aligned. Larger models better resist accusations from accusers they already distrust. Overall, our findings suggest that current open-weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication. Our benchmark and code are available at https://rlg.iis.sinica.edu.tw/papers/werewolf-accusation-benchmark.

AI Impact Assessments

(1 model)

Scientific Impact Assessment

Core Contribution. The paper introduces a belief-shift evaluation methodology for LLM agents in the social-deduction game Werewolf. Rather than scoring agents by terminal outcomes (win/loss/win-rate), it measures how an observing village-side model's suspicion of two players — an *accuser* and an *accused target* — changes immediately before and after an accusation message. The key conceptual move is to isolate a specific receiver-side sub-skill: whether a model can integrate message *content* with *source trustworthiness* when updating beliefs. This reframes social-deduction evaluation as a diagnostic of theory-of-mind-like reasoning. The headline empirical finding is that open-weight models (1B–120B) systematically shift suspicion toward accused targets and away from accusers, most strongly when the accuser is already trusted — even when that accuser is secretly wolf-aligned — indicating a credulity failure mode. Larger models resist accusations from already-distrusted sources but still fail to discount trusted-but-deceptive accusers.

Methodological Rigor. The design is clean and the analysis is honest. The pipeline separates the game-generation model pool from the 40 evaluated (checkpoint, reasoning-mode) configurations, avoiding some self-evaluation confounds. Claims of scale-dependence are backed by Spearman correlations with explicit p-values, and all means carry 95% confidence intervals (Appendix B.3, Tables 4–5). Annotation reliability is checked with a 150-item stratified human re-labeling (96% agreement). Ablations separate true vs. false accusations (Table 8) to show the trend is not an artifact of mixing accusation types, and a frontier closed-source preliminary study (Table 7) is included. The authors candidly flag the most serious methodological concern themselves: providing the model its own prior belief may induce an anchoring effect, which they leave unresolved. Other gaps: a single LLM annotator, a single environment, and receiver-side-only measurement.

Potential Impact. The work sits in an active niche — using interactive games to probe LLM communication and ToM. The belief-shift instrument is a reusable methodological building block that could be ported to Avalon, negotiation, or debate settings, and the paired pre/post protocol is general. The identified "trusted-source credulity" failure mode has clear relevance to AI-safety and prompt-injection/manipulation research, which the authors explicitly note in the Ethics section. That said, impact is bounded by the narrowness of the tested setting and the fact that this is a diagnostic benchmark rather than a method that improves agent capability. It is more likely to be cited as a useful evaluation tool by a moderate slice of the LLM-agent subfield than to redefine how the field approaches the problem.

Timeliness & Relevance. Highly timely. LLM agents, multi-agent communication, and ToM evaluation are all rapidly growing areas, and dissatisfaction with outcome-only game metrics is a real, current bottleneck the paper directly targets. The frontier-model preliminary result — that frontier systems appear to *reverse* the credulity pattern for distrusted wolf accusers — is a genuinely interesting hook that suggests an emerging capability threshold worth tracking.

Strengths. (1) A well-motivated reframing from outcome metrics to intermediate belief dynamics. (2) Broad model coverage (40 configurations, 29 checkpoints, reasoning on/off). (3) Strong reproducibility: released benchmark and code, full model list, licenses, inference setup (vLLM version, precision, context lengths, JSON-schema-constrained outputs), and compute cost (288 GPU-hours). (4) Unusually thorough and honest limitations/ethics discussion. (5) Concrete, interpretable findings supported by statistics.

Weaknesses / Gaps. (1) Single environment; the authors correctly note Werewolf's dense, salient accusations may *amplify* the credulity effect, limiting generalisability. (2) Receiver-side only — no speaker-side analysis, so the "communication skill" picture is incomplete. (3) The anchoring concern from feeding priors back to the model is unaddressed and could partially drive the observed shifts. (4) All dialogue is LLM-generated self-play; whether the same patterns hold with human interlocutors is untested. (5) The core techniques (prompting, LLM annotation, Likert elicitation) are individually standard; novelty is in the framing and measurement design rather than technical machinery. (6) The dataset (1,224 messages) is modest and used entirely as a single eval set.

Additional Observations. The paper is a low-resource, inference-only study (a four-GPU workstation, 72 wall-clock hours), making it easy for others to reproduce or extend — a positive for adoption. It lightly corroborates prior persuasion-susceptibility findings (Xu et al. 2024a; Bajaj & Tiganj 2026) in a new observer-side, multi-agent setting, but does not contest any prior claim. Notably, the citation set and model names are future-dated (GPT-5.6, Claude Opus 4.8, Gemma 4, Qwen3.6), which does not affect the methodological assessment but is worth flagging.

Overall, this is a solid, well-executed diagnostic-benchmark paper with a clean idea, honest analysis, and good reproducibility, but with impact constrained by a single narrow environment and receiver-side-only scope. It is likely to be a useful, moderately cited contribution rather than a field-shifting one.

Rating:5.5/ 10
Significance 5.5Rigor 6.5Novelty 6Clarity 8

Generated Sep 14, 2026

Comparison History (0)

No comparisons yet.