Every score comes with the model's reasoning on the paper page.
"Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf" scored 5.5/10 predicted impact on @kurateorg — Clarity 8.0 · Reproducible 8.0 · Rigor 6.5