Every score comes with the model's reasoning on the paper page.
"K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments" scored 7.0/10 predicted impact on @kurateorg — Reproducible 9.0 · Rigor 8.5 · Significance 7.5