Every score comes with the model's reasoning on the paper page.
"K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments" scored 7.5/10 predicted impact on @kurateorg — Rigor 8.5 · Reproducible 8.5 · Novelty 8.0