Every score comes with the model's reasoning on the paper page.
"Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories" scored 6.5/10 predicted impact on @kurateorg — Reproducible 8.0 · Clarity 7.5 · Rigor 7.0