Every score comes with the model's reasoning on the paper page.
"When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation" scored 7.0/10 predicted impact on @kurateorg — Clarity 8.0 · Rigor 7.5 · Significance 7.0