
Every score comes with the model's reasoning on the paper page.
"MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education" scored 5.0/10 predicted impact on @kurateorg — Reproducible 7.0 · Clarity 6.5 · Rigor 6.0