Back to Rankings

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng

Sep 16, 2026arXiv:2609.19088v1
cs.AIcs.CLcs.CV
Share
Scorecard· 16/16
5.0/10 impact

A well-constructed, timely benchmark filling a real niche in artistic/educational VLM evaluation, but moderate scale and confirmatory findings cap its likely field-wide influence to a specialized subcommunity.

Abstract

Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must interpret artistic imagery, understand its semantic, affective, and cultural content, and reason about visual context to support meaningful interaction. However, existing benchmarks primarily focus on real-world images or domain-specific educational reasoning, providing limited coverage of artistic educational content. To address this gap, we introduce MUSE, a benchmark for evaluating large vision-language models on artistic image understanding in situated educational applications. MUSE decouples image annotation from question generation, enabling diverse tasks with controllable difficulty while reducing annotation effort. It comprises twelve tasks spanning visual perception, semantic and affective interpretation, culture understanding, and compositional reasoning, together with diverse artistic images deliberately curated to center Singaporean and Southeast Asian multicultural contexts alongside Western art traditions, covering multiple themes and difficulty levels. Evaluation of open-source and proprietary models reveals substantial disparities across capability dimensions, particularly in affective interpretation and compositional reasoning. Our analysis further identifies common failure modes and key challenges for developing trustworthy multi-modal models for education. We hope MUSE will serve as a standardized benchmark for advancing multi-modal understanding in situated educational applications.

AI Impact Assessments

(1 model)

Scientific Impact Assessment: MUSE

1. Core Contribution

MUSE is a vision-language benchmark targeting an under-served niche: understanding of *artistic* imagery (paintings, illustrations, cartoons) in *situated educational* contexts, particularly image-based language learning. The paper makes two interlocking contributions. First, a dataset of 2,400 questions over 1,174 commissioned artworks spanning 12 tasks organized into five capability dimensions (visual perception, semantic understanding, affective interpretation, compositional reasoning, cultural understanding), deliberately weighted toward Singaporean/Southeast Asian cultural content alongside Western traditions. Second, a methodological construction framework—"annotation-first, task-generative"—that decouples reusable structured image annotations from task-specific question instantiation, enabling annotation reuse, semantic consistency across tasks, and controllable difficulty/format. The paper evaluates 30 open-source and proprietary VLMs and reports that models are competent at scene/activity recognition but weak at affective interpretation, viewpoint-aware spatial reasoning, and evidence grounding.

2. Methodological Rigor

The construction pipeline is a genuine strength: 127 trained annotators, one-annotator-plus-two-reviewer consensus protocol, standardized Label Studio interface, pilot-stage guideline refinement, and documented annotation cost/time. The evaluation is broad (30 models across scales/families), uses deterministic decoding, a unified parser, and reports supplementary analyses that add credibility—inter-task Spearman correlations (demonstrating non-redundancy), cross-benchmark correlations (arguing complementarity to MMBench, AI2D, BLINK, AICA-Bench), invalid-response-rate analysis, taxonomy-dimension breakdowns, and a human-performance comparison establishing the ceiling. The correlation analysis showing that Jigsaw Puzzle is nearly orthogonal to other tasks, and that general-benchmark performance transfers unevenly, is a thoughtful validation of the benchmark's added value.

Weaknesses temper this. The human study is thin—only two annotators on 20 sampled questions per task—which is too small to serve as a robust human baseline despite being framed as a "performance envelope." The Relative Position task yields near-floor scores across all models (0–10%), raising a question of whether the task/answer format is well-posed rather than merely hard; a model scoring 0–4 out of 100 may reflect an ambiguous ground-truth specification (the paper itself notes forced-relation bias against a "None" ground truth). Open-ended tasks are scored purely by embedding cosine similarity, a metric with known weak correlation to human judgment of correctness. No inter-annotator agreement statistics are reported. Some evaluated model names (e.g., "GPT-5.6-Sol," "InternVL3-78b" appearing inconsistently with "InternVL3-38b") suggest naming/labeling inconsistencies that reduce confidence in careful proofreading.

3. Potential Impact

The benchmark fills a real gap: existing multimodal benchmarks (MMMU, MathVista, ScienceQA, MMBench) center natural images and disciplinary reasoning, while art/affect/culture benchmarks (ArtEmis, ArtELingo, VQArt-Bench, AICA-Bench, CVQA) each address a single dimension. MUSE's joint, multi-dimensional coverage over stylized imagery is distinctive. The most transferable contribution is arguably the annotation-first/task-generative construction philosophy, which other benchmark builders could adopt to reduce annotation cost and improve cross-task consistency—though this idea is a sensible engineering refinement rather than a conceptual breakthrough. Real-world relevance to AI tutoring and culturally-situated educational tools is plausible, and the Southeast-Asian cultural centering has value for representation in a Western-dominated benchmark landscape. However, benchmark impact is heavily gated by adoption: whether the community treats MUSE as a standard reference, and whether the released code/dataset are maintained and used. Educational-VLM benchmarks are a crowded space, and datasets of ~2,400 questions are moderate in scale, which may limit its staying power against larger, more heavily-marketed benchmarks.

4. Timeliness & Relevance

Highly timely. VLM evaluation is a fast-moving, actively contested area, and the specific concerns MUSE raises—affective reasoning, cultural grounding, spatial/compositional reasoning, evidence-based grounding—are precisely the frontier weaknesses the field is currently probing. The educational-AI application framing aligns with strong current interest in AI tutoring. The commissioned-artwork approach also sidesteps the training-data-contamination problem afflicting scraped benchmarks, which is a meaningful contemporary advantage.

5. Strengths & Limitations

Strengths: original commissioned imagery (contamination-resistant); reusable annotation framework; broad 30-model evaluation; rich diagnostic analyses (correlation, error mode, cascading-failure case studies); clear identification of concrete failure modes (region-activity binding, viewpoint-aware spatial uncertainty, affective grounding cascade); cultural diversity emphasis.

Limitations: moderate scale; weak human baseline; reliance on embedding similarity for open-ended scoring; a possibly ill-posed spatial task producing near-zero scores; no IAA reporting; findings (models weak at affective/compositional reasoning) largely confirm what the field already suspects rather than overturning beliefs; some presentation inconsistencies. The core empirical conclusions are useful but unsurprising.

Reproducibility: Strong on the evaluation side—dtypes, generation configs, model list, and metrics are specified, and code/dataset links are advertised. Reproducing the *dataset construction* would be far harder given the bespoke commissioned artwork and annotator pool.

Overall: A solid, well-executed benchmark paper contributing a genuinely under-covered evaluation niche and a reusable construction methodology. Its impact will likely be as a useful specialized resource cited within multimodal-evaluation and educational-AI subcommunities rather than a field-reshaping benchmark. The methodological framing (annotation decoupling, contamination-free commissioned art) is its most durable idea.

Rating:5/ 10
Significance 5Rigor 6Novelty 6Clarity 6.5

Generated Sep 17, 2026

Comparison History (0)

No comparisons yet.