Back to Rankings

AI for Games in the Foundation Model Era

Meng Luo, Yanlin Li, Hao Li, Hongzhan Lin, Pengfei Zhou, Tianjie Ju, Ran Zhang, Yeying Jin

Sep 15, 2026arXiv:2609.16679v1
cs.AI
Share
Scorecard· 13/16
6.0/10 impact

Comprehensive, rigorous, timely survey with a genuinely useful cross-role framing, but conceptual rather than method-defining and overlapping prior role-based surveys.

Abstract

Foundation models, alongside advances in learned game-world models, are reshaping AI across the game lifecycle. Beyond playing games, recent systems model players and game dynamics, support design and development, adapt player-facing experiences at runtime, and evaluate resulting artifacts. Yet these directions have evolved largely separately, obscuring which capabilities transfer across settings and which remain tied to particular games, engines, interfaces, or player populations. We organize the literature into six roles according to the immediate use of AI output: playing and acting; modeling players and games; designing games; building and maintaining games; generating and adapting at runtime; and testing and evaluating games. For each role, we examine what structure is supplied by the game or workflow, what AI learns or produces, which capabilities and artifacts transfer across settings and roles, and what evidence supports the claims. We identify cross-role connections: trajectories train world models, learned environments provide experience for agents, design specifications drive executable implementations, and play or testing feedback guides revision. However, control schemes, rules, engine interfaces, state representations, and player contexts often remain setting-specific, so downstream claims require validation in the target setting. Evaluation is most standardized for bounded game playing and selected learned environments, while persistent state in learned worlds, repeated software revision, validated player modeling, sustained runtime adaptation, and representative automated testing remain less established. The central challenge is to reuse or transfer outputs and capabilities across roles while re-establishing evidence for effectiveness in the game-specific contexts where they are used.

AI Impact Assessments

(1 model)

Scientific Impact Assessment

Paper type: Survey / literature synthesis (large-scale, ~120 pages, extensive taxonomy, tables, and curated resource repository).

1. Core Contribution

The paper organizes the sprawling literature on foundation models in games into six functional roles keyed to the *immediate use* of AI output: playing/acting, modeling players and games, designing, building/maintaining, generating/adapting at runtime, and testing/evaluating. Its genuine novelty over prior role-based surveys (e.g., Gallotta et al.) is twofold: (1) an explicit cross-role connection analysis tracing how artifacts pass between roles (trajectories→world models, learned environments→agent training, specs→executable code, test feedback→revision), and (2) a disciplined evidence-interpretation framework that repeatedly separates "artifact reuse" from "capability transfer," and insists downstream claims be re-validated in target settings. The central thesis—that broader interfaces expand what can be connected but do not remove game-specific structure, and that evidence is strongest for bounded play and weakest for persistent state, sustained adaptation, and player-experience claims—is a coherent organizing argument rather than a mere catalog.

2. Methodological Rigor

For a survey, rigor lies in comprehensiveness, critical stance, and evaluation discipline, and here it is unusually strong. The paper consistently interrogates *what evidence supports each claim*, distinguishing perceptual quality from mechanics correctness from persistent state; progress from completion; diversity from representativeness; coverage from correctness. Section 9's five cross-role interpretive dimensions (standardization, execution grounding, scope/transfer, horizon/revision, external validation) and its attention to statistical support (resampling units, confidence intervals, judge independence, contamination) exceed typical survey standards. The taxonomy assignment rules (primary/secondary roles) are principled. The main weakness is inherent to the genre: the synthesis is qualitative and the authors acknowledge it does not rank prevalence or maturity.

3. Potential Impact

As a reference map for a fast-growing, fragmented area, this is likely to be cited as a useful entry point and taxonomy by researchers spanning RL agents, world/video models, PCG, LLM development agents, and automated game testing. The "reuse vs. transfer" distinction and the role-specific evaluation-target tables could shape how future papers frame claims and design benchmarks. Industry relevance is real—game studios and tooling teams face exactly these lifecycle questions—but the paper is descriptive, not a deployable method, so translational value is indirect. It is unlikely to *change* how the field approaches a problem in a paradigm sense; its impact will be as a well-organized, critically-minded reference and a source of concrete open-problem framings.

4. Timeliness & Relevance

Highly timely. The paper explicitly addresses the fragmentation created by rapid, parallel development across the game lifecycle and the difficulty of comparing capabilities across settings. It captures very recent systems and benchmarks (learned game-world models, development agents, model-based judges). One caveat: it cites near-future/placeholder-seeming models (GPT-6 Astra, Claude Opus 5, Gemini 3), suggesting either a forward-dated draft or synthetic references; this does not undermine the framework but slightly complicates verification of some quantitative examples.

5. Strengths & Limitations

Strengths: exceptional breadth; disciplined skepticism about evidence; the cross-role framing is a genuinely useful lens; strong evaluation/benchmark synthesis with reproducibility and statistical caveats; a public curated repository lowers the barrier to building on it. Limitations: substantial overlap with existing role-based surveys; the contribution is conceptual rather than empirical; extreme length and density reduce accessibility; findings, while well-argued, are largely confirmatory of what informed practitioners suspect (evidence strong for bounded play, weak for persistence/experience). It does not contest a specific prior empirical result so much as caution against a common inferential shortcut.

Other observations: The dataset/resource contribution (awesome-list, system index, benchmark tables) adds durable practical value. The foundationality is moderate—the six-role taxonomy could become a shared vocabulary, but taxonomies in fast-moving areas often get superseded.

Overall, this is a strong, rigorous, timely survey that will serve as a valuable reference and framing device for a meaningful slice of the game-AI subfield, without being a field-redefining contribution.

Rating:6/ 10
Significance 6Rigor 7.5Novelty 5Clarity 7.5

Generated Sep 16, 2026

Comparison History (0)

No comparisons yet.