Back to Rankings

Preventing Model Collapse: A Fisher-Rao Perspective on the Dynamics of Training with Synthetic Data

Matteo Marchi, João Pedro Silvestre, Bahman Gharesifard, Paulo Tabuada

Sep 16, 2026arXiv:2609.18878v1
cs.LGeess.SY
Share
Scorecard· 15/16
5.0/10 impact

A technically sound, timely, but narrow refinement of one prior paper's bound via a metric substitution, with no empirical validation and limited practical reach.

Abstract

Large Language Models (LLMs) are now routinely trained using synthetic data, since high-quality human data has been exhausted by the ever increasing needs of larger and larger models. However, recursive training on synthetic data frequently induces model collapse, a degenerative feedback loop where models progressively forget the true underlying data distribution. Training on a mixture of synthetic and fresh human data is a logical countermeasure and can prevent model collapse. However, it is an open question as to what is the exact minimum required ratio of human-to-synthetic data to maintain training stability. In this paper, we establish rigorous theoretical guarantees on the minimum rate of human data required to prevent model collapse. Although previous work established a formal lower bound for this ratio, such bound can be vacuous for very high dimensions, as the analysis relies on the usual Euclidean metric in R^n and is not adapted to the space of categorical probability distributions. Instead, in this paper we explicitly leverage the information-geometric structure of the probability simplex by analyzing the dynamics of the process under the Fisher-Rao metric. We derive quantitative contraction and invariance bounds that are stable and do not become trivial as the dimensions increase. Thus, we show that the effective required data ratio to prevent model collapse is different than previously implied.

AI Impact Assessments

(1 model)

Scientific Impact Assessment

Core Contribution

This paper refines the theoretical characterization of the minimum human-to-synthetic data ratio (μ) needed to prevent "model collapse" during recursive training of generative models. It builds directly on a specific prior framework (Gharesifard & Tabuada, CDC 2025), which modeled iterative training as a closed-loop stochastic process converging to a continuous-time ODE, and which derived a convergence bound using the Euclidean metric on the probability simplex. The paper's central insight is that Euclidean bounds become vacuous in high dimensions — because Euclidean distances between distributions collapse toward zero as n→∞ — and that the information-geometric (Fisher-Rao) structure of the simplex is the correct setting. Using a KL-divergence Lyapunov function, the authors derive contraction and invariance bounds that remain meaningful as dimension grows, and conclude that the required data ratio scales far more aggressively than previously implied (μ ~ n^{5/2+γ} rather than μ ~ n for O(1/n) error control).

Methodological Rigor

The mathematical development is sound and self-contained. The proof proceeds through carefully constructed lemmas: a uniform lower bound keeping trajectories off the simplex boundary (essential since Fisher-Rao is undefined on the boundary), a Lie-derivative computation for the KL Lyapunov function, Fisher-Rao-Lipschitz continuity of the softmax/temperature map, and two-sided bounds relating KL divergence to a weighted log-distance. Standard tools (Young's inequality, Hölder, Bregman-divergence representation, Hellinger-KL inequalities) are applied correctly. The disjoint-distribution example (Euclidean distance O(1/√n) vs. Fisher-Rao constant π/2) is a compelling and honest motivation. The main weakness in rigor is not internal but external: the analysis inherits a heavily stylized model (categorical one-hot outputs, a deterministic ODE surrogate for stochastic training, an abstracted perturbation term ε representing "training accuracy," and an assumption that the flow preserves the simplex). The connection between these idealizations and actual LLM training dynamics is asserted rather than demonstrated.

Potential Impact

The topic — data exhaustion and model collapse — is highly salient, and a rigorous handle on "how much human data is enough" is genuinely valuable. However, the practical reach of this specific result is limited. The takeaway (a superlinear n^{5/2} scaling requirement) is abstract and pessimistic, and it is not obvious how a practitioner would map n (the categorical vocabulary/outcome dimension) or the parameters δ, η, T to a real training pipeline. The contribution is best characterized as a methodological sharpening within a niche theoretical lineage (the Marchi/Tabuada "heat death" line of work), likely to be cited and extended by the small community working on formal dynamical-systems analyses of self-consuming loops, rather than by the broad LLM-training community.

Timeliness & Relevance

Very timely. Model collapse (Shumailov et al., Nature 2024; Alemohammad et al., ICLR 2024) and projected exhaustion of human text are active concerns. The information-geometric framing is a natural and currently under-exploited lens for this problem, giving the paper good positioning.

Strengths & Limitations

Strengths: (1) A clean conceptual argument for why Euclidean analysis is the wrong geometry, with a concrete illustrative example; (2) technically careful, complete proofs; (3) a clear, quantitative comparison (Proposition 1) that isolates exactly how the two metrics disagree in their dimensional scaling; (4) it meaningfully qualifies a recent prior result, showing its bound is over-optimistic.

Limitations: (1) No empirical validation whatsoever — even a synthetic numerical simulation of the ODE across dimensions would have strengthened the claims and demonstrated the practical gap between the two bounds; (2) the model's abstraction from real training limits actionable impact; (3) the scope is narrow — it is essentially a metric-substitution refinement of one prior paper, not a new framework; (4) the practical prescription (need dramatically more human data) is stated but its real-world plausibility is not examined.

Additional Observations

This is a competent, incremental theoretical contribution of the type typical at a strong control-systems venue (CDC). Its scientific value lies in correcting a scaling misconception rather than opening a new research direction. Reproducibility of the proofs is high for a specialist reader; the barrier to entry is low (pen-and-paper theory), which aids extension but also caps the resource-driven prestige. The interdisciplinary blend (information geometry + dynamical systems + ML) is appealing but confined to overlapping theory communities.

Rating:5/ 10
Significance 4.5Rigor 7Novelty 6Clarity 7.5

Generated Sep 17, 2026

Comparison History (0)

No comparisons yet.