Back to Rankings

From Tsallis to KL: Convergence and Error Estimates for Tsallis-Regularized Optimal Transport

Takeshi Suguro, Toshiaki Yachimura

Sep 6, 2026arXiv:2609.06432v1
cs.ITmath.FAmath.OCmath.PR
Share
Scorecard· 15/16
4.0/10 impact

Rigorous, clean theoretical result quantifying the Tsallis-to-KL limit, but confined to a narrow subfield with no algorithmic or empirical contribution and expected qualitative conclusions.

Abstract

We study the Tsallis-to-Kullback--Leibler (KL) limit for entropy-regularized optimal transport with nonnegative bounded continuous costs. Fixing the regularization parameter ε>0\varepsilon > 0, we first derive an exact variational reformulation of Tsallis-regularized optimal transport in terms of the Tsallis information projection onto the set of couplings. The formula isolates an explicit correction term and thereby explains why, unlike in the KL case, the regularized transport problem and the corresponding information projection problem do not coincide exactly. We also establish existence and uniqueness for the Tsallis information projection. We then prove, with respect to the narrow topology, the ΓΓ-convergence of the Tsallis-regularized functionals to the KL-regularized functional as q1q\downarrow1, together with narrow convergence of their unique minimizers. Finally, we obtain explicit error estimates of order O(q1)O(q-1) for both the regularized optimal transport values and the associated information projection values. These results quantify the passage from Tsallis regularization to the classical KL setting and clarify the relation between entropic regularization and information projection for 1<q21 < q \leq 2.

AI Impact Assessments

(1 models)

Scientific Impact Assessment

Core Contribution

This paper provides a rigorous analytical treatment of the Tsallis-to-KL limit for entropy-regularized optimal transport (OT). Three contributions stand out. First, the authors derive an *exact variational reformulation* (Proposition 1.1) expressing Tsallis-regularized OT in terms of a Tsallis information projection onto the coupling set, isolating an explicit correction term. This cleanly explains *why* — unlike the special KL case — the regularized transport problem and the information projection problem do not coincide for general Tsallis divergences. Second, they establish existence and uniqueness for the Tsallis information projection. Third, they prove Γ-convergence (in the narrow topology) of Tsallis-regularized functionals to the KL functional as q↓1, with narrow convergence of minimizers, and — most valuably — explicit O(q−1) error estimates for both regularized OT values and information projection values, with constants depending explicitly on ‖c‖∞ and ε.

The work fits a niche: it complements the authors' prior study of the ε↓0 limit [22] by instead fixing ε and taking q↓1. The correction-term formula is the conceptual centerpiece, giving structural insight into the divergence between f-divergence regularization and f-projection.

Methodological Rigor

The mathematical development is sound and self-contained. Proofs are careful: the Γ-convergence argument correctly handles liminf/limsup inequalities with an explicit recovery-sequence construction (truncation + marginal correction) and a diagonal-sequence argument to combine n→∞ and q_k↓1. Lower semicontinuity, narrow compactness of the coupling set, and strict convexity of the generator ϕ_q are appropriately invoked for existence/uniqueness. The error estimates rely on clean pointwise inequalities (mean value theorem bounds on t^q−t, log₂₋q(x)−log x) and uniform L∞ bounds on optimal couplings via Schrödinger-potential bounds imported from Di Marino–Gerolin [11]. The constructions are convincing, and the constants are explicit rather than abstract, which is a strength. The results are restricted to 1<q≤2 and nonnegative bounded continuous costs — reasonable and clearly stated assumptions. No gaps are apparent in the derivations.

Potential Impact

The impact is likely to be modest and confined to the theoretical OT / information-geometry community. Entropy-regularized OT is a heavily used tool in ML, imaging, and graphics, and generalized (Tsallis / f-divergence) regularizations have appeared in the ML literature (Muzellec et al. AAAI 2017, Bao–Sakaue, Terjék–González-Sánchez AISTATS 2022). Quantifying the Tsallis-to-KL passage with explicit rates provides a principled justification for using Tsallis regularization as a tunable interpolation and for understanding numerical behavior as q→1. However, the paper contains no algorithms, no experiments, and no direct computational contribution. Its influence will be primarily as a cited theoretical reference for those working on generalized regularizers, sparsity via q-entropy, or the projection-vs-regularization distinction. The O(q−1) rate is a useful concrete takeaway but is not surprising — first-order convergence in the deformation parameter is the expected outcome.

Timeliness & Relevance

Generalized-entropy OT is an active but specialized topic. The paper addresses a genuine conceptual question (why regularized OT and information projection diverge for non-KL divergences) that had been noted but not sharply quantified. It is timely within its subfield but does not address a widely-felt bottleneck; the machine-learning demand for Tsallis OT specifically is limited compared to KL/Sinkhorn.

Strengths & Limitations

Strengths:

  • Clean, exact variational identity (Prop. 1.1) with clear conceptual payoff.
  • Explicit, constant-tracked error bounds rather than mere asymptotic statements.
  • Rigorous Γ-convergence with minimizer convergence, carefully proven.
  • Well-organized, precise exposition typical of a strong analysis paper.
  • Limitations:

  • Purely theoretical; no numerical validation of the O(q−1) rate or of the practical size of the (potentially large, exponential in ‖c‖∞/ε) constants.
  • Constants blow up exponentially in ‖c‖∞/ε, limiting practical usefulness for small ε — a caveat the paper does not discuss in application terms.
  • Restricted to 1
  • The results, while solid, are somewhat expected extensions of known machinery (Γ-convergence, f-divergence projection theory, Schrödinger potential bounds), leaning heavily on imported results [11].
  • Narrow audience; limited cross-disciplinary reach beyond OT/information geometry.
  • Other Observations

    This is a "complete-in-itself" analytical result rather than a reusable building block or framework. It corroborates the intuitive expectation that Tsallis→KL smoothly, adding quantitative precision. Reproducibility in the theoretical sense is high — all proofs are given in full. There is no dataset, code, or empirical component. The paper does not contest any prior claim; it clarifies and quantifies a known structural distinction. Its foundational value is limited: others may cite the correction-term formula and rate, but few will build machinery directly on top of it.

    Overall, a technically clean, well-executed piece of specialized mathematical analysis with clear but bounded significance.

    Rating:4/ 10
    Significance 3.5Rigor 8Novelty 5Clarity 8

    Generated Sep 9, 2026

    Comparison History (0)

    No comparisons yet.