Rubén Moreno-Bote
A clean theoretical clarification challenging the free energy variational approach to homeostasis, but narrow validation and an unaddressed scalability limitation cap its broad impact.
A common formalization of homeostasis is the free energy principle, a framework that defines a set of desired observation values, or critical states, that the agent should reach or remain close to. Under the free energy principle, an agent should act to maximize the probability of receiving the desired observations. Here we revisit the common approach of solving the problem of maximizing the log probability of the desired observations by maximizing a variational lower bound, the so-called negative free energy. We show that, instead, an approach directly maximizing that probability under the agent's policy is better suited to, and provides a better solution for, the original homeostatic control problem. This is done using hidden Markov model (HMM) control by allowing the policy to act over hidden states or noisy versions thereof while trying to maximize the probability of repeatedly having the desired observations. HMM control largely improves performance over the variational, or free energy, approach. We also show that the optimal policy is strictly deterministic, while the variational approach leads to a stochastic policy approximation. We finally provide a homeostatic reinterpretation of the maximum occupancy principle -a principle proposing that agents ought to maximize the occupancy of action-state path space -by defining homeostatic states as any states that do not immediately entail the termination or death of the agent.
The paper reformulates homeostatic control by challenging the dominant variational (free energy principle) approach used in active inference. Its central claim is that homeostasis—framed as maximizing the probability of obtaining desired observations over a finite horizon—should be solved *directly* via hidden Markov model (HMM) control rather than by maximizing a variational lower bound (negative free energy). Two analytical results anchor the paper: (1) direct HMM control yields a multiplicative backward message-passing recursion that exactly solves the objective, whereas the variational approach gives a generally suboptimal policy; and (2) the true optimal policy is strictly deterministic (because the objective is linear in each per-timestep policy on a simplex), while the variational/control-as-inference route produces a stochastic policy. The paper also offers a homeostatic reinterpretation of the author's own Maximum Occupancy Principle (MOP), treating homeostatic range boundaries as terminal (death) states.
The theoretical derivations are sound and cleanly presented. The deterministic-optimality argument follows directly from linearity over policy simplexes—correct, though somewhat elementary. The Section 3 analytical comparison of HMM control, control-as-inference, and variational inference for T=1 and T=2 is careful and yields a genuinely useful insight: control-as-inference exhibits "wishful thinking" by implicitly assuming control over state transitions the agent cannot actually steer. The empirical support, however, is thin: a single toy grid-world example (Fig. 2) with a hand-constructed reward landscape. There is no scaling study, no comparison across environments, and no engagement with the central practical reason variational methods exist—tractability of marginalization in large state spaces. The paper explicitly notes variational inference is used because exact marginalization is intractable, yet then advocates exact HMM control (which requires that same intractable marginalization) as "better" without resolving this tension. This is the most significant gap: the "advantage" is demonstrated only where exact solution is already feasible.
The work targets the active inference / free energy principle community in computational neuroscience, an active and sometimes contentious subfield. Its argument that the free energy variational bound is *suboptimal by construction* for the homeostatic objective, and its clean articulation of the deterministic-vs-stochastic distinction, could sharpen ongoing methodological debates. However, the impact is bounded by: (a) the niche scope, (b) heavy reliance on the author's own prior framework (MOP, refs 6, 7, 12 are self-citations), and (c) the unaddressed scalability limitation that undercuts practical adoption. It reads as a conceptual clarification and extension of the author's research program rather than a broadly enabling new tool.
Timely. Critiques and reformulations of the free energy principle are a live topic, and homeostatic reinforcement learning is an active area (refs 1–5). The paper's precise identification of where control-as-inference diverges from genuine control is relevant to a real, current confusion in the literature.
Strengths: Clear theoretical argument; a genuine and correct conceptual distinction (deterministic optimum vs. variational stochastic approximation); the "wishful thinking" critique of control-as-inference is illuminating and well-illustrated; connects several frameworks (free energy, MOP, empowerment, control-as-inference) coherently.
Limitations: Single toy experiment; scalability/tractability—the raison d'être of variational methods—is essentially sidestepped; the deterministic result, while correct, is mathematically mild; heavy self-referencing; no code released; the practical claim that HMM control is "better" is only validated where exact inference is trivially affordable. The paper somewhat overstates generality ("regardless of the specific problem considered") given the narrow demonstration.
This is a single-author theoretical paper with minimal resource requirements (analytical + laptop-scale simulation). Its refutation posture toward the variational free energy approach is real but partial—it qualifies rather than overturns, and it does so within a specific objective formulation. The connection between inference and control it exploits is acknowledged as previously known (Levine 2018, Todorov, Kappen), so the contribution is a targeted application/clarification to homeostasis rather than a new paradigm. The arXiv identifier (2026) suggests very recent/forward-dated posting.
Overall, a competent, conceptually clean theoretical contribution that will interest a specific subfield and may be cited in active inference debates, but whose impact is constrained by narrow validation and an unaddressed tractability limitation.
Generated Sep 9, 2026
A clean theoretical clarification challenging the free energy variational approach to homeostasis, but narrow validation and an unaddressed scalability limitation cap its broad impact.