Back to Rankings

Spread of Chronic Wasting Disease under Stochastic Environmental Conditions and its Control using Deep Reinforcement Learning

Wei Yin, Wesley J. Marrero, Kamal Jnawali, Lale Asik, Michael G. Tyshenko, Tamer Oraby

Sep 3, 2026arXiv:2609.05566v1
q-bio.PEmath.PR
Share
Scorecard· 16/16
5.5/10 impact

Technically solid dual contribution (rigorous reflected-SDE theory plus DRL control) on an important wildlife disease, but limited by non-spatial modeling, assumed parameters, weak RL evaluation, and niche applicability.

Abstract

Chronic wasting disease (CWD) is a fatal prion disease affecting deer, elk, moose, reindeer, muntjac, and other cervids. Because free-ranging cervid populations face environmental variability and randomness, deterministic models may miss important dynamics like stochastic fade-out. We develop a stochastic Susceptible-Infectious-Environmental model using differential equations with reflection to ensure the susceptible class remains non-negative. We examine how environmental variability influences cervid populations as CWD pressure and control measures increase. For the deterministic model, we derive the basic reproduction number as the sum of direct and environmental contributions, showing the endemic phase arises at R0=1. For the stochastic system, we establish local well-posedness, positivity, and the disease-free law. The top Lyapunov exponent for invasion remains unaffected by reflection. We evaluate CWD mitigation using a deep reinforcement learning agent trained with Proximal Policy Optimization in a hybrid action space, comparing hunting, decontamination, and combined strategies. In the deterministic case, hunting alone can control the disease but reduces the population by about 58%, while decontamination requires sustained effort. The combined policy more than doubles the cervid population and nearly eliminates infection and contamination. In the stochastic case, the policy contains the disease in about 80% of runs, with 10% experiencing large outbreaks; effectiveness decreases as noise increases. Across all scenarios, the agent consistently emphasizes environmental decontamination, the key control method.

AI Impact Assessments

(1 models)

Scientific Impact Assessment

Core Contribution. This paper couples two distinct threads: (1) a rigorous stochastic compartmental model of chronic wasting disease (CWD) with a susceptible-infectious-environmental (SIU) structure, formulated as reflected SDEs using Skorokhod (normal) reflection to keep the susceptible class nonnegative; and (2) a deep reinforcement learning (PPO-based, memory-augmented) controller over a hybrid discrete-continuous action space to optimize hunting versus environmental decontamination strategies. The analytical contributions include decomposing the basic reproduction number into direct and environmental components, establishing a forward transcritical bifurcation at R0=1, proving local well-posedness/positivity of the reflected system, and deriving a top Lyapunov exponent invasion criterion that is shown to be unaffected by reflection. The central empirical finding is that combined hunting+decontamination dominates either alone, and that environmental decontamination is the key control lever — a conclusion mechanistically explained by R_env ≫ R_dir.

Methodological Rigor. The mathematical analysis is the strongest part. The next-generation matrix derivation, bifurcation analysis, and the reflected-SDE well-posedness/Lyapunov exponent theory are carefully executed and internally consistent, drawing on standard but nontrivial tools (Tanaka–Lions–Sznitman theory, ergodic theorems, Itô calculus). The demonstration that reflection leaves the disease-free invariant measure and transverse linearization unchanged is a clean, non-obvious result. The RL side is weaker: it uses a standard PPO implementation (Stable-Baselines3), and several key results are drawn from "a single independent testing phase," with stochastic evaluation on only 30 realizations. There are no error bars/statistical significance tests on many claims, no baseline comparison against classical optimal control or MPC, and no ablation of the memory-based architecture versus a memoryless agent. Parameter values are heavily "assumed" (shedding rate, decontamination rate, noise intensities), which the authors candidly acknowledge.

Potential Impact. CWD is an escalating and economically consequential wildlife disease with no treatment, so a decision-support framework is genuinely useful. The most transferable contribution is the modeling template: coupling reflected stochastic dynamics with environmental-reservoir transmission and learned adaptive control, which the authors argue generalizes to other environmentally-transmitted diseases. However, practical impact is limited by acknowledged constraints: the model is non-spatial (CWD is fundamentally a hotspot/spatial problem), landscape-scale environmental decontamination of wild cervids is largely impractical (the authors themselves note this), and the RL policy's reliability degrades sharply as noise increases. Thus the "decontamination is dominant" recommendation is more relevant to farmed/captive settings than the free-ranging populations the model targets. The R_e(t) discussion — showing effective reproduction number can rise while disease is controlled under adaptive population-reduction — is a nice cautionary insight for practitioners.

Timeliness & Relevance. RL for epidemiological/epizootic control has grown rapidly since 2020, and applying it to wildlife disease with explicit environmental reservoirs is a reasonable and somewhat underexplored niche. The integration of stochasticity into CWD decision-making addresses a genuine gap the authors identify (prior CWD modeling relied on deterministic/regression approaches).

Strengths. (1) Unusual and genuine breadth — rigorous SDE theory plus modern DRL in one paper; (2) the reflection formulation and its invariance proofs are technically elegant and correct; (3) the mechanistic linkage between the analytical R0 decomposition and the RL agent's learned preference for decontamination gives the empirical results interpretive grounding; (4) careful design constraint capping the hunting rate at the analytically-derived extinction threshold, preventing physically absurd policies.

Limitations & Gaps. (1) Non-spatial well-mixed assumption is a serious simplification for a spatially clustered disease; (2) heavy reliance on assumed parameters undermines quantitative claims; (3) weak RL evaluation protocol (single testing phase, 30 runs, no competing control baselines, no statistical rigor); (4) no code/data release mentioned, limiting reproducibility despite fairly complete hyperparameter reporting; (5) the practical infeasibility of the recommended decontamination lever in wild settings creates tension between the model's target population and its conclusions; (6) claims of generalizability to "a broader class of infectious diseases" are asserted rather than demonstrated.

Additional Observations. The theoretical results (bifurcation, Lyapunov exponent, reflection invariance) are self-contained and citable independently of the RL component; a reader could reuse the modeling machinery. The resource barrier is low (single GPU, modest simulation), which aids extension. Surprisingness is modest — stochastic fade-out, Itô-corrected mean below deterministic equilibrium, and environmental-route dominance are consistent with prior theory and expectation. The paper does mildly qualify the standard use of R_e(t) as a control metric, which has some refutation-adjacent value but is framed as nuance rather than a challenge to an established belief.

Overall, this is a competent, technically solid interdisciplinary paper whose mathematical analysis exceeds its empirical control evaluation. Its likely impact is as a useful methodological template cited by a modest slice of the mathematical-ecology/disease-control community, rather than a field-changing contribution.

Rating:5.5/ 10
Significance 5Rigor 6Novelty 6.5Clarity 6.5

Generated Sep 9, 2026

Comparison History (0)

No comparisons yet.