Lefebvre Renard Clément, Lébé Vincent, Da Silva Ribeiro Pereira Ricardo, Sundell Johan, Jaoul Arnaud Saiah Kenza, Mijatovíc Nenad
Timely, well-synthesized industry roadmap for railway AI trustworthiness, but with no original technical contribution, experiments, or resolution of the core SIL-vs-AI-accuracy problem it raises.
Artificial Intelligence (AI) is currently only applied to non-safety critical applications due to the strict standards and regulations for railway industries. We propose to review the three main fields necessary to increase trust in data science and AI algorithms and reach compliance: robustness, Operational Design Domain (ODD), and explainability. Robustness is the ability of an AI system to maintain its level of performance under any circumstances (ISO24029). ODDs allow the explicit definition of operating conditions under which a system is intended to operate, according to the recently published DIN DKE SPEC 99004. Explainability is the property of an AI system to express important factors influencing the AI system results in a way that humans can understand. Those 3 domains of research are already well investigated by nonrailway actors, with algorithms and methods ready to use for railway applications. A system view is necessary to ensure all trustworthy requirements interact continuously in a safe MLOps environment thereby fostering acceptance from regulators, operators and the public. Beyond safeguarding safety-critical applications, we aim to show that fostering deep trust in AI, as now required by regulatory frameworks worldwide, will unlock its full potential and transform the pace of adoption across mission-critical domains.
This is a position/review paper authored by industry practitioners at Alstom (a major railway rolling-stock manufacturer) that argues AI can be brought into railway safety-critical applications if three "levers of trust" are systematically addressed: robustness, Operational Design Domain (ODD), and explainability. The paper's central claim is that these three domains are already well-developed by non-railway actors (autonomous driving, general ML research) and can be transferred into the railway domain, integrated within a "safe MLOps" system view. The core contribution is not a new algorithm, theorem, or dataset — it is a *framing and synthesis* that maps existing regulatory standards (EU AI Act, ISO/IEC TR5469, ISO24029, DIN DKE SPEC 99004, EN5012x/SIL frameworks) onto three actionable technical research areas, and proposes a system-level integration diagram (Figure 4) tying these into a classical ML workflow.
The problem it solves is essentially organizational/conceptual: it bridges the gap between fragmented regulatory obligations and concrete engineering practices for a specific vertical (rail) where AI is currently barred from safety-critical use.
As a review/position paper, there are no original experiments, proofs, or empirical validations. The paper cites appropriate and current standards and foundational ML robustness/xAI literature (Szegedy, Goodfellow, Madry, Carlini-Wagner, Cohen randomized smoothing, Lipschitz networks, SHAP, GradCAM). The one formal element — the definition of Lipschitz continuity and the certified-robustness inequality — is standard textbook material, correctly stated but not novel.
The argument is coherent but the treatment of each of the three pillars remains at a survey depth. There is no case study demonstrating the proposed system approach on an actual railway safety function, no quantitative evaluation, and the crucial gap it itself identifies (AI accuracy of 80–99% vs. SIL4's 10⁻⁹ hazard rate) is flagged but not resolved. This is an honest but significant limitation: the paper motivates the problem sharply but offers no evidence that the proposed combination actually achieves SIL-level guarantees.
The impact is primarily industrial/practical rather than scientific-generative. For the railway domain specifically, the paper serves as a useful orientation document that aligns terminology across standards and points practitioners toward mature tooling (DEEL-Lip, Xplique, Confiance.AI outputs). It could be cited by railway AI teams, standardization committees, and safety engineers as a reference framing. However, none of the underlying techniques are new, so the paper is unlikely to influence the core ML research community. Its influence is bounded to a fairly narrow vertical (railway AI safety), with possible spillover to adjacent regulated-transport domains (metro, signaling).
The affiliation with Alstom lends credibility and signals genuine industrial intent, which increases the chance the framing gets adopted in practice, but the paper's function is more advocacy/roadmap than technical breakthrough.
This is the paper's strongest dimension. It sits precisely at the confluence of the EU AI Act (2024), newly published DIN DKE SPEC 99004 (2025), ISO/IEC TR5469 (2024), and rising pressure to deploy AI in predictive maintenance and autonomous train operations. The regulatory-compliance angle is genuinely timely and addresses a real bottleneck: the field has active regulation but immature guidance for how to comply in rail. The synthesis of very recent standards is valuable and current.
This is a competent, timely industry position/review paper that will serve as a useful orientation and advocacy document within the railway AI safety community. Its scientific novelty is low — it recombines known techniques and standards — but its practical relevance and timing are high for its target vertical. It is unlikely to be a widely-cited foundational reference outside rail, but may become a go-to framing document for railway AI trustworthiness discussions and standardization efforts.
Generated Sep 17, 2026
Timely, well-synthesized industry roadmap for railway AI trustworthiness, but with no original technical contribution, experiments, or resolution of the core SIL-vs-AI-accuracy problem it raises.