Back to Rankings

Building Trust in Artificial Intelligence: A Necessity for Railway Applications

Lefebvre Renard Clément, Lébé Vincent, Da Silva Ribeiro Pereira Ricardo, Sundell Johan, Jaoul Arnaud Saiah Kenza, Mijatovíc Nenad

Sep 16, 2026arXiv:2609.18278v1
cs.AI
Share
Scorecard· 14/16
3.5/10 impact

Timely, well-synthesized industry roadmap for railway AI trustworthiness, but with no original technical contribution, experiments, or resolution of the core SIL-vs-AI-accuracy problem it raises.

Abstract

Artificial Intelligence (AI) is currently only applied to non-safety critical applications due to the strict standards and regulations for railway industries. We propose to review the three main fields necessary to increase trust in data science and AI algorithms and reach compliance: robustness, Operational Design Domain (ODD), and explainability. Robustness is the ability of an AI system to maintain its level of performance under any circumstances (ISO24029). ODDs allow the explicit definition of operating conditions under which a system is intended to operate, according to the recently published DIN DKE SPEC 99004. Explainability is the property of an AI system to express important factors influencing the AI system results in a way that humans can understand. Those 3 domains of research are already well investigated by nonrailway actors, with algorithms and methods ready to use for railway applications. A system view is necessary to ensure all trustworthy requirements interact continuously in a safe MLOps environment thereby fostering acceptance from regulators, operators and the public. Beyond safeguarding safety-critical applications, we aim to show that fostering deep trust in AI, as now required by regulatory frameworks worldwide, will unlock its full potential and transform the pace of adoption across mission-critical domains.

AI Impact Assessments

(1 model)

Scientific Impact Assessment

1. Core Contribution

This is a position/review paper authored by industry practitioners at Alstom (a major railway rolling-stock manufacturer) that argues AI can be brought into railway safety-critical applications if three "levers of trust" are systematically addressed: robustness, Operational Design Domain (ODD), and explainability. The paper's central claim is that these three domains are already well-developed by non-railway actors (autonomous driving, general ML research) and can be transferred into the railway domain, integrated within a "safe MLOps" system view. The core contribution is not a new algorithm, theorem, or dataset — it is a *framing and synthesis* that maps existing regulatory standards (EU AI Act, ISO/IEC TR5469, ISO24029, DIN DKE SPEC 99004, EN5012x/SIL frameworks) onto three actionable technical research areas, and proposes a system-level integration diagram (Figure 4) tying these into a classical ML workflow.

The problem it solves is essentially organizational/conceptual: it bridges the gap between fragmented regulatory obligations and concrete engineering practices for a specific vertical (rail) where AI is currently barred from safety-critical use.

2. Methodological Rigor

As a review/position paper, there are no original experiments, proofs, or empirical validations. The paper cites appropriate and current standards and foundational ML robustness/xAI literature (Szegedy, Goodfellow, Madry, Carlini-Wagner, Cohen randomized smoothing, Lipschitz networks, SHAP, GradCAM). The one formal element — the definition of Lipschitz continuity and the certified-robustness inequality — is standard textbook material, correctly stated but not novel.

The argument is coherent but the treatment of each of the three pillars remains at a survey depth. There is no case study demonstrating the proposed system approach on an actual railway safety function, no quantitative evaluation, and the crucial gap it itself identifies (AI accuracy of 80–99% vs. SIL4's 10⁻⁹ hazard rate) is flagged but not resolved. This is an honest but significant limitation: the paper motivates the problem sharply but offers no evidence that the proposed combination actually achieves SIL-level guarantees.

3. Potential Impact

The impact is primarily industrial/practical rather than scientific-generative. For the railway domain specifically, the paper serves as a useful orientation document that aligns terminology across standards and points practitioners toward mature tooling (DEEL-Lip, Xplique, Confiance.AI outputs). It could be cited by railway AI teams, standardization committees, and safety engineers as a reference framing. However, none of the underlying techniques are new, so the paper is unlikely to influence the core ML research community. Its influence is bounded to a fairly narrow vertical (railway AI safety), with possible spillover to adjacent regulated-transport domains (metro, signaling).

The affiliation with Alstom lends credibility and signals genuine industrial intent, which increases the chance the framing gets adopted in practice, but the paper's function is more advocacy/roadmap than technical breakthrough.

4. Timeliness & Relevance

This is the paper's strongest dimension. It sits precisely at the confluence of the EU AI Act (2024), newly published DIN DKE SPEC 99004 (2025), ISO/IEC TR5469 (2024), and rising pressure to deploy AI in predictive maintenance and autonomous train operations. The regulatory-compliance angle is genuinely timely and addresses a real bottleneck: the field has active regulation but immature guidance for how to comply in rail. The synthesis of very recent standards is valuable and current.

5. Strengths & Limitations

Strengths:

  • Excellent, up-to-date synthesis of a fast-moving regulatory landscape specific to rail.
  • Clear practitioner perspective with concrete pointers to industrial tooling and existing railway ML deployments (track circuits, point machines, signal detection).
  • Honest acknowledgment of gaps (post-hoc xAI unreliability under correlated features, SIL vs. AI accuracy chasm, immaturity of causality).
  • The system-view integration (Figure 4) is a modest but useful conceptual contribution.
  • Limitations:

  • No original technical contribution, no experiments, no case study, no quantitative validation.
  • Does not resolve the central tension it raises (how AI reaches SIL4 rates).
  • Coverage is broad but shallow; each pillar is treated at survey level without depth.
  • The claim that methods are "ready to use for railway applications" is asserted rather than demonstrated.
  • Limited generalizability of any concrete finding, since there are no findings per se — it is a roadmap.
  • Explicitly not exhaustive (omits cybersecurity, ethics, data management by its own admission).
  • Overall

    This is a competent, timely industry position/review paper that will serve as a useful orientation and advocacy document within the railway AI safety community. Its scientific novelty is low — it recombines known techniques and standards — but its practical relevance and timing are high for its target vertical. It is unlikely to be a widely-cited foundational reference outside rail, but may become a go-to framing document for railway AI trustworthiness discussions and standardization efforts.

    Rating:3.5/ 10
    Significance 3.5Rigor 4Novelty 3Clarity 7

    Generated Sep 17, 2026

    Comparison History (0)

    No comparisons yet.