Machine Learning Integration in Nuclear Safety Analysis Reports: A Regulatory Framework for Nuclear Safety Applications
The use of machine learning (ML) in nuclear power plant safety analysis is becoming an important development in the assessment of risk, the detection of abnormal conditions, and the improvement of operational safety. Rather than replacing established safety analysis methods, ML is increasingly being examined as a supporting tool that can strengthen data interpretation, pattern recognition, and decision-making within the nuclear safety framework. This article develops a structured framework for incorporating ML-based methods into Safety Analysis Reports (SARs) for nuclear power plants, with particular attention to International Atomic Energy Agency (IAEA) safety standards and recent academic research. Drawing on recent technical discussions, relevant safety guidance, and peer-reviewed studies, the methodological foundations, regulatory expectations, and practical applications of ML in nuclear safety are examined. Particular attention is given to challenges that are especially significant in nuclear applications, including limited data availability, model transparency, and the verification and validation of ML-based tools. A structured approach is proposed for the inclusion of physics-informed machine learning, probabilistic risk assessment methods, and human factors considerations within SAR documentation. In addition, emerging applications in piping integrity assessment, fault diagnosis, and reactor performance monitoring are reviewed, and future research needs are identified in relation to the IAEA's developing safety framework for artificial intelligence (AI) in nuclear installations.
The International Atomic Energy Agency (IAEA) has recognized that the adoption of AI and ML in nuclear applications introduces new challenges for safety assurance and regulatory oversight. In particular, AI-related issues such as data dependency, uncertainty, algorithmic transparency, system integration, and human–machine interaction require further consideration. The planned revisions of IAEA Safety Standards Series No. SSG-39 on instrumentation and control systems and No. SSG-51 on human factors engineering are expected to address emerging challenges associated with advanced digital technologies, including AI and ML, in nuclear safety-related applications (IAEA, 2019, Cha, 2024).
The established regulatory principles of specific verification, conservative engineering assumptions, and traceable design remain essential, but they should be extended for ML-based systems. Unlike conventional software, ML performance is shaped by training data, model architecture, operating context, and potential model drift during service. These characteristics require Safety Analysis Reports (SARs) to document data provenance and governance, model development, validation under uncertainty, explainability commensurate with safety significance, and the procedures by which algorithms will be monitored, maintained, and updated throughout their operational lifecycle (Park et al., 2023; Park & Koo, 2026).
These requirements are especially relevant for small modular reactors (SMRs), which are intended for factory manufacture, deployment at multiple sites, and operation with smaller staffing levels. Such features change the safety and security context and increase reliance on digital instrumentation and control (I&C), including advanced human-system interfaces (HSIs). The way operators perceive plant status, manage cognitive workload, and respond to abnormal conditions may therefore differ from conventional nuclear power plant practice (Ameyaw et al., 2025). As a result, human reliability analysis (HRA) should incorporate emerging methods, including biometric monitoring and Bayesian belief networks, to assess the determinants of performance under stress. This approach is consistent with the human factors engineering principles outlined in IAEA SSG-51 (IAEA, 2019; Ameyaw et al., 2025). The graded approach already used for safety classification, including in the NuScale design certification process (Alamsyah et al., 2023), should also be applied to ML-based safety functions so that algorithms with higher safety significance receive proportionately stronger regulatory scrutiny without unnecessarily restricting technical innovation (Onwuatuegwu and Nwagu, 2021; Alamsyah et al., 2023).
Verification and validation (V&V) of ML components is among the most demanding issues in applying AI to nuclear safety systems. Conventional software V&V is not sufficient for adaptive or self-learning algorithms because the behaviour of an ML model depends on the data to which it is exposed and on the environment in which it operates. Therefore more robust assurance strategies are required. Promising approaches combine physics-informed neural networks (PINNs), surrogate modelling, and knowledge distillation from high-fidelity simulators (Chae et al., 2024; Jinia et al., 2024; Lye et al., 2025). These methods can improve model reliability while maintaining the level of confidence expected in safety-critical applications. Digital twins provide a further mechanism for continuous comparison between model predictions and plant data and support the development of a living probabilistic safety assessment (PSA) that is updated as operating experience accumulates (Mendoza et al., 2026; Zhang et al., 2024).
Cybersecurity issues should also be considered when including ML in SAR. Modern digital I&C systems, particularly those using wireless communication or cloud-based analytics, expand the potential attack surface relative to traditional architectures (Edhi Harianto et al., 2023; Lopez Pulgarin et al., 2024). Accordingly, Integrated Management of Safety and Security (IMSS) should be embedded within SAR documentation (Ylönen & Björkman, 2023). The SAR should show how safety and security requirements are balanced in design decisions, how fail-safe modes will maintain the plant in a safe state if an ML component becomes unavailable or behaves unexpectedly, and how defense-in-depth separates safety- or security-important channels from non-essential digital networks. Such separation reduces the likelihood that cyber compromise in one part of the architecture will propagate to functions that are important to safe operation (Ylönen & Björkman, 2023; Azam et al., 2026).
Regulatory harmonization is another important enabler for the international deployment of SMRs. Although considerable effort has been made to improve consistency, licensing terminology and evidentiary expectations still differ among jurisdictions and continue to impede approval across multiple regulatory systems (Cusmanri & Lim, 2024; Sam et al., 2023). A harmonized SAR template for ML-enabled safety and security functions would help by linking the evidence required for regulatory decisions to the safety principles established by the IAEA. Such a framework should maintain traceability from the top-level safety or security objective to each significant ML design decision. It should also define methods for uncertainty quantification, performance monitoring, and retraining criteria based on model-drift detection so that ML-enabled systems remain dependable over their service life (Claghorn et al., 2025).
So, updating SARs for AI and ML is not merely a documentation exercise; it is a necessary step in demonstrating that advanced nuclear technologies can be licensed, operated, and reviewed with appropriate confidence. A mature SAR should give regulators transparent evidence that ML-based functions satisfy applicable safety and security expectations and should support more consistent licensing decisions across national boundaries. This requires alignment with IAEA safety standards from the earliest design stages, disciplined use of physics-aware AI, and explicit treatment of cybersecurity and human factors throughout the system lifecycle. When these elements are brought together within an integrated safety framework, the nuclear sector can benefit from AI while preserving its established commitment to conservative and demonstrable nuclear safety. Fig. 1 presents a conceptual overview of AI/ML integration in nuclear facilities, including applications, safety benefits, cybersecurity, and regulatory considerations.
Fig. 1: Conceptual overview of AI/ML integration in nuclear facilities, including safety applications, I&C, predictive maintenance, cybersecurity, and regulatory challenges.
Foundational Concepts and Regulatory Context
The IAEA safety standards provide the principal international framework for demonstrating the safety of nuclear facilities and activities. Developed over several decades through operational experience and lessons learned from major accidents, including Three Mile Island, 1979; Chornobyl, 1986; Fukushima Daiichi, 2011), the framework is organized as a hierarchy of principles, requirements, and guidance (IAEA, 2006, 2016a, 2016b; Burke, 2022; Duliba, 2023). At the highest level, the Fundamental Safety Principles (SF-1) define the objectives and core principles of nuclear safety. These principles are implemented through Safety Requirements, including SSR-2/1 (Rev. 1) for nuclear power plant design and SSR-2/2 (Rev. 1) for commissioning and operation. Safety Guides provide further regulatory recommendations and practical means of meeting those requirements. SSG-39, which addresses instrumentation and control systems for nuclear power plants, and SSG-51, which addresses human factors engineering in nuclear facility design, are particularly relevant to the integration of ML technologies (IAEA, 2019, 2006, 2016a, 2016b; Rakitin & Chebyshov, 2022; Wahlström, 2015). The IAEA has begun developing broader guidance on the deployment and safety demonstration of AI-enabled technologies in nuclear power applications. This work reflects the fact that ML systems can be sensitive to distributional change, may offer limited interpretability, and can behave unpredictably when exposed to inputs outside the training distribution. These properties were not fully anticipated by frameworks developed for deterministic digital technologies (Meschini, 2022; Park et al., 2023).
ML, a subfield of AI, comprises methods that improve task performance by learning from data rather than relying exclusively on explicitly programmed rules for every operating condition (Jinia et al., 2024). In nuclear safety applications, ML is increasingly being explored for reliability analysis, severe-accident modelling, real-time fault and anomaly detection, digital twins, prognostics and health management, and thermal-hydraulic optimization (Meschini, 2022; Nguyen & Diab, 2023; Sallehhudin & Diab, 2021). The key distinction from conventional safety-sensitive software is that ML models infer internal decision boundaries from training data. When operating conditions or data distributions differ from those represented during training, model behaviour may become uncertain and difficult to trace to predefined requirements. This weakens the direct applicability of traditional structural software testing and requirements-traceability methods (Borg et al., 2023; Park et al., 2023). Traditional software V&V, requirements traceability and structural testing remain necessary, but they may not be sufficient on their own to provide assurance for data-dependent ML behaviour.
This distinction challenges established V&V practices in nuclear control systems. Methods designed for deterministic software cannot, by themselves, provide sufficient assurance for data-driven models. New strategies are therefore needed, with greater emphasis on data governance, robustness testing, uncertainty quantification, and continuous performance monitoring during operation. These measures are required if ML-based safety functions are to remain reliable under changing plant conditions (Park et al., 2023).
The Safety Analysis Report (SAR) is a key licensing document for a nuclear facility. It demonstrates that the design and operating arrangements satisfy applicable safety requirements, provides the technical basis for operating limits and conditions, and presents the integrated safety case supporting emergency preparedness and severe-accident management (IAEA, 2016a, 2016b, 2021; Martin, 2012). Within the IAEA framework, SSR-2/2 (Rev. 1) establishes requirements for operational safety, while complementary Safety Guides address maintenance, in-service inspection, surveillance testing, and emergency operating procedures (Rakitin & Chebyshov, 2022; Wahlström, 2015). When ML-based systems are introduced into nuclear I&C architectures, several SAR chapters are affected. The report should specify ML lifecycle requirements, including data sources, data quality assurance, model training, validation, robustness against adverse or unusual inputs, in-service performance monitoring, and retraining controls where appropriate (Park et al., 2023). It should also classify ML functions within the I&C hierarchy, evaluate human-system interactions in accordance with SSG-51, and assess the influence of ML-based decision support on deterministic safety analyses and probabilistic safety assessments (PSA). These additions are necessary to keep the safety case consistent with evolving IAEA expectations (Gisquet et al., 2021; Martorell et al., 2018; Rakitin & Chebyshov, 2022).
Fig. 2: ML Development Lifecycle for Safety Analysis Reports (SAR).
Machine Learning Applications in Nuclear Safety Analysis
The use of ML in nuclear safety analysis has expanded the analytical tools available for risk assessment, system integrity evaluation, and operational monitoring. Probabilistic risk assessment (PRA) remains a foundation of nuclear safety analysis because it evaluates both the likelihood and consequences of accident scenarios. ML can strengthen PRA by addressing persistent limitations, including rare data, complex model interpretation, and the need to analyse dynamic plant states. One important development is physics-enhanced or physics-based machine learning (PEML) (Xu et al., 2023; Zhu et al., 2023). By embedding conservation laws, constitutive relationships, and governing equations in the learning process, PEML constrains predictions to remain physically meaningful while still using available data. This reduces dependence on very large datasets, improves estimation in sparse or variable operating environments, and mitigates the black-box character of purely data-driven models. For safety-sensitive nuclear applications, these qualities are valuable because they combine predictive accuracy with a stronger physical basis (Xu et al., 2023; Zhu et al., 2023). Reliability analysis, a core element of PRA, also benefits from ML through improved estimation of component failure probability and degradation behaviour in large, complex datasets.
Recent studies indicate that survival analysis combined with ML can produce more informative time-dependent failure models for pipeline systems (Xiao et al., 2024). Such models describe how failure risk evolves and can capture degradation patterns that are difficult to represent with conventional statistical methods. Neural networks have likewise been applied to predict failure times of atmospheric storage tanks exposed to external fires, where multiple stressors interact and influence structural integrity (Tamascelli, Scarponi, et al., 2024). In accident modelling, ML-based surrogate models for thermal-hydraulic and severe-accident codes can reproduce event sequences far faster than high-fidelity simulations while retaining the capacity for uncertainty analysis (Hossny et al., 2023; Song & Kim, 2024). This computational efficiency allows a wider range of accident progressions to be examined during both routine safety studies and emergency decision support.
Piping integrity assessment is another area in which ML has substantial relevance for nuclear power plant safety. ML techniques have therefore been investigated for piping inspection and maintenance, particularly in three applications: corrosion-defect detection using convolutional neural networks (CNNs) trained on ultrasonic testing data, prediction of wall-thinning rates under varying operating conditions, and estimation of remaining useful life (RUL) for risk-informed inspection planning (Akhlaghi et al., 2023; Allah et al., 2025). Although these applications have produced encouraging results, deployment is limited by the scarcity and cost of representative inspection data. Nuclear plants operate within narrow safety and security margins, so large and diverse degradation datasets are difficult to obtain. To address this limitation, synthetic data generated with Generative Adversarial Networks (GANs) can be used to balance datasets and improve model training (Mazumder et al., 2023). Transfer learning provides another route by transferring knowledge from related non-destructive evaluation tasks to the target application, thereby reducing the amount of labelled data required while maintaining predictive performance (Sandhu et al., 2024).
Deep learning has also improved fault and anomaly detection. Autoencoders and recurrent neural networks (RNNs), trained on data from normal plant operation, learn baseline system behaviour and can identify small deviations before conventional alarm limits are exceeded (Arshad et al., 2025; Qian & Liu, 2022). Because these models evaluate multivariate time-series data from plant sensors and instrumentation, they can detect fault patterns that would be missed when parameters are monitored individually (Zhou et al., 2023; Zubair & Akram, 2023). Their diagnostic reliability can be strengthened by incorporating physics-based constraints so that identified anomalies remain consistent with known failure mechanisms, thereby reducing false alarms (Wu et al., 2024). Digital twins extend this capability by combining real-time sensor data with simulation models that track the plant state during operation. This allows PRA to be updated as conditions change (Mengyan et al., 2024) and supports predictive maintenance and operator decision-making during abnormal or emergency conditions (Jeon et al., 2024).
Despite these advances, several issues should be resolved before ML can be relied upon in nuclear safety applications. First, V&V methods should demonstrate reliable performance despite sensor drift, missing data, and operating conditions that differ from the training domain, including unforeseen plant states (Meschini, 2022). Second, human-machine interfaces should present warnings and recommendations in forms that are clear, procedurally consistent, and actionable, so that operators can respond with confidence during normal and abnormal conditions (Sethu et al., 2023). Third, uncertainty should be communicated rather than hidden. ML systems should provide confidence information or other uncertainty measures that help operators judge the reliability of predictions, reduce unnecessary alarms, avoid alarm fatigue, and preserve adequate safety margins (Tamascelli, Campari, et al., 2024). Future work is therefore likely to focus on hybrid models that combine data-driven flexibility with physics-based knowledge, enabling more accurate, credible, and proactive approaches to risk management.
Safety Considerations for ML Integration
Incorporating ML into nuclear safety systems requires a reassessment of traditional engineering assurance. ML models differ from conventional software because they are probabilistic, dependent on training data, and often only partially explainable. Safety-sensitive software normally follows specified logic that can be checked against formal requirements and expected test cases. ML decisions, by contrast, are shaped by the quality, coverage, and representativeness of the training dataset. A central concern is the response of a model to out-of-distribution (OOD) inputs. In such cases, predictions may become unreliable or difficult to interpret. This is particularly important in nuclear power plants because data cannot be collected for every possible condition, especially rare events and severe accidents. Even high-quality simulations cannot capture all circumstances that may arise in operation. The resulting uncertainty is a major challenge in constructing a credible safety case for ML-based nuclear applications (Askarpour, 2021).
Recent technical discussions, including those associated with IAEA activities, have therefore emphasized three priorities for ML V&V. The first is to extend conventional software V&V across the entire ML lifecycle. Verification should cover not only source code, but also training data, data preparation, model architecture, training process, and deployment environment, since each can affect performance (Mattioli, 2022). The second priority is rigorous data quality and governance. In safety-related applications, poor data can lead directly to unreliable predictions and unsafe decisions; therefore, data provenance, quality, completeness, and representativeness should be documented at each lifecycle stage (Meschini, 2022). The third priority is evaluation under abnormal and unexpected conditions, including degraded operation, rare events, and cases not represented during training. Stress testing, scenario-based testing, and adversarial evaluation should be used to identify failure modes and to ensure that models fail in controlled rather than catastrophic ways (Khadka et al., 2024). These priorities are closely tied to data availability. Serious safety events are rare in nuclear power plants, which is desirable operationally but problematic for ML training because limited real-world data are available for early fault signatures (Alsafadi, 2024). Sensor noise, missing values, calibration drift, and design differences among plants further limit generalization across facilities (Soomro et al., 2026). Hybrid approaches that combine opertional data with physics-based simulations, including physics-informed machine learning (PIML) and synthetic data from validated nuclear codes, can expand training coverage. Nevertheless, simulation models introduce their own uncertainty, so uncertainty quantification remains essential before ML systems are deployed in safety-critical applications (Akins et al., 2024; Mao & Jin, 2024).
Fig. 3: Defense-in-Depth Architecture for ML Systems.
The opacity of advanced ML models, particularly deep neural networks, is another major concern. High predictive accuracy alone is insufficient in nuclear safety because regulators and operators require technically defensible explanations for important recommendations (Najar & Wang, 2024b). Physics-enhanced ML can improve interpretability by incorporating physical laws as constraints or prior knowledge, but it cannot by itself solve all explainability problems (Rengasamy et al., 2021). A complete explainability framework should describe data provenance, representativeness, hidden bias, model limitations, conditions under which predictions may become unreliable, and uncertainty estimates accompanying the output. It should also present information through human-machine interfaces that support decision-making without overloading operators. ML also affects the defense-in-depth principle. Functions that appear independent may share hidden vulnerabilities if they use the same algorithm, similar training data, or common development tools. A single unexpected condition could then produce correlated failures across several models (Ee, 2024). Safety assessment should therefore examine the entire ML lifecycle and verify diversity in data sources, preprocessing methods, training datasets, model architectures, learning approaches, and software toolchains. Only by managing these dependencies can defense-in-depth be preserved in ML-enabled nuclear systems.
Integrating ML into Safety Analysis Reports: A Methodological Framework
A structured methodology is needed to incorporate ML into nuclear power plant SARs. The framework should remain consistent with international regulatory expectations, particularly those of the IAEA, while addressing the distinctive features of data-driven systems. It should define the function of each ML component, its limitations, the uncertainty associated with its outputs, and the circumstances under which performance may change. The SAR should also describe how the ML system interacts with operators and with conventional safety systems. These elements are necessary to demonstrate that ML can be introduced without weakening the standards required for safe and reliable plant operation.
System description and safety classification form the starting point of the framework. Each ML application should be described in terms of intended function, system scope, connectivity, and safety role so that reviewers can understand where the model operates within the plant. Under a graded approach, classification should consider task criticality, consequences of failure, and the availability of independent backup or non-ML alternatives (Meschini, 2022). The analysis should address whether incorrect, unavailable, or delayed outputs could reduce safety margins or damage operational decisions, and whether operators can detect and manage ML failure (Najar & Wang, 2024c). The resulting classification determines the level of rigor required for development, V&V, and documentation, ensuring that regulatory effort is proportionate to the safety significance of the application (Burton & Herd, 2023).
The complete ML pipeline should be treated as a safety-related configurable item. Documentation should cover data collection, dataset compilation, model development, testing, validation, deployment, and maintenance because each stage can influence system safety. Records should identify training, validation, and test datasets; assess data quality and representativeness for normal, transient, and accident conditions; and disclose known biases, missing data, and dataset limitations (Burton & Herd, 2023). The selected algorithm should be justified for the intended application, including model architecture, hyper parameter tuning, regularization against overfitting, and measures for numerical stability under challenging conditions (Ma et al., 2026). Evaluation criteria should be appropriate for nuclear applications, and predictions should, where possible, be compared with physics-based benchmarks. Testing should extend beyond nominal operation to boundary conditions and abnormal transients, including loss-of-coolant accidents (LOCAs), to demonstrate performance across the expected operating envelope (Kaminski & Diab, 2024; Sallehhudin & Diab, 2021).
Uncertainty quantification (UQ) is central to ML-based safety analysis. Unlike deterministic simulation codes, ML introduces several uncertainty sources that should be identified and evaluated: epistemic uncertainty from limited or non-representative training data, aleatoric uncertainty from inherent stochastic variability in physical processes, model-structural uncertainty from architectural and modelling choices, and distributional-shift uncertainty when plant conditions diverge from the training domain (Burton & Herd, 2023; Di Maio et al., 2024). Modern UQ methods provide practical mechanisms for addressing these issues, including frameworks that couple DAKOTA with reactor simulation codes (Pang et al., 2023) and Bayesian physics-informed neural networks that estimate prediction uncertainty while incorporating physical knowledge (Nascimento et al., 2023). The SAR should state how each uncertainty source is identified, quantified, and communicated through the ML system. It should also explain how uncertainty is presented to decision-makers. Where feasible, outputs should include probability bounds, confidence intervals, or equivalent quantitative measures that support informed and risk-aware safety decisions.
Human factors integration is essential because ML outputs can shape how operators interpret plant conditions and make decisions. Consistent with the human-factors engineering principles of IAEA SSG-51 and the SAR content recommendations of IAEA SSG-61, the SAR should describe how ML-generated recommendations and supporting information are presented through the human–machine interface and should document human-factors evaluations of their effects on operator situation awareness, decision-making performance, and cognitive workload during normal operation and representative abnormal and accident scenarios (IAEA, 2019, 2021; Najar & Wang, 2024c; Simonsen et al., 2020). Operator training is equally important. Personnel must understand when ML recommendations can be relied upon, when they should be questioned, and how predefined fallback strategies are to be used if the ML system fails (Najar & Wang, 2024c). Interface design should also guard against automation complacency by keeping operators engaged in active monitoring rather than encouraging passive acceptance of automated advice.
V&V should therefore cover the entire ML lifecycle rather than only conventional software tests. Independent reviewers who were not involved in model development should assess performance under normal operation, transients, and accident scenarios to demonstrate reliability across the intended operating range (Herer, 2023). After deployment, performance must be monitored continuously to detect drift, model degradation, or unexpected behaviour. Configuration management should control model retraining, software updates, and data revisions. Version control, cryptographic hashing of model artefacts, and regression testing in a digital-twin environment should be completed before an updated model is released. This approach ensures that, even when internal model logic is difficult to interpret, the processes used to build, validate, maintain, and update the model remain transparent, traceable, and auditable.
Fig. 4: Physics-Enhanced Machine Learning (PEML) Architecture.
Experience and Applications
The integration of physics-enhanced machine learning (PEML) into probabilistic risk assessment (PRA) represents a shift from purely data-driven modelling toward hybrid methods grounded in physical laws. By embedding conservation of mass, momentum, and energy in the model structure or loss function, PEML discourages physically unrealistic predictions and improves reliability when the operating state differs from the training data. Surrogate models for accident sequence analysis illustrate this benefit.
High-fidelity thermal-hydraulic simulations, including RELAP 5/SCDAP/MOD3.4, can be used to train neural networks that reproduce transient system responses (Nguyen & Diab, 2023; Spisak & Diab, 2024). Once trained, these surrogates can evaluate scenarios much faster than full simulations while maintaining useful accuracy, including for station blackout accidents and steam generator tube ruptures where historical operational data are limited (Yarizadeh-Bene et al., 2026). Because the model is constrained by physical principles, it can generalize more effectively to off-normal conditions. PEML also improves interpretability: safety analysts can relate predictions to plant physics rather than treating them as unexplained black-box outputs. For regulatory review, this separation between physics-based constraints and data-learned parameters makes the model easier to validate, explain, and justify in safety documentation (Lye et al., 2025; Xiao et al., 2024; Karim et al., 2026).
Fig. 5: Human-Al Interaction Framework.
ML has also matured sufficiently for application in piping integrity assessment, a central aspect of nuclear power plant safety. Reactor coolant and associated piping systems are exposed to flow-accelerated corrosion and erosion, which can undermine structural integrity if not detected early (Sandhu et al., 2023). ML-based condition monitoring uses nondestructive examination (NDE) data, including vibration signals and ultrasonic testing (UT) measurements, to predict wall-thinning rates, identify early defects, and estimate remaining useful life (Soomro et al., 2026). CNNs are widely used because they can extract spatial and temporal features from sensor data that correlate with piping condition and can support high prediction accuracy (Sandhu et al., 2024). However, adoption remains constrained by limited failure data, since serious piping failures are rare and extreme-condition data are scarce. Physics-consistent synthetic data can expand training coverage without creating unrealistic degradation scenarios, while transfer learning can adapt knowledge from one piping system to another with limited labelled data (Al-Adly & Kripakaran, 2024; Zhu et al., 2023). In SARs, the reliability and robustness of these predictions must be demonstrated, particularly outside the training domain.
Conservative margins, independent analytical checks, and clearly defined application domains are needed before ML outputs are relied upon in safety-related integrity assessments. Internationally, a balanced approach is encouraged to AI in nuclear applications, recognizing benefits such as predictive maintenance, early fault detection, and real-time anomaly identification while also emphasizing concerns over transparency, lifecycle data quality, adaptive V&V, architectural separation, human-AI interface design, and the continuing authority of operators to override automated recommendations. These developments indicate broad agreement that AI may strengthen nuclear safety, but only when deployed under rigorous engineering controls, regulatory oversight, physical consistency, and sustained human involvement.
The integration of ML into nuclear power plant safety analysis represents a significant development in nuclear engineering, moving beyond strictly deter-ministic software workflows toward adaptive diagnostic and decision-support systems. Techniques such as PEML, deep auto encoders, and real-time digital twins offer improved computational speed, multivariate pattern recognition, and predictive maintenance capability, but they also introduce distinct validation and regulatory challenges. In particular, OOD behaviour shows that final-model black-box testing is not sufficient for licensing confidence. A defensible safety case requires a lifecycle assurance framework that includes disci-plined data governance, rigorous V&V, uncertainty quantification, human-centred interface design, and continuous monitoring in accordance with the intent of IAEA SSG-39 and SSG-51. Defense-in-depth must also be reconsidered so that common-cause algori-thmic failures are reduced through diversity in datasets, model architectures, and software toolchains. The most credible path is offered by physics-informed and hybrid approaches that combine empirical operating data with the constraints of established physical laws. Under independent validation and strict configuration management, such systems can improve resilience, maintenance planning, and operational reliability while preserving the conservative safety principles on which nuclear regulation depends.
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
M.G.A.: Study conception and design, Data collection, Analysis, and Manuscript preparation, K.Z.I.: Study conception and design, Data collection, Analysis, and Manuscript preparation, S.M.: Study conception and design, Data collection, Analysis, and Manuscript preparation.
We would like to acknowledge our colleagues in our field who consistently encourage us to advance and develop our scientific research.
The authors have no conflicts to disclose.
UniversePG does not own the copyrights to Supplemental Material that may be linked to, or accessed through, an article. The authors have granted UniversePG a non-exclusive, worldwide license to publish the Supplemental Material files. Please contact the corresponding author directly for reuse.
Academic Editor
Dr. Wiyanti Fransisca Simanullang, Assistant Professor, Department of Chemical Engineering, Universitas Katolik Widya Mandala Surabaya, East Java, Indonesia
Nuclear Security Division, Bangladesh Atomic Energy Regulatory Authority, Agargaon, Dhaka-1207, Bangladesh
Azam MG, Islam KZ, and Mahbub S. (2026). Machine learning integration in nuclear safety analysis reports: a regulatory framework for nuclear safety applications, Int. J. Mat. Math. Sci., 8(1), 185-200. https://doi.org/10.34104/ijmms.026.01850200