Introduction
Neutron transport equations are fundamental to nuclear reactor physics. The primary objective of solving this problem is to determine the eigenvalues and spatial distribution of the neutron angular flux densities in the reactor core. These play crucial roles in reactor engineering design, development, and operational support. The conventional methods for solving this problem include deterministic and statistical methods such as the discrete ordinates (SN) method [1], method of characteristics [2-4], and Monte Carlo method [5]. Although these have been widely used with considerable success, these have limitations in terms of accuracy, resolution ratio, and efficiency. Therefore, the improvement and enhancement of numerical methods for solving neutron transport equations remain key areas of research in the industry.
In recent years, the research on artificial intelligence and deep-learning methods for solving complex partial differential equations has become popular in computational physics [6-11]. Significant breakthroughs have been achieved in various fields, such as solving thermal–hydraulic equations [12-17]. In reactor physics, numerical computation methods based on deep learning have advanced significantly [18-21]. This research provides a strong foundation for a deep-learning-based numerical solution of the neutron transport equation. It is gaining increasing attention from the industry [22-27].
However, the neutron transport equation, which is a differential integral equation, involves multiple integrals that represent neutron scattering and fission sources depending on the energy and angular flux. The application of existing deep learning methods to solve this equation presents significant challenges. A previous study adopted a discrete approach for angular variables based on the conventional SN method [28] by constructing a loss function for machine learning that approximates and discretizes definite integral terms. However, the computational results obtained using this method have limited accuracy. Furthermore, a separate study proposed the theory of differential-order transformation [29] to solve the neutron transport equation using deep learning, thereby transforming it into a higher-order differential form to enhance accuracy. However, the method presented in this study is limited to isotropic sources.
In this paper, we propose the multi-antiderivative transformation alternating iterative deep learning method (M-AIM) to solve anisotropic scattering neutron transport equations. The integral terms are transformed into multiple exact differential equations using antiderivative functions. Neural networks represent both antiderivative and angular flux functions. The numerical solution of the neutron transport equation is approximated gradually using alternating iterations of deep learning. Additionally, the accuracy of the M-AIM using examples of typical geometric transport equations is verified.
Neutron transport equations in differential-integral format
The neutron transport equation has multiple equivalent mathematical forms. The most common of these are the differential–integral, exact integral, adjoint, and even-parity forms [1, 2, 5]. The most widely studied and applied form of the neutron transport equation is the differential integral. It is expressed as follows:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M001.png)
An anisotropic scattering source [5] can be expressed by Eq. 2:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M002.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M003.png)
The common form of the transport equation for discrete energy groups under steady-state conditions without considering external sources is [5]_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M004.png)
Simultaneously, anisotropic scattering sources (Eq. 2) are usually expanded using Legendre polynomials. The multigroup forms of their anisotropic scattering sources can be represented as_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M005.png)
In practical applications, the specific transport equation forms are usually represented by combining the geometric form of the equation, number of energy group discretization, order of anisotropic scattering sources, and presence of external sources.
The specific forms of Eq. 4 for first-order anisotropic scattering sources in a single-group flat-plate geometry are as follows:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M006.png)
The specific form of Eq. 4 for first-order anisotropic scattering sources in a spherical geometry is as follows:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M007.png)
Therefore, by moving the terms in Eq. 4 from right to left, various specific neutron transport equations such as Eq. 6 and Eq. 7 can be described using a unified partial differential integral equation._2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M008.png)
Principles of Multi-antiderivative Transformation Alternating Iterative Deep Learning Method
The M-AIM operates based on the following principle: Mathematical methods are employed to convert all the integral terms in the neutron transport equation into their corresponding antiderivative functions. This conversion transforms the neutron transport equation from its differential–integral form into an exact differential equation. Specific methods are provided for identifying unique solutions for these transformed antiderivative functions. The unknown angular flux density function and each term of the antiderivative functions are represented by separate neural networks. Substituting these neural network functions into the transport equation, a loss function is generated based on the residuals of the equation. When necessary, we formulate additional loss functions for the constraints, boundary conditions, initial conditions, and eigenvalues. These loss functions are weighted and combined into a single unified loss function. We progressively minimize the weighted loss functions by alternating iterative training by applying the M-AIM to the neural networks. After all the loss functions reduce below a set threshold, the output from the neural networks for the angular flux density function approximates the solution to the transport equation.
This section discusses in detail the transformation methods for antiderivative functions.
The angular flux ϕ(r, Ω, E, t) is a continuous function of the angular variable Ω and energy variable E in Eqs. 2 and 3. The integrand
Antiderivative function transformation of the scattering source term
With regard to Eq. 2, the integral of the scattering sources can be calculated as follows:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M009.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M010.png)
According to the principle of calculus, _2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M011.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M012.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M013.png)
Another function _2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M014.png)
Antiderivative function transformation of the fission source term
The fission-source term in Eq. 3 can be derived as_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M015.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M016.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M017.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M018.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M019.png)
Under this condition,_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M020.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M021.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M022.png)
Antiderivative function transformation of the multi-group transport equations
For the multigroup neutron transport Eq. 4, the neutron angular flux moment in the scattering source terms Eq. 5 can be transformed into an antiderivative function. Owing to the separation of the angle variables Ω, Ω′, the transformation form is simplified significantly. By setting _2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M023.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M024.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M025.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M026.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M027.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M028.png)
Therefore, substituting Eq. 26 and Eq. 27 into the multigroup neutron transport equation Eq. 4 yields_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M029.png)
Thus, Eqs. 24, 28, and 29 construct the neutron transport equations in an exact differential form under multigroup conditions.
Transformation of the two-dimensional angular parameter
When the angular variable Ω is a two-dimensional variable, it can be expressed using the directional cosines μ and azimuthal angle φ [1, 2, 5] in the neutron transport equation. Thus, the angular integral of the neutron angular flux can be expressed using Eq. 30:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M030.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M031.png)
The constraint conditions for a definite solution are as follows: μ=μ0,
Referring to Sect. 3.3, for the high-order scattering multigroup neutron transport equation with a two-dimensional angular variable, the high-order scattering antiderivative function can be defined as _2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M032.png)
Furthermore,_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M033.png)
To summarize, by antiderivative function transformation of the integral term, the neutron transport equation can be converted into an exact differential equation consisting of multiple antiderivative function terms. This is shown in Eq. 34:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M034.png)
For practical problems such as Eq. 34, if several integral terms with an identical form exist, these can be combined and simplified. The discussion is as follows:
Use _2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M035.png)
Generation of the loss function for machine learning
Equations 21, 29, and 34 were obtained by transforming the neutron transport equation. These can be solved based on various basic theories and methods for solving differential equations using deep learning. An important step was to use a neural network function as the trial function, which was substituted into the neutron transport control equation. This equation was derived from antiderivative functions, constraints of definite solutions, and constraints of boundary/initial conditions to generate the loss function. Considering Eq. 34 as an example, the function represented by a fully connected deep neural network can be expressed as [31]_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M036.png)
Additionally, the neural network can also have multiple outputs. These can be used to represent a group of functions with similar forms. For example, for a multigroup transport equation, different outputs of the same neural network can be used to represent the angular flux densities of different energy groups.
Loss function of the neutron transport control equation
Set
These equivalent relationships are substituted into Eq. 34 to obtain the loss function of the transport equation in the residual form:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M037.png)
Note that when g=g’, the input and output of the
Loss function of the equations transformed from antiderivative functions
As for the transformation equation of the scattering source integral term, shown in Eq. 24, set _2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M038.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M039.png)
Loss function of constraints of the definite solution for the antiderivative function
The constraint on the definite solution for the antiderivative function can be represented as a data-driven loss function.
With regard to the constraint condition Ω′=Ω0, _2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M040.png)
Similarly, for the constraint condition Ω′=Ω0, _2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M041.png)
Loss function of boundary condition
The neutron transport equation should have additional boundary and initial conditions to obtain a solution. In this case, we can refer to existing deep-learning methods for solving the neutron diffusion and transport equations. Based on the various forms of boundary conditions, additional boundary/initial condition loss functions should be set for the corresponding neural network functions and their corresponding differential forms [12, 13, 18, 19]. If the equation contains a time term, the corresponding initial condition should be set. The multi-group transport equation has the following form:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M042.png)
As an example, the form of a nonnegative boundary loss function can be expressed as_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M043.png)
Weighted loss function
Different neural network functions are used as objects to construct the weighted loss functions.
(1) Weighted loss function of the angular flux density
With regard to the neural network _2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M044.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M045.png)
(2) Weighted loss function of the antiderivatrtiove function
Similarly, according to the different neural network functions
The weighted loss function for the scattering source integral term neural network _2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M046.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M047.png)
If Eqs. 46, and 47 use multiple outputs of a neural network function, these can be represented as_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M048.png)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M049.png)
Framework of deep learning
After transforming the transport equation into an exact differential equation using an antiderivative function and defining the loss functions for each neural network, we can apply automatic differentiation techniques [32] to perform iterative deep learning in an alternating manner. This process is based on standard architectures such as PINN [12] and extends to improved architectures designed for solving various differential equations [33, 34]. It systematically minimizes the loss functions of all the neural networks involved. Consequently, the neural network output, which represents the neutron flux density, serves as a numerical solution to the neutron transport equation.
Please note that the eigenvalues in Eq. 34 can be derived using conventional source iteration methods. The approach presented in this study is designed to be integrated into an iterative process, specifically within the inner loop. Furthermore, we may utilize specialized deep-learning-based indirect methods to determine the eigenvalues [19]. In addition, the direct search method proposed in [35] was designed to concurrently optimize the eigenvalues and weights of the neural network. In the following discussion, we use the conventional source iteration method to solve for keff.
In the multi-group transport equation, the deep learning process utilizes weighted loss functions, as detailed in Eqs. 44–49. It involves iteratively adjusting the neural network functions
In the case of a neural network with a single output, the convergence criterion dictates that all the loss functions and iterative parameters of the source-iteration method should be below a predefined threshold ε. When this condition is satisfied, the training procedure is as illustrated in Fig. 1.
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-F001.jpg)
In this figure, the loss functions for the unknown function of the angular flux, as well as for the antiderivative function of the scattering and fission sources, are defined by Eqs. 44, 46, and 47.
Numerical Results
Critical homogeneous single-group problem of slab geometry
To facilitate a comparison with the theoretical values, we first verified the critical form of the single-group planar geometric first-order anisotropic scattering source Eq. 6. keff was set equal to 1.0. A set of geometric and cross-sectional parameters that are theoretically critical for verification were identified.
The isotropic form of Eq. 6 for a single group is_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M050.png)
Problem setup: We assumed slab material properties ∑t=0.050 cm-1, ∑s0=0.030 cm-1, ∑s1=0.001705 cm-1, and
For the single-group form of the first-order scattering equations (Eq. 6) according to the M-AIM described in Sect. 3.3 and the simplified form Eq. 32, we obtain_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M051.png)
It is observed that in Eq. 51, both zeroth-order scattering and fission sources are isotropic, and
The machine learning method and parameter selection were as follows: the neural network was set to (n = 20) and (l = 9), the tanh activation function was applied, the network was initialized with a Gaussian random distribution, and the ADAM machine learning algorithm was employed [12, 19, 29]. The learning rate started at 1 × 10-4 and decreased gradually to 5 × 10-7 as the number of learning iterations increased. It stopped after 1 × 10-4 iterations. To accelerate the convergence, referring to the literature [19, 29], a feature value constraint ϕ(0,±1)=0.2 was set. The number of sample points for the control equation was 8000, and 100 sample points were considered under the boundary conditions. The weight coefficient of the loss function is an empirical parameter. The strategy of this study was to use Pϕ=1 as a benchmark and periodically replace different values of Pb, PF,s, PF,f, and Pd,s with values ranging from 10 to 50 during the training process. The hardware was equipped with an Intel Core i9-11900K CPU and NVIDIA RTX A2000 GPU.
Figure 2 shows the flux density equation ϕ(x,μ), antiderivative function of the zeroth-order scattering source/fission source
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-F002.jpg)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-F003.jpg)
As shown above, the application of deep-learning techniques successfully produces flux solutions that align well with the theoretical values. Furthermore, it provides a continuous flux distribution under anisotropic scattering conditions. It is noteworthy that the impact of the higher-order scattering terms
| Method | Different values for x/b | |||||
|---|---|---|---|---|---|---|
| 0(center) | 0.25 | 0.50 | 0.75 | 1.00 (boundary) | Time(min) | |
| Actual calculation value (M-AIM) | 0.49994469 | 0.47717363 | 0.40185571 | 0.27500153 | 0.13455909 | 505 |
| Normalized calculation value (M-AIM) | 1 | 0.95445284 | 0.80380034 | 0.55006390 | 0.26914796 | / |
| Normalized theoretical value | 1 | 0.94714400 | 0.79372641 | 0.55329025 | 0.21419206 | / |
| Normalized VNM value | 1 | 0.94965777 | 0.79561143 | 0.55840989 | 0.23590973 | 0.03 |
| Normalized OpenMC value | 1 | 0.94440551 | 0.81270187 | 0.55785209 | 0.25272043 | 2.5 |
| Normalized ResNet value | 1 | 0.88963412 | 0.77045493 | 0.54105856 | 0.26854225 | 450 |
| M-AIM: mean squared error relative to theoretical value | 7.59151071×10-4 | |||||
| M-AIM: mean squared error relative to VNM value | 3.58673270×10-4 | |||||
| M-AIM: mean squared error relative to OpenMC value | 2.56903029×10-4 | |||||
| ResNet: mean squared error relative to theoretical value | 1.54500122×10-3 | |||||
Non-critical multi-material single-group problem of spherical geometry
Problem setup: Two materials are used for spherical fuels (Fig. 4) with vacuum boundary conditions on the outer surface.
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-F004.jpg)
The spherical geometry consists of two regions composed of two materials. The material parameters are listed in Table 2.
| Material number | ∑t (cm-1) | ∑s,0 (cm-1) | ∑s,1 (cm-1) | |
|---|---|---|---|---|
| ① | 0.05 | 0.03 | 0.001705 | 0.0255 |
| ② | 0.05 | 0.03 | 0.001705 | 0.023625 |
The basic equation for the anisotropic transport problem of a single-group spherical scattering source is given by Eq. 7. Applying the antiderivative function transformation method described in Sect. 3.3, we obtain_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M052.png)
The machine learning method and parameter selection were as follows: the neural network was set to n=40 and l=9, and the eigenvalue constraint was ϕ(0,μ)=0.5, and the other parameters were identical to those in Sect. 6.1.
Figure 5 presents the numerical results of the equation’s flux density ϕ(x,μ), zeroth-order scattering source/fission source antiderivative function
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-F005.jpg)
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-F006.jpg)
| Method | keff | keff/relative error | Mean squared error of scalar flux | Time (min) | Lossϕ,all | |
|
|---|---|---|---|---|---|---|---|
| M-AIM | 1.0218 | -0.0053 | 1.9200×10-3 | 764 | 6.7260×10-4 | 3.2938×10-6 | 1.3420×10-9 |
| OpenMC | 1.0272 | / | / | 2.0 | / | / | / |
Non-critical homogeneous two-group problem of slab geometry
Problem setup: For the two-group single-material region problem in a slab geometry, with geometric parameters as in Sect. 6.1 and the material parameters listed in Table 4, the basic forms of the two-group slab theory and transport equations are given by Eq. 6 and a detailed explanation in references [1, 2, 5]. Referring to Eq. 29 for the antiderivative function transformation equation, we obtain_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M053.png)
| g | ∑t (cm-1) | |
χg | ∑s,0,g-1 (cm-1) | ∑s,0,g-2 (cm-1) | ∑s,1,g-1 (cm-1) | ∑s,1,g-2 (cm-1) |
|---|---|---|---|---|---|---|---|
| 1 | 0.33285210 | 0.01266922 | 1 | 0.31876071 | 0.00093413 | 0.00730600 | 0.00004500 |
| 2 | 0.44925527 | 0.22110247 | 0 | 0 | 0.31466259 | 0 | 0.00129800 |
The machine learning method and parameter selection were as follows: The eigenvalue constraint for the fast group was
Because this problem does not have an analytical solution or theoretical estimate, the results of the M-AIM were compared with the normalized scalar flux calculation results obtained using conventional methods (both normalized based on the scalar flux value at x = 0 for the fast group). This is shown in Fig. 7 and Table 5. The results of the loss functions are listed in Table 6.
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-F007.jpg)
| Method | keff | keff/relative error | MSE of flux | Time (min) | |
|---|---|---|---|---|---|
| fast group | thermal group | ||||
| OpenMC | 0.9710 | -0.0076 | 5.1988×10-4 | 1.3964×10-3 | 6 |
| VNM | 0.9700 | -0.0037 | 4.9479×10-4 | 1.3614×10-3 | 0.06 |
| M-AIM | 0.9636 | / | / | / | 2360 |
| Energy group | Lossg,all | |
|
|---|---|---|---|
| fast(g=1) | 5.0392×10-4 | 4.7992×10-7 | 6.6793×10-9 |
| thermal(g=2) | 1.4554×10-5 | 3.5532×10-4 | 1.4946×10-8 |
Fixed-source homogeneous single-group problem of infinite cylindrical geometry
Problem setup: For the fixed-source single-group infinite cylindrical geometry single-material region problem, the parameter radius R=1.08225766 m. The material parameters are identical to those described in Sect. 6.1 under the zero-flux boundary condition. According to literature [1, 2, 5], the basic form of a transport equation with first-order anisotropy is as follows:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M054.png)
The above equation was transformed using the method described in Sect. 3.4 and the simplified transformation form of Eq. 34 to obtain:_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-M055.png)
In Eq. 55, it is given that
At this point, the following are the constraints of the definite solution for the antiderivative function:
The angular flux density ϕ(r,μ,φ) is represented by neural network
The machine learning method and parameter selection were as follows: The neural network was set with n=100 and l=5. The other parameters were identical to those in Sect. 6.1.
The zeroth-order scattering source antiderivative function
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-F008.jpg)
Table 7 compares the results and final loss function. MSEMC represents the mean squared error with respect to OpenMC. MSESN represents the mean squared error with respect to the SN-DFEM. The results indicate that the M-AIM and DL method display comparable accuracies. However, the training time of the M-AIM was approximately 93% shorter than that of the DL method. Compared with the DL method for solving the original differential-integral form of the transport equation (Eq. 54), the M-AIM for solving Eq. 55 has a significant advantage in terms of computational efficiency.
| Method | MSEMC | MSESN | Time (min) | |
|
|
|
Lossϕ,all |
|---|---|---|---|---|---|---|---|---|
| M-AIM | 6.4125 × 10-6 | 6.6431 × 10-6 | 144.2 | 1.0242×10-6 | 1.1836×10-6 | 1.0453×10-6 | 2.0745×10-6 | 1.8186 × 10-5 |
| DL | 7.9087 × 10-6 | 5.7365 × 10-6 | 2005.7 | / | / | / | / | 1.7098 × 10-5 |
To analyze the performance of the adjustment strategies, several numerical experiments were conducted. We set the weight of the governing equation loss term to one. The other loss terms used different weight values. The first five sets used fixed weight values, whereas the last three sets employed a dynamic adjustment strategy. All the cases were trained 500 times. Table 8 presents the results for different PB weight values. Figure 9 shows the weight values of the PB in the fixed-source homogeneous single-group problem with an infinite cylindrical geometry. The results indicate that the weight values of the loss terms affect the accuracy of the model. Experiments 1–5 could not achieve optimal computational results using fixed weight values. In contrast, experiments 6–8, which employed adaptive adjustment strategies, dynamically adjusted the weight values within the range of 10–50. This yielded a higher precision. Therefore, in actual hyperparameter settings, a dynamic adjustment strategy can be adopted to achieve better computational accuracy. The actual dynamic adjustment strategies depend on specific problems, and the hyperparameters can be determined and optimized based on experiments and experience. No “one-size-fits-all” adjustment strategy exists for any problem.
| Number | Weight | MSEmc | MSEsn | Lossϕ,all |
|---|---|---|---|---|
| 1 | 100 | 1.1176×10-4 | 1.1459×10-4 | 4.0336 × 10-6 |
| 2 | 80 | 3.3023 × 10-4 | 3.2613 × 10-4 | 2.0588 × 10-6 |
| 3 | 50 | 1.0796×10-4 | 1.0657×10-4 | 1.9826 × 10-6 |
| 4 | 30 | 2.0696 × 10-5 | 1.5756 × 10-5 | 1.0382×10-6 |
| 5 | 10 | 4.1371 × 10-5 | 4.0113 × 10-5 | 7.6040 × 10-7 |
| 6 | strategy 1 | 1.3919 × 10-5 | 1.0784×10-5 | 1.4721 × 10-6 |
| 7 | strategy 2 | 1.6479 × 10-5 | 1.1058×10-5 | 7.1233 × 10-7 |
| 8 | strategy 3 | 1.5255 × 10-5 | 1.0176×10-5 | 7.6950 × 10-7 |
_2026_05/1001-8042-2026-05-83/alternativeImage/1001-8042-2026-05-83-F009.jpg)
To summarize, the above verification demonstrates that the proposed M-AIM is applicable to various critical/noncritical neutron transport problems under different conditions including different energy groups, typical geometries (slab, sphere, and cylinder), and one/two-dimensional angular variables. It is also suitable for fixed-source and eigenvalue problems that involve higher-order scattering. It provided continuous numerical solutions for the neutron transport equation under anisotropic scattering conditions. The numerical calculation process was standardized, and the results showed good precision.
However, the aforementioned examples indicate that compared with conventional numerical methods, deep learning requires more time. This is a common challenge confronted by the current numerical methods using deep learning. The directions for improvement include using higher-performance GPUs and adopting large-scale parallel technology to reduce the training time. In addition, improvements can be achieved using transfer learning, training differential operators, and optimizing the structure of deep neural networks to reduce the training time.
Summary
This study addressed the challenges confronted by the current deep learning techniques when solving the neutron transport equation with anisotropic scattering sources. It introduced the multi-antiderivative transformation alternating iterative deep learning method (M-AIM). The method employs multiple antiderivative functions to transform the transport equation. This transformation converts the integral–differential form of the transport equation into an exact differential equation suitable for deep-learning-based solutions. Subsequently, multiple neural network functions were used to map the unknown angular flux and antiderivative functions. Deep learning was then performed using an alternating iterative method to obtain numerical solutions for each antiderivative function and unknown angular flux density function. Finally, the correctness of the theory and related methods was verified through the numerical results of typical problems. This research has laid the foundation for solving more complex neutron transport equations in geometry and energy groups using deep learning methods. New technical approaches for numerical calculation methods of anisotropic scattering neutron transport equations were evaluated.
The current numerical solution methods for transport equations using PINN are still in the early stages of research. Numerous unknown challenges need to be addressed urgently. The areas that deserve particular attention include the following:
Enhancing the training efficiency of deep learning models: As described in the numerical experiments in this paper, the current training time for deep learning models is relatively long, and the convergence speed should be improved significantly. Adopting advanced neural network structures, machine learning optimization algorithms, and large-scale parallel technology based on specialized GPU hardware is recommended to enhance the computational efficiency.
Strengthening the generalization capability of deep learning models: Currently, deep learning models exhibit good generalization capabilities for non-training data samples within the solution space. However, for regions outside the solution space, particularly for problems with different cross-sections and geometries, the generalization capability of the model is weak. It is necessary to achieve improvements through methods such as transfer learning and the training of differential operators.
There is an urgent need to perform uncertainty analyses and assessments of deep learning models to enhance their interpretability. Numerical solution methods and the related software for neutron transport, which play a role in reactor engineering, are crucial for nuclear safety. Thus, these require rigorous verification and uncertainty analysis. However, the “black box” characteristics of current deep neural networks complicate this process. This necessitates a thorough uncertainty analysis and assessment before these can be applied in practice. A feasible approach is to provide a comprehensive uncertainty assessment for deep-learning models by referring to the evaluation methods of conventional numerical methods. Furthermore, it is important to develop new methods that can improve the interpretability of deep-learning models and ensure their reliability and safety.
Heterogeneous parallel high-order scattering MOC and its application to simulation of critical experiment
. At. Energy Sci. Technol. (in Chinese) 58(01), 135–143 (2024). http://dx.chinadoi.cn/10.7538/yzk.2023.youxian.0099Development and verification of 2D kinetic MOC code based on GPU
. At. Energy Sci. Technol. (in Chinese) 56(1), 87–95 (2022). http://dx.chinadoi.cn/10.7538/yzk.2021.youxian.0502DPM: A deep learning PDE augmentation method with application to large-eddy simulation
. J. Comput. Phys. 423,Deep PDE solution to BSDE
. Digit. Finance. 6, 727–758 (2024). https://doi.org/10.1007/s42521-023-00098-6Solving multi-dimensional optimal stochastic control problems with deep learning and evolution algorithm
. Adv. Appl. Math. (in Chinese) 11(3), 1222–1241 (2022). https://doi.org/10.12677/AAM.2022.113133A compound neural network for learning partial differential equations from noisy data
. Chin. J. Comput. Phys. (in Chinese) 39(2), 223–232 (2022). https://doi.org/10.19596/j.cnki.1001-246x.8364FDM-PINN: Physics-informed neural network based on fictitious domain method
. Int. J. Comput. Math. 100(3), 511–524 (2023). https://doi.org/10.1080/00207160.2022.2128674Physics-informed encoder-decoder network based on ConvLSTM for solving time-dependent partial differential equations
. Adv. Appl. Math. (in Chinese) 13(2), 714–722 (2024). https://doi.org/10.12677/AAM.2024.132069Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
. J. Comput. Phys. 378(1), 686–707 (2019). https://doi.org/10.1016/j.jcp.2018.10.045Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators
. Nat. Mach. Intell. 3(3), 218–229 (2021). https://doi.org/10.1038/s42256-021-00302-5Method of predicting transient thermal hydraulic parameters of the core based on the gated recurrent unit model of soft attention
. Nucl. Tech. (in Chinese) 47(1),Study on calculation method of Courant limit in thermal hydraulic system analysis code
. Nucl. Power Eng. (in Chinese) 42(4), 63–67 (2021). https://doi.org/10.13832/j.jnpe.2021.04.0063Development and application of three-dimensional nuclear steam generator thermal-hydraulic analysis code STAF
. At. Energy Sci. Technol. (in Chinese) 56(11), 2239–2252 (2022). https://doi.org/10.7538/yzk.2022.youxian.0590Progress of convolution neural networks in flow field reconstruction
. Chin. J. Mech. (in Chinese) 54(9), 2343–2360 (2022). http://dx.chinadoi.cn/10.6052/0459-1879-22-130Neural network based deep learning method for multi-dimensional neutron diffusion problems with novel treatment to boundary
. J. Nucl. Eng. 2(4), 533–552 (2021). https://doi.org/10.3390/jne2040036Solving multi-dimensional neutron diffusion equation using deep machine learning technology based on PINN model
. Nucl. Power Eng. (in Chinese) 43(2), 1–8 (2022). https://doi.org/10.13832/j.jnpe.2022.02.0001A data-enabled physics-informed neural network with comprehensive numerical study on solving neutron diffusion eigenvalue problems
. Ann. Nucl. Energy. 183,Surrogate modeling for neutron diffusion problems based on conservative physics-informed neural networks with boundary conditions enforcement
. Ann. Nucl. Energy. 176,Data-enabled physics-informed machine learning for reduced-order modeling digital twin: Application to nuclear reactor physics
. Nucl. Sci. Eng. 196(6), 668–693 (2022). https://doi.org/10.1080/00295639.2021.2014752On the uncertainty analysis of the data-enabled physics-informed neural network for solving neutron diffusion eigenvalue problem
. Nucl. Sci. Eng. 198(5), 1075–1096 (2024). https://doi.org/10.1080/00295639.2023.2236840Physics-constrained neural network for solving discontinuous interface K-eigenvalue problem with application to reactor physics
. Nucl. Sci. Tech. 34, 161 (2023). https://doi.org/10.1007/s41365-023-01313-0Automatic boundary fitting framework of boundary dependent physics-informed neural network solving partial differential equation with complex boundary conditions
. Comput. Method. Appl. M. 414,Neural network study of the nuclear ground-state spin distribution within a random interaction ensemble
. Nucl. Sci. Tech. 35, 64 (2024). https://doi.org/10.1007/s41365-024-01424-2Research on inversion method for complex source-term distributions based on deep neural networks
. Nucl. Sci. Tech. 34, 195 (2023). https://doi.org/10.1007/s41365-023-01327-8Solving the linear transport equation by a deep neural network approach
. Discrete. Cont. Dyn-S. 15(4), 669–686 (2022). https://doi.org/10.3934/dcdss.2021070Differential transform order theory for solving neutron transport equation by deep learning method
. At. Energy Sci. Technol. (in Chinese) 57(05), 946–959 (2023). https://doi.org/10.7538/yzk.2023.youxian.0002Automatic differentiation in machine learning: a survey
. J. Mach. Learn. Res. 18(153), 1–43 (2018). http://jmlr.org/papers/v18/17-468.htmlB-PINNs: Bayesian physics-informed neural networks for forward and inverse PDE problems with noisy data
. J. Comput. Phys. 425,Self-adaptive physics-informed neural networks
. J. Comput. Phys. 474,The deep learning method to search effective multiplication factor of nuclear reactor directly
. Nucl. Power Eng. (in Chinese) 44(05), 6–14 (2023). https://doi.org/10.13832/j.jnpe.2023.05.0006The openMC Monte Carlo particle transport code
. Ann. Nucl. Energy. 51, 274–281 (2013). https://doi.org/10.1016/j.anucene.2012.06.040An efficient parallel algorithm of variational nodal method for heterogeneous neutron transport problems
. Nucl. Sci. Tech. 35, 69 (2024). https://doi.org/10.1007/s41365-024-01430-4A comprehensive study of non-adaptive and residual-based adaptive sampling for physics-informed neural networks
. Comput. Method. Appl. M. 403,Solving the linear transport equation by a deep neural network approach
. Discrete Cont. Dyn-S. 15(4), 669–686 (2022). https://doi.org/10.3934/dcdss.2021070Solving Fredholm integral equations using deep learning
. Int. J. Appl. Comput. Math. 8(2), 87 (2022). https://doi.org/10.1007/s40819-022-01288-3TARS: A parallel tetrahedral discontinuous finite element code for the solution of the discrete ordinates neutron transport equation
. Ann. Nucl. Energy. 196,The authors declare that they have no competing interests.

