Introduction
Muons were discovered in cosmic rays by Anderson and Neddermeyer in 1937. These muons are generated when primary cosmic rays collide with atomic nuclei in the upper atmosphere, initiating nuclear electromagnetic cascades. Most cosmic-ray muons originate from the decay of charged pions and kaons generated by these interactions. Below 1 GeV, the cosmic-ray muon energy spectrum is nearly flat, but it steepens in the energy range of 10–100 GeV, closely following the primary cosmic-ray spectrum. Above 100 GeV, the spectrum becomes even steeper because high-energy pions are more likely to interact with the atmosphere before decaying into muons [1]. At sea level, they are the most abundant charged particles, with an intensity of approximately 1 cm-2 min-1 [2, 3]. Similar to other charged particles, muons interact with atomic matter, leading to energy loss and multiple scattering. However, their interactions with matter are purely electroweak, resulting in significantly lower energy loss compared to most other particles, which grants them exceptional penetration capabilities. Consequently, over the past few decades, muon tomography has played an important role in detection and imaging [4-7].
In 1970, muon transmission was first developed to discover new chambers inside a pyramid by Alvarez et al. [8]. Since then, muon techniques have been widely applied in many fields, such as nuclear safeguards [9], volcano studies [10], and underground tunneling [11]. Cheng et al. applied muon radiography to investigate the internal density distribution of the Laoheishan volcanic cone [12]. In 2003, the Los Alamos National Laboratory (LANL) first introduced the application of muon scattering tomography to the field of security detection and material identification [13-15], underscoring its immense potential in detecting special nuclear materials, such as illicit uranium, concealed within cargo and containers. Leveraging the distinctive physical properties of muons, this technology has become an effective method for detecting large-scale and high-density objects. As a grouping of materials based on their atomic number, typically divided into low-Z, mid-Z, and high-Z categories, Z classification reflects both physical characteristics (e.g., scattering behavior) and practical needs in inspection and nuclear verification tasks [16]. Z-class identification of materials based on muon scattering is a crucial task for security screening and industrial applications. Xiao et al. presented a modified multigroup model to improve the image resolution of high-Z materials [17]. A method for imaging materials using the ratio of secondary particles produced by muons was proposed by Ji et al. [18]. A novel imaging reconstruction method with a large voxel and angle capping was proposed to reduce the time and storage consumption [19] at the Institute of Modern Physics, CAS. Most traditional identification methods rely on the complicated reconstruction of the muon event and track fitting process, which significantly increases the design and calculation costs of the algorithm.
Compared with traditional physics-based reconstruction methods, deep learning models can automatically extract and learn complex features from data. This end-to-end learning paradigm significantly improves computational accuracy and efficiency [20, 21]. Deep learning techniques have demonstrated outstanding performance in compressed sensing technique, which is well-suited for image reconstruction and array processing [22-24]. Meanwhile, in the field of physics, deep learning has demonstrated great potential, especially in the nuclear physics and particle physics areas [25-27]. Common deep learning methods rely on supervised learning, which is based on labeled training samples. Gao et al. proposed a convolutional neural network model for feature extraction to classify and recognize materials based on muon scattering [28]. Our previous work also explored a feasible solution for muon-based material identification using supervised deep learning [29, 30]. However, in practical identification scenarios, obtaining labeled scattering angle data for coated materials is a major challenge. The scarcity of labeled data limits the practical application of supervised learning in the prediction of coated materials.
As a critical deep learning strategy, transfer learning enables knowledge acquired from one domain (the source domain) to be transferred to another related domain (the target domain), effectively mitigating the problem of limited labeled data in the target domain [31-33]. This inspired the proposal of a lightweight neural network model based on transfer learning for Z-class identification of coated materials using muon scattering data. In this study, we defined the Z-class identification of bare materials as the source domain task and that of coated materials as the target domain task. This formulation is inspired by realistic application scenarios, such as cargo inspection [34], nuclear safeguard verification, and arms control [35, 36], where the internal composition is unknown or only partially known. By introducing a novel data preprocessing and sampling method, the model transfers the feature-label mapping learned from the bare material scattering angles to the Z-class prediction of the coated materials. We designed two coated material scenarios, where Al and PE served as coating materials (Fig. 1). Two transfer learning paradigms, fine-tuning [37] and Domain-Adversarial Neural Network (DANN) [38] to achieve Z category classification for nine coated materials, including three materials from each of the high, mid, and low-Z categories.
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-F001.jpg)
To evaluate the effectiveness of our proposed approach, we conducted a series of simulations on the Z-class classification task with muon scattering angle data using the Geant4 Monte Carlo simulation [39]. The results demonstrate that this approach effectively transfers scattering angle features from bare materials, enabling the accurate classification of coated materials. Our method not only reduces the reliance on large-scale labeled data for coated materials but also maintains excellent Z-class classification accuracy in a few minutes, particularly achieving superior accuracy in the more application-critical high-Z class identification. The contributions of this study are as follows:
Proposing a novel sampling method combining inverse cumulative distribution function (CDF), integration, and interpolation improves the feature expression ability of muon scattering angle samples, optimizing the training effect of neural network.
Development of a fine-tune-based transfer learning model for fast Z-class prediction in the case of scarcity of coated material labels, a DANN transfer learning model based on an adversarial idea for stable Z-class prediction when the coated material labels are completely unknown.
The comprehensive result analysis and physical correlation interpretation. Demonstration of the potential application scenarios of each method, which provides a diversified scheme for the application and expansion of transfer learning in the field of muon techniques, such as cargo inspection and nuclear safeguards.
To the best of our knowledge, this is the first study to apply transfer learning strategies to muon tomography. This study provides an alternative machine learning scheme to the traditional identification method. This demonstrates the great potential of transfer learning in mitigating the high-cost reconstruction and data scarcity challenges in muon-based applications with similar scenarios.
Muon scattering simulation and sampling method
Simulation setups
The scattering angle data were provided by a Geant4 simulation incorporating the Cosmic-Ray Shower (CRY) Library [40], which generates muons with the energy and angular distribution of cosmic muons at sea level. The dimension of the object is 10 cm × 10 cm × 10 cm; it may be coated on all six sides with a 1 cm thick layer, resulting in an overall size of 12 cm × 12 cm × 12 cm for the coated object. Four position-sensitive detectors of size 30 cm × 30 cm are placed above and below the object to measure the trajectories of the incident and emergent muons for subsequent angle calculations. A spacing of 20 cm is placed between the second and third detectors to place the tested materials at the center of this gap. The distance between detectors 1 and 2 and detectors 3 and 4 was the same (35 mm). It is worth noting that the detector used in the simulation is an ideal detector, and detector characteristics such as detector noise are not considered. This setup is shown in Fig. 2.
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-F002.jpg)
First, we set up a muon source with an energy of 1 GeV and conducted scattering angle simulations for nine bare materials (Mg, Al, Ti, Fe, Cu, Zn, W, Pb, and U) to verify the reliability of the simulated data. Figure 3 presents the probability distribution statistics of the simulated scattering angle data using the kernel density estimation (KDE), categorized by both Z classes and different materials. The results indicated significant differences in the scattering angle distributions among different Z-class materials and among materials within the same Z-class. In the formal experiment, we performed simulations based on the energy and angular distributions of cosmic-ray muons at the sea level. For each material in different scenarios (bare, Al-coated, or PE-coated), 500,000 muon scattering angle data were simulated. The scattering angle in our simulation is the included angle between the incoming and outgoing muon tracks in 3D space using the following formula:_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-M001.png)
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-F003.jpg)
Inverse CDF sampling
As shown in Fig. 3, the primary distinguishing feature of muon scattering through different materials lies in the variations of their scattering angle distributions. Additionally, high-quality training samples contribute to the training of more effective neural networks [41, 42]. When employing a neural network as a mapping function between scattering angle data and material properties, it is desirable for the model to learn features that facilitate differentiation between various materials based on their respective scattering characteristics. To effectively capture the distribution characteristics of the simulation data, a sampling method called the inverse cumulative distribution function was introduced. First, we uniformly selected a set of quantile points within the range of the overall scattering angle distribution for a given material. Non-parametric methods are used to estimate the probability density function (PDF) of the overall data at these quantile points. The corresponding CDF values were then computed based on these PDF values. Finally, we obtained the training samples through inverse CDF sampling. The complete sampling process is shown in Fig. 4, and the process pseudocode is referred to Appendix 1. Compared with the traditional random sampling (RS) method, the samples generated using this method shared the same probability distribution as the overall simulation data.
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-F004.jpg)
We employed two non-parametric estimation methods: KDE [43] and histogram estimation (HE), to compute the probability density values for selected quantile points within the overall data. The fundamental idea of the KDE is to use a smooth kernel function to perform weighted averaging of the total data points, thereby obtaining an estimate of the probability density. In other words, the weighted contribution of data points at a given quantile point is calculated. Selecting specific quantile points rather than drawing the entire data ensures computational efficiency and accurate distribution estimation, mitigating the impact of fine-grained details that could disrupt the smoothness of the overall probability distribution function. For a selected quantile point Xq, the kernel density estimation of the PDF is given by:_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-M002.png)
In contrast, the HE method is more straightforward and intuitive. It uniformly divides the overall data range into b bins and counts the number of data points that fall within each bin:_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-M003.png)
To convert discrete PDF values into continuous CDF values, the compound trapezoidal rule (Eq. 4) and cubic spline interpolation (Eq. 5) are employed during the converting process. The integral interval is divided into multiple sub-intervals, and the trapezoidal rule is applied within each sub-interval to improve the integration accuracy._2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-M004.png)
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-M005.png)
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-M006.png)
| Parameter | Value | Description |
|---|---|---|
| tot_mat | 9 | Total number of materials |
| sim_data | 500,000 | Simulated muon scattering data for each material |
| smp_num (Ns) | 1,000 | The number of samples generated for a material |
| scat_num (n) | 500 | The number of muon scattering data contained in each sample |
| qtl_point | 1,000 | The number of quantile points where KDE calculating PDF values |
| bin_num | 1,000 | The number of bins where histogram counting PDF values |
| bw | Silverman | Bandwidth adjustment mode in KDE |
| sim_thd | 5 × 10-3 | The threshold of similarity discrimination between samples |
In different scenarios (bare, Al, or PE coated), 1,000 samples were acquired for each material. While in the specific training and prediction phases, a consistent train-test split was applied across different materials. The specific data partitioning ratios for the various phases are detailed in the corresponding parameter tables. In addition, the traditional random sampling method, which is directly sampled from the raw simulated scattering angle data (denoted as RS), was also applied in this study. The sample size generated by the RS method was the same as that of the inverse CDF sampling. The superiority of the inverse CDF sampling method can be demonstrated by comparing the training accuracies.
Transfer learning-based Z-class identification
There are some implicit correlation features between source domain tasks and target domain tasks, which constitute the practical feasibility basis for transfer learning [31-33]. In this study, we adopted two transfer learning paradigms: fine-tuning learning and adversarial transfer learning with DANN. Based on a pre-trained model trained in the source domain, fine-tuning performs limited parameter adaptation in the target domain and enables the efficient learning of feature-label relationships, even when the target domain data are scarce. Adversarial transfer learning, on the other hand, is applicable when target domain labels are completely unknown. By extracting shared discriminative features, the feature distributions between the source and target domains are aligned, enabling unsupervised classification.
Pre-training and Fine-tuned transfer
Fine-tuning is an essential technique for transferring neural network tasks and is widely applied in transfer learning because of its low computational cost and high training efficiency under limited target domain data conditions. In this study, we constructed a unified lightweight neural network with two hidden layers for both the pre-training and fine-tuning processes (P&F model). The scattering angle sample data were first received by the input layer, then processed through two hidden layers for feature extraction, and finally classified by the output layer. The detailed network structure is illustrated in Fig. 5.
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-F005.jpg)
Because the P&F model is designed for multi-classification tasks, we used cross-entropy as the loss function [44]:_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-M007.png)
| Parameter | Value | Description |
|---|---|---|
| input_dim | 500 | The number of nodes in the P&F model input layer |
| hidden_dim1 | 128 | The number of nodes in the P&F model first hidden layer |
| hidden_dim2 | 64 | The number of nodes in the P&F model second hidden layer |
| output_dim | 3 | The number of nodes in the P&F model output layer |
| tt_ratiop* | 7:3 | Ratio of source samples in training set to test set during pre-training |
| tt_ratiof* | 3:7 | Ratio of target samples in training set to test set during fine-tuning |
| batch_size | 128 | The number of samples trained for each material in one iteration |
| epochp* | 200 | The number of epochs in the pre-training process |
| epochf* | 100 | The number of epochs in the fine-tuning process |
| lr | 5×10-5 | The learning rate of the P&F model in the training process |
For the pre-training process, we trained the P&F model with feature-label pairs (xs, ys) samples obtained from the source domain using three different sampling methods. The training process is illustrated in Fig. 6. The goal was to learn the mapping between the scattering angle data of bare materials and their Z categories. After pre-training, the hidden layers, which serve as the key structures for feature extraction, have their parameters optimized to effectively extract high-order features from the original scattering-angle data. In the subsequent fine-tuning process, the parameters of the input and hidden layers are frozen, while only the parameters between the last hidden layer and the output layer are trained using a smaller number of target domain training samples.
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-F006.jpg)
In the pre-training stage, we aimed at nine different materials from the high-, mid-, and low-Z categories and conducted supervised training on the training dataset. The training results in Table 3 indicate that the pre-trained model achieves high prediction accuracy on the test dataset, confirming the effectiveness of neural networks in learning the mapping between scattering angle data and Z categories. The accuracy metric is the ratio of correct predictions to the total sample number in their respective scenarios.
| Sampling method | Z categories | Total | ||
|---|---|---|---|---|
| low-Z | mid-Z | high-Z | ||
| RS | 0.972 | 0.958 | 0.997 | 0.976 |
| KDE | 0.996 | 0.999 | 1.000 | 0.998 |
| HE | 0.999 | 0.999 | 1.000 | 0.999 |
Before applying the fine-tuned transfer learning method, we first directly evaluated the pre-trained model on the two target domains, where Al and PE served as coating materials. The test results are shown in Table 4 (Pre-training). When trained with samples obtained using two inverse CDF sampling methods (KDE and HE), the classification accuracy in the Al-coated target domain decreased by approximately 12%, with the most significant drop occurring in the prediction accuracy of low-Z materials (approximately 30%). This phenomenon can be attributed to the fact that, as a low-Z material, the relative influence of the Al coating on the scattering angle distribution of the coated material is inversely proportional to the intrinsic Z value of the coated material. Meanwhile, in the PE-coated target domain, the prediction accuracy of the pre-trained model remained almost unchanged (a decrease of only approximately 1%). Because PE, as a hydrocarbon compound, can be considered an ultra-low-Z material, its coating has a minimal impact on the scattering angle distribution of metallic materials. In contrast, when the model was trained with RS data, the randomness of the sampling process hindered the prediction results from following the physically consistent patterns observed in the inverse CDF-based training.
| Training stage | Dataset | Sampling method | Z categories | Total | ||
|---|---|---|---|---|---|---|
| low-Z | mid-Z | high-Z | ||||
| Pre-train | Al | RS | 0.980 | 0.784 | 0.816 | 0.860 |
| KDE | 0.697 | 0.933 | 1.000 | 0.877 | ||
| HE | 0.643 | 0.980 | 0.999 | 0.874 | ||
| PE | RS | 0.996 | 0.794 | 0.938 | 0.909 | |
| KDE | 0.943 | 0.997 | 1.000 | 0.980 | ||
| HE | 0.987 | 0.977 | 1.000 | 0.988 | ||
| Fine-tune | Al | RS | 0.914 | 0.923 | 0.985 | 0.941 |
| KDE | 0.981 | 0.977 | 0.998 | 0.985 | ||
| HE | 0.957 | 0.943 | 0.989 | 0.963 | ||
| PE | RS | 0.949 | 0.957 | 0.996 | 0.968 | |
| KDE | 0.995 | 0.993 | 1.000 | 0.996 | ||
| HE | 0.985 | 0.984 | 0.996 | 0.988 | ||
In our P&F model architecture, the first hidden layer was designed to extract fundamental features reflecting low-level scattering physics. These features are largely invariant across domains, given the shared nature of the input data structures between the source and target domains. Therefore, freezing the first hidden layer preserves these transferable and physically meaningful representations. Regarding the second hidden layer, although it captures more domain-specific cues, our empirical studies showed that fine-tuning this layer, while the first remains frozen, often leads to overfitting owing to the limited size of the target dataset. In contrast, freezing both hidden layers stabilizes performance by retaining shared representations and reducing the influence of domain-specific noise. Additionally, freezing both layers reduces the number of trainable parameters, which is especially beneficial in low-cost applications. For the fine-tuning process, we freeze the corresponding parameters and fine-tune the P&F network with fewer target domain training samples (xt, yt). This enables the pretrained model to be adapted efficiently to the target domain tasks. For the two different target domains of Al-coated and PE-coated materials, the classification accuracy of the fine-tuned model on the target test dataset is presented in Table 4 (fine-tuning). Benefiting from the well-optimized parameters obtained during pre-training, the fine-tuned P&F model demonstrated excellent predictive performance across both tasks. The detailed training process is illustrated in Fig. 7.
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-F007.jpg)
It should be noted that, owing to the globally shared parameters of the neural network, the fine-tuning process aims to enhance the overall classification accuracy rather than optimize each class individually. As the model improves its discriminative ability for certain categories, the performance of others may deteriorate slightly, resulting in parameter competition and trade-offs. This phenomenon may cause a minor decrease in the prediction accuracy of specific classes. However, the overall classification performance of the target task improved. Moreover, the observation that the training loss exceeds the testing loss, with the training curve showing greater fluctuation while the test curve remains smooth, can be primarily attributed to the limited size of the training dataset. Because only 30% of the total data were used for training, the model became more sensitive to outliers and difficult samples, leading to a higher average loss and increased variance during optimization. In contrast, the larger testing set (70%) provides a more reliable and representative performance estimate, effectively averaging out the anomalies and resulting in a smoother and lower loss curve. Furthermore, because parameter updates occur solely on the training set, its loss is expected to fluctuate across batch iterations, whereas the test loss remains relatively stable.
Transfer learning with DANN model
The identification of the Z-class of an unknown coated material, as an unlabeled target domain problem, presents a more significant challenge. However, by identifying common scattering angle features between the coated material and its bare counterpart, we can train a neural network to achieve superior classification performance in an unknown domain. Traditional domain alignment methods typically compute specific mathematical relationships between the source and target domains and incorporate them into the training process as part of the loss function [45, 46]. Given that different transfer learning tasks exhibit distinct data characteristics, determining the optimal mathematical relationship as a training objective remains highly challenging.
The concept of adversaries in neural networks was first introduced in Ref. [47], where adversarial models generated adversarial samples to enhance model robustness. The DANN extends adversarial training to transfer learning by incorporating a domain discriminator that enforces feature distribution alignment between the source and target domains through adversarial training, thereby facilitating unsupervised learning in the target domain. This approach fully exploits the fitting capabilities of neural networks and enables effective feature alignment without requiring an explicit definition of the feature relationships between a specific source and target domains.
Our DANN model is illustrated in Fig. 8, consists of three main components: a feature extractor, a classifier, and a domain discriminator. During training, the labeled scattering angle data of the bare materials and unlabeled scattering angle data of the coated materials were both fed into the feature extractor. The total extracted features are then passed into the domain discriminator, and the extracted source domain features and corresponding labels serve as inputs to the classifier. First, the domain discriminator determines whether the input features originate from the source or target domain and simultaneously computes the loss function _2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-M008.png)
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-M009.png)
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-F008.jpg)
| Parameter | Value | Description |
|---|---|---|
| input_dimf | 500 | The number of nodes in the feature extractor input layer |
| hidden_dimf1 | 256 | The number of nodes in the feature extractor first hidden layer |
| hidden_dimf2 | 256 | The number of nodes in the feature extractor second hidden layer |
| hidden_dimc | 256 | The number of nodes in the classifier input layer |
| output_dimc | 3 | The number of nodes in the classifier output layer |
| hidden_dimd1 | 256 | The number of nodes in the discriminator input layer |
| hidden_dimd2 | 256 | The number of nodes in the discriminator hidden layer |
| output_dimd | 1 | The number of nodes in the discriminator output layer |
| tt_ratio | 7:3 | Ratio of samples in training set to test set |
| batch_size | 64 | The number of samples trained simultaneously in one iteration |
| epoch | 200 | The number of epochs in the training process |
| lr_f | 5×10-5 | The learning rate of the feature extractor in the training process |
| lr_c | 5×10-5 | The learning rate of the classifier in the training process |
| lr_d | 1×10-4 | The learning rate of the discriminator in the training process |
| grad_rev (λ) | 5 | Parameters for gradient reversal |
The training process and final results of DANN are presented in Fig. 9 and Table 6. The training results indicate that when trained with inverse CDF sampled data, the DANN exhibits slightly lower prediction accuracy for low-Z materials than for mid-Z and high-Z materials. This phenomenon is consistent with the results obtained when the pretrained model performs inference directly in the target domain, which can be attributed to the significant change in the scattering angle distribution of low-Z materials after being coated with Al, making feature alignment more challenging than that for mid-Z and high-Z materials. Meanwhile, the introduction of a large gradient reversal coefficient -λ causes the overall training loss to become negative because it includes the adversarial loss from the domain discriminator. However, this does not affect the evaluation of the test set, where the domain discriminator and gradient reversal are not involved. Therefore, the test loss, which was calculated solely based on the standard cross-entropy loss, remained positive. This behavior is expected and does not affect the assessment of the model performance on the target classification task.
_2026_05/1001-8042-2026-05-77/alternativeImage/1001-8042-2026-05-77-F009.jpg)
| Dataset | Sampling method | Z categories | Total | ||
|---|---|---|---|---|---|
| low-Z | mid-Z | high-Z | |||
| Al | RS | 0.941 | 0.941 | 0.980 | 0.954 |
| KDE | 0.861 | 0.991 | 0.998 | 0.950 | |
| HE | 0.917 | 0.982 | 0.996 | 0.965 | |
| PE | RS | 0.962 | 0.961 | 0.981 | 0.968 |
| KDE | 0.974 | 0.999 | 1.000 | 0.991 | |
| HE | 0.966 | 0.992 | 0.999 | 0.986 | |
Results and Discussion
The overall prediction accuracies of the pre-trained, fine-tuned, and DANN models in the target domain are summarized in Table 7. The training results indicate that the introduction of the inverse CDF sampling method effectively improved the sample quality, thereby enhancing the prediction accuracy. Owing to the randomness of the RS method and the inherent black-box nature of neural networks, it is difficult to provide a clear physical explanation for the variation in the training results in the source and target domains using RS methods. However, because the inverse CDF sampling method effectively captures the scattering angle data distribution, it offers stronger interpretability to the result. Additionally, as stated in Ref. [48], 1,400 scattering instances are sufficient to construct a statistically reliable muon scattering angle probability distribution. Although we computed the CDF at selected data points, our global interpolation acted on 500,000 instances, which is far beyond this threshold. Therefore, in practical applications, as long as sample diversity is maintained, the total amount of training data required can be reduced.
| Training method | Dataset | Sampling method | ||
|---|---|---|---|---|
| RS | KDE | HE | ||
| Pre-train | Al | 0.860 | 0.877 | 0.874 |
| PE | 0.909 | 0.980 | 0.988 | |
| Fine-tune | Al | 0.941 | 0.985 | 0.963 |
| PE | 0.968 | 0.996 | 0.988 | |
| DANN | Al | 0.954 | 0.950 | 0.965 |
| PE | 0.968 | 0.991 | 0.986 | |
For the two target domain tasks considered, because the PE coating has a minimal impact on the muon scattering angle distribution, the prediction accuracy of the model in the PE task only slightly improved (close to 100%) after transfer learning. However, for the Al task, transfer learning improved the prediction accuracy by approximately 10%. When comparing the two transfer learning methods, fine-tuning and DANN, the training results showed that the fine-tuned model achieved slightly higher prediction accuracy than the DANN model. As seen in the results summarized in Sect. 3, fine-tuning benefits from the supervised learning process, leading to a more balanced prediction accuracy across different Z-class materials. In contrast, although the DANN improves the prediction accuracy of low-Z materials by more than 20% compared with the pre-trained model without transfer learning, its accuracy remains slightly lower than that of fine-tuning owing to the unsupervised training in the target domain. However, this also highlights one of the key advantages of DANN. It does not require any label from the target domain; however, it achieves a prediction accuracy comparable to that of fine-tuning. Moreover, because the DANN-based transfer learning approach focuses on extracting shared features between the source and target domains, whereas fine-tuning involves learning the entire feature space of the source domain before transferring it to the target domain, it may capture irrelevant features that do not contribute to the target domain task. The training process indicated that the DANN achieved greater stability and robustness than fine-tuning. This makes DANN particularly valuable for real-world applications in which fully unsupervised transfer learning is often required.
The parameters of the two models are in the 105~106 magnitude (the P&F model contains 72,579 parameters, and the DANN model contains 323,332 parameters). In our implementation, the two transfer learning methods require a running memory on the order of 102~103 MB, and run in a few minutes on a local Intel Core i9-14900 K (24-core, 32-thread, up to 6.0 GHz) computer. The model architectures, training process, and data loading were implemented using the PyTorch libraries. The above analysis proves that our model and algorithm have relatively low deployment and training costs.
Conclusion
We developed a novel transfer learning method for identifying the material Z-class using muon scattering angle data, which is an alternative to traditional identification methods based on complex physical model reconstruction. First, Monte Carlo simulations were conducted using Geant4 to obtain scattering angle data for specified materials in both the bare and coated states. A series of fitting techniques, including computation and interpolation, were employed to derive the probability distribution of the material-scattering angles. Based on this distribution, we generated a sampled dataset that better conforms to the overall distribution, resulting in an approximately 4% improvement in prediction accuracy in the target domain and enhanced physical interpretability of the training process.
In real-world application scenarios, the correspondence between the scattering angle data and the coated material is often unknown. To address this challenge, we introduced two novel lightweight neural networks trained using transfer learning. By employing either fine-tuned supervised learning or adversarial unsupervised learning on the coated material, these models transfer the source-domain knowledge learned from the bare material data to the target domain of the coated materials. In the PE target domain task, where the scattering angle distribution remained largely unchanged before and after coating, the prediction accuracy reached 99%. In contrast, for the more challenging Al-coated task, the prediction accuracy improved by approximately 10% compared to the pre-transfer learning model. The results demonstrate that our method achieves high prediction accuracy, even when the mapping between the coated material and scattering data is scarce or completely unknown. We analyzed the results for different tasks and scenarios, verifying that Z-class identification based on machine learning aligns with the physical principles of muon interactions, validating the feasibility of Z-class prediction for coated materials via transfer learning. Furthermore, this study revealed that features learned from data through machine learning exhibit transferability rather than merely relying on the repeated application of domain-specific expertise across different scenarios. This suggests that machine learning methods based on transfer learning can serve as a cost-effective training approach for conducting physics research in similar situations.
In future research, incorporating additional physical variables, such as muon momentum beyond the scattering angle into the training data, is expected to further improve the accuracy and robustness of the Z classification. Real-world detection challenges include detector noise and electronic fluctuations that bias scattering angle measurements, resolution limitations that reduce the model’s ability to distinguish materials with similar scattering properties, and system inefficiencies, such as dead time and low detection efficiency, which degrade the overall data quality. We may introduce noise-aware training to boost the robustness and generalization of the model in real environments. Furthermore, by introducing more statistics and enhancing the generalization ability of the model, transfer learning methods are expected to be extended to Z-value identification tasks in more complex scenarios in the future.
Energy and angular distributions of atmospheric muons at the Earth
(2016). arXiv:1606.06907Cosmic-ray particles of inter- mediate mass
. Phys. Rev. 54, 88-89 (1938). https://doi.org/10.1103/PhysRev.54.88.2Review of particle physics: particle data group
. Phys. Rev. D. 98,A Comparison of Muon Flux Models at Sea Level for Muon Imaging and Low Background Experiments
. Front. Energy. Res. 9,Muon scattering tomography: review
. Appl. Opt. 61, C154-C161 (2022). https://doi.org/10.1364/AO.445806Muography
. Nat. Rev. Methods Primers 3, 88 (2023). https://doi.org/10.1038/s43586-023-00270-7Search for hidden chambers in the pyramids
. Science. 167, 3919 (1970). https://doi.org/10.1126/science.167.3919.832Analysis of spent nuclear fuel imaging using multiple coulomb scattering of cosmic muons
. IEEE. T. Nucl. Sci. 63, 2866-2874 (2016). https://doi.org/10.1109/TNS.2016.2618009Muography as a new complementary tool in monitoring volcanic hazard: implications for early warning systems
. Proc. R. Soc. A. 477, 2255 (2021). https://doi.org/10.1098/rspa.2021.0320Cosmic muon flux measurement and tunnel overburden structure imaging
. JINST. 15,Imaging internal density structure of the Laoheishan volcanic cone with cosmic ray muon radiography
. Nucl. Sci. Tech. 33, 88 (2022). https://doi.org/10.1007/s41365-022-01072-4Radiographic imaging with cosmic-ray muons
. Nature. 422, 277 (2003). https://doi.org/10.1038/422277aDetection of high-Z objects using multiple scattering of cosmic ray muons
. Rev. Sci. Instrum. 74, 4294–4297 (2003). https://doi.org/10.1063/1.1606536Image reconstruction and material Z discrimination via cosmic ray muon radiography
. Nucl. Instrum. Meth. A. 519, 687-694 (2004). https://doi.org/10.1016/j.nima.2003.11.035Fundamental limitations of dual energy X-ray scanners for cargo content atomic number discrimination
. Appl. Radiat Isotopes. 206,A modified multi-group model of angular and momentum distribution of cosmic ray muons for thickness measurement and material discrimination of slabs
. Nucl. Sci. Tech. 29, 28 (2018). https://doi.org/10.1007/s41365-018-0363-7A novel 4D resolution imaging method for low and medium atomic number objects at the centimeter scale by coincidence detection technique of cosmic-ray muon and its secondary particles
. Nucl. Sci. Tech. 33, 2 (2022). https://doi.org/10.1007/s41365-022-00989-0A new efficient imaging reconstruction method for muon scattering tomography
. Nucl. Instrum. Meth. A. 1069,Nuclear mass based on the multi-task learning neural network method
. Nucl. Sci. Tech. 33, 48 (2022). https://doi.org/10.1007/s41365-022-01031-zBeam based alignment using a neural network
. Nucl. Sci. Tech. 35, 75 (2024). https://doi.org/10.1007/s41365-024-01436-yA review of deep learning methods for compressed sensing image reconstruction and its medical applications
. Electronics. 11(4), 586 (2022). https://doi.org/10.3390/electronics11040586A comparative study of sparse recovery and compressed sensing algorithms with application to AoA estimation
.A Newton-type Forward Backward Greedy method for multi-snapshot compressed sensing
.Machine learning at the energy and intensity frontiers of particle physics
. Nature. 560, 41 (2018). https://doi.org/10.1038/s41586-018-0361-2Machine learning in the search for new fundamental physics
. Nat. Rev. Phys. 4, 399–412 (2022). https://doi.org/10.1038/s42254-022-00455-1Colloquium: Machine learning in nuclear physics
. Rev. Mod. Phys. 94,Convolutional neural network algorithm for material discrimination in muon scattering tomography
. Atom. Energy. Sci. Techno. 57, 353-361 (2023). https://doi.org/10.7538/yzk.2022.youxian.0055Material discrimination using cosmic ray muon scattering tomography with an artificial neural network
. Radiat. Detect. Technol. Methods. 6, 254–261 (2022). https://doi.org/10.1007/s41605-022-00319-3A Survey on Transfer Learning
. IEEE. T. Knoml. Data. En. 22, 1345 (2010). https://doi.org/10.1109/TKDE.2009.191A Comprehensive Survey on Transfer Learning
. in Proceedings of the IEEE. 109, 43 (2021). https://doi.org/10.1109/JPROC.2020.3004555Survey on Transfer Learning Research
. J. Softw. 26, 26-39 (2015). https://doi.org/10.13328/j.cnki.jos.004631Statistical Reconstruction for Cosmic Ray Muon Tomography
. IEEE. T. Image Process. 16, 1985-1993 (2007). https://doi.org/10.1109/TIP.2007.901239Characterising encapsulated nuclear waste using cosmic-ray Muon Tomography (MT)
.Cosmic Ray Muon Radiography Applications in Safeguards and Arms Control
(2018). arXiv:1808.06681Geant4—a simulation toolkit
. Nucl. Instrum. Meth. A. 506, 250-303 (2003). https://doi.org/10.1016/S0168-9002(03)01368-8Cosmic-ray shower generator (CRY) for Monte Carlo transport codes
.The unreasonable effectiveness of data
. IEEE. INTELL. SYST. 24, 8-12 (2009). https://doi.org/10.1109/MIS.2009.36ImageNet: A large-scale hierarchical image database
.A tutorial on kernel density estimation and recent advances
. Biostatistics & Epidemiology, 1, 161–187(2017). https://doi.org/10.1080/24709360.2017.1396742A Kernel Two-Sample Test
. J. MACH. LEARN. RES. 13, 723-773 (2012). http://jmlr.org/papers/v13/gretton12a.htmlGenerative adversarial networks
. arXiv:1406.2661 (2014) https://doi.org/10.48550/arXiv.1406.2661Experimental study on material discrimination based on muon discrete energy
. Acta Phys. Sin. 72,The authors declare that they have no competing interests.

