Page 231 - Read Online
P. 231

Yuan et al. J. Mater. Inf. 2026, 6, 17                                            Page 9 of 13





               Using the trained weights, a single prediction was performed, and the results shown in Figure 5B
               demonstrate clear and accurate identification of the corresponding molecular structures. The predicted
               patterns align well with the molecular structures in Figure 5A, indicating that the model effectively captures
               and retains the features of all six categories without significant performance degradation. To further quantify
               this performance, the evaluation metrics indicate a precision of 0.959, recall of 0.931, mAP@0.5 of 0.976, and
               mAP@0.5:0.95 of 0.792, confirming the model’s robustness across classes. Specifically, the per-class
               mAP@0.5:0.95 values for M1 through M6 are 0.698, 0.728, 0.685, 0.794, 0.904, and 0.945, respectively. These
               results highlight the model’s strong generalization capability and its ability to maintain high accuracy and
               localization precision across diverse molecular categories, even after multiple stages of incremental updates.


               These findings further highlight the strength of the incremental learning strategy, which enables the model to
               continuously learn new molecular categories while retaining knowledge of previously learned ones. The
               predictions indicate that the model can handle the complexity of molecular assemblies and accurately detect
               and classify molecules, thus illustrating its robustness and generalization capability after incremental
               learning. The clear patterns and accurate bounding boxes in the predictions reflect the model’s ability to
               integrate diverse molecular features, making it a powerful tool for analyzing high-resolution images of
               complex molecular arrangements.


               Batch effects arising from different STM imaging environments, such as variations in tunneling current
               stability, tip geometry, or detector sensitivity, may introduce distributional bias into the image data. To
               mitigate this issue, the dataset includes STM images collected under diverse experimental conditions on
               Au(111), Ag(111), and Cu(111) substrates. By incrementally introducing data from these different imaging
               environments while replaying earlier samples, the model continuously calibrates itself across multiple
               acquisition domains, thereby reducing the risk of bias. In addition to the experiments discussed above, the
               model’s performance was further evaluated on images of the same molecular species acquired on three
               different substrates, namely Au(111), Ag(111), and Cu(111) [Figure 6]. The results show that the incremental
               learning strategy not only allows the model to continuously adapt to new data but also exhibits strong
               generalization across diverse imaging conditions. Despite variations in surface structures, noise levels, and
               contrast among different substrates, the model successfully identified the target molecule with no noticeable
               degradation in detection accuracy. This outcome indicates that the replay mechanism effectively preserves
               the essential features of the learned molecule, enabling robust recognition across heterogeneous datasets. The
               ability to adapt to different substrates highlights the strong transferability and generalization capability of the
               incremental learning framework. In STM imaging, the substrate type plays a critical role in determining the
               overall image characteristics. For example, Au(111), Ag(111), and Cu(111) surfaces differ in lattice constants
               and electronic structures, leading to significant variations in background textures, contrast levels, and noise
               patterns in the resulting images. Conventional deep learning detection models are often sensitive to such
               differences; when trained exclusively on data from a single substrate, their performance typically declines
               significantly when tested on another . To overcome this issue, the aforementioned incremental learning
                                              [32]
               strategy combined with a replay mechanism was adopted.


               In our experiments, we progressively introduced STM images of the target molecule obtained on Au(111),
               Ag(111), and Cu(111). At the initial stage, the model was trained solely on Au(111) data and achieved high
               detection accuracy. We then incrementally incorporated images from Ag(111) and Cu(111) substrates while
               simultaneously replaying Au(111) images to reinforce the retention of previously acquired knowledge. This
               strategy ensured that the model maintained strong recognition performance for Au(111) molecules while
               adapting to the unique imaging characteristics of Ag(111) and Cu(111). The final evaluation results
               demonstrate that the model maintains consistently high accuracy across all three substrates (with only minor
               variations in mAP@0.5), thereby validating the effectiveness of the incremental learning approach.
   226   227   228   229   230   231   232   233   234   235   236