Page 70 - Read Online
P. 70

Wu et al. J. Mater. Inf. 2025, 5, 14  https://dx.doi.org/10.20517/jmi.2024.77    Page 7 of 15

               Therefore, after considering both the predictive accuracy of MLFFs and the prediction speed of deep
               learning models, we designed an integrated solution that uses MLFF as a data augmenter for the adsorption
               energy prediction model. Initially, for low-index surfaces such as Cu(100), Cu(110), and Cu(111) depicted
               in Figure 1B, where the enumerated configuration count is low, we randomly select 20% of the unoptimized
               adsorption configurations for DFT calculations. For higher-index surfaces, we perform a primary sampling
               based on adsorption site combination types, followed by a secondary sampling based on simplified
               adsorption site combinations. The configurations obtained from secondary sampling undergo DFT
               structural optimization. This sampling process yielded a total of 1,592 structures. The resulting structure-
               energy pairs through structural optimization were obtained for training the MLFF [Supplementary Table 5].
               Subsequently, we train the MLFF for the Cu-CO system with accuracy close to DFT calculations using the
               open-source and user-friendly deep learning architecture DPMD, which has been proven to have high
               simulation accuracy and efficiency in vast Si atomic systems [32,50] . This force field model is then applied to
               optimize the remaining low-index surface configurations and those from the primary sampling. By
               expanding the training dataset with adsorption configurations using MLFFs, we can significantly address
               the issue of data scarcity faced when constructing deep learning models, thereby substantially improving
               model quality .
                           [31]

               Our trained force field exhibits high accuracy, with root mean square error (RMSE) values for energy and
               force being < 1 meV/atom [Figure 3A] and < 0.03 eV/Å [Figure 3B], respectively. The optimization of initial
               configuration guesses in the secondary sampling space using the well-fitted force field further assesses the
               force field quality: the steady-state adsorption configurations obtained from force field optimization almost
               perfectly match those obtained from DFT optimization [Figure 3C]. Moreover, analyzing the distribution of
               Cu–C bond lengths in the steady-state adsorption configuration set reveals that the distribution from MLFF
               optimization closely aligns with that from DFT optimization [Figure 3D], with an average Cu–C bond
               length discrepancy of < 0.01 Å. This demonstrates that the MLFF can achieve near-DFT calculation
               accuracy for adsorption configurations across eight Cu surfaces. The high-quality force field is attributed
               not only to the superior MLFF architecture but also to the effective sampling method - ensuring the
               sampled configuration guesses, i.e., the starting points for DFT geometric optimization, are as diverse as
               possible, facilitating a rich optimization trajectory that benefits model learning of more complex multi-body
               interactions. We used the MFLL to optimize ~186,000 structures for the following graph neural network
               (GNN) model and the distribution of structures was presented in Supplementary Table 6. The specific
               training parameters of MLFF are shown in Supplementary Table 7.


               Graph-based model for adsorption energy prediction
               In recent years, GNN models have shone brightly in material design and performance prediction, leveraging
               graph embeddings of materials for various predictive tasks [51-54] . One of their key advantages over other deep
               learning models lies in the natural suitability of graph structures for describing chemical and material
               structures, with nodes and edges representing atoms and their interactions, respectively. Hence, extracting
               high-quality graph data from materials is crucial. In our study, the interactions among CO molecules affect
               their adsorption energy on Cu surfaces, especially at high coverages. Previous research often overlooked
               non-bonding interactions beyond hydrogen bonds when constructing graph data, failing to accurately
                                             [18]
               characterize adsorption structures . We propose an adsorption configuration graph data extraction
               method that effectively captures the impact of CO interactions on their adsorption energy, as shown in
               Figure 4A and Supplementary Figure 2. The first step involves searching for atoms in contact with a CO
               molecule on the Cu surface, considering it as the central CO, within the van der Waals radius to form a
               neighbor atom set, including first-order neighbor COs in contact. The second step continues the search for
               atoms in contact with these first-order neighbor COs and builds local subgraphs for multiple neighbor sets
               centered on CO. The third step merges the central CO and its neighboring COs’ local subgraphs into a
   65   66   67   68   69   70   71   72   73   74   75