Page 72 - Read Online
P. 72
Wu et al. J. Mater. Inf. 2025, 5, 14 https://dx.doi.org/10.20517/jmi.2024.77 Page 9 of 15
Figure 4. GNN model for CO Adsorption Prediction. (A) The graph data extraction method for CO adsorption configurations, detailing
the steps from detecting neighboring atoms under van der Waals conditions to merging local subgraphs into a feature graph that
accurately represents the adsorption environment of CO molecules on Cu surfaces. (D) The comparison of the GNN model’s predictive
performance between using only the DFT calculation dataset corresponding to (B) and using the DFT + MLFF calculation dataset
corresponding to (C). (E) The computational efficiency gains of the DFT + MLFF + GNN workflow compared to traditional DFT and DFT
+ MLFF methods, showcasing a significant reduction in computational cost and time, thus enabling the study of vast adsorption
configuration spaces with enhanced efficiency. DFT: Density functional theory; MLFF: machine-learning force field; GNN: graph neural
network; RMSE: root mean square error; GEN: graph embedding network model; MAE: mean absolute error; MAPE: mean absolute
percentage error.
[Figure 4D]. Moreover, compared to direct DFT calculations or combined DFT + MLFF approaches for
exploring target configuration spaces, our DFT + MLFF + GNN methodology significantly reduces
computational costs by three and one orders of magnitude, respectively, greatly enhancing research
efficiency in vast adsorption configuration spaces [Figure 4E]. This acceleration allows for exploring
extensive configuration spaces within a foreseeable short period, a capability previously unattainable [25,29] .
This dual-speed framework, integrating high-precision MLFFs with advanced graph representation
learning, is also applicable to other catalytic systems with large search spaces, such as catalytic reaction path
searches, stable adsorbate motif determination and protein-ligand structure prediction [55-57] . The primary
objective is to identify the global minimum adsorption configurations and their corresponding energies
across varying surface coverages. Such critical information is unlikely to be fully captured within the
smaller, randomly sampled set of ~186,000 configurations. In contrast, the comprehensive dataset of
~7 million configurations exhaustively enumerates nearly all possible adsorption states, thereby ensuring
that the most stable adsorption configurations are accurately identified.
The generality of the proposed framework primarily manifests in two stages. Firstly, in the process of
independent adsorption configuration enumeration, the method is applicable to all high-coverage
adsorption configurations involving monodentate intermediates, such as H , N , O , and OH . For specific
*
*
*
*
adsorbates, only minor adjustments are required. However, for configurations involving multidentate
intermediates, our method still needs further improvement. Secondly, in the process of adsorption energy
prediction, our approach is expected to be easily extended to metal surfaces with a fcc lattice. However, for
metal surfaces with other lattice structures, some adjustments may be necessary based on the corresponding
system.

