Page 69 - Read Online
P. 69
Page 6 of 15 Wu et al. J. Mater. Inf. 2025, 5, 14 https://dx.doi.org/10.20517/jmi.2024.77
Figure 2. Enumeration of unique CO Adsorption Configurations. A three-step method for defining unique CO adsorption configurations
on Cu surfaces. Initially, Step 1 identifies 68 distinct adsorption sites using graph theory to capture surface atom arrangements. Step 2
demonstrates the systematic filling of CO on these sites while adhering to spatial constraints, yielding approximately 44 million
preliminary configurations. Step 3 applies symmetry operations to distill these down to around 7 million unique configurations,
streamlining the dataset for further computational exploration.
[Supplementary Table 2]. Compared to the graph-theoretical deduplication algorithm, this method not only
has a crushing advantage in comparison speed but also can theoretically overcome the limitations of the
graph isomorphism algorithm in comparing periodic crystal structures . Overall, this enumeration process
[25]
of the configuration space allows for a more comprehensive consideration of the adsorbate sequence space,
thereby reducing errors associated with energetics-based configuration sampling [25,29] .
Machine-learning force field
Given the vast adsorption configuration space spanning eight different Cu surfaces, it is imperative to
leverage artificial intelligence to navigate this extensive configuration space. Two research approaches are
considered viable: The first involves constructing MLFFs through the rich trajectory data obtained during
the first-principles geometric optimization of a small set of adsorption configurations, followed by using the
trained force field model to optimize the remaining adsorption configurations . The second approach
[30]
entails developing a deep learning model capable of directly predicting the steady-state adsorption energy
from initial configuration guesses [18,49] . Given that the construction of force fields can significantly utilize
data from the structural optimization process, the initial DFT calculations required for the first strategy are
considerably less. However, using the force field model to optimize the remaining configurations is also a
time-consuming task, especially considering the configuration space volume is close to seven million.
Additionally, selecting a small number of representative configurations for DFT calculations from a vast and
unevenly distributed sample space, particularly among high-index surface configurations, poses a
challenging problem.

