Page 27 - Read Online
P. 27
Page 4 of 18 Liu et al. J. Mater. Inf. 2025, 5, 27 I http://dx.doi.org/10.20517/jmi.2024.105
section “Results and Discussions”, we illustrate the construction of FT DP, and its application of the Fe-FTS at
2
the DFT precision level with efficiency, including exploring the reaction mechanism of Fe-FTS reactions and
revealing the morphology of reconstructed FeC surfaces with steps. The last section summarizes the main
findings of this work and remarks on the possible further developments and applications of the FT DP model
2
in the future.
MATERIALS AND METHODS
DFT calculations
All spin-polarized DFT calculations were carried out by using Atomic-orbital Based Ab-initio Computation
at UStc (ABACUS) package [44,45] . The SG15-optimized Norm-Conserving Vanderbilt (SG15-ONCV) multi-
projectorpseudopotentials [46,47] wereemployedandthevalenceconfigurationswere[H]1s ,[C]2s 2p ,[O]2s 2p 4
2
1
2
2
and [Fe]3d 4s . The generalized gradient approximation (GGA) in the Perdew-Burke-Ernzerhof (PBE) vari-
2
6
ant [48] was adopted for the exchange-correlation functional. The second generation of numerical atomic or-
bitals (NAOs)inthedouble- pluspolarizationfunction(DZP)form [49] wasused asthebasisset. Theperiodic
boundary condition (PBC) and the Γ-centered Monkhorst–Pack scheme [50] for sampling the Brillouin zone
were adopted in the DFT calculations, with an automated mesh determined by k-spacing = 0.14 Bohr and
−1
only one k-point for the direction without and with vacuum layers, respectively. The dipole correction perpen-
diculartothesurfacewasappliedforallDFTcalculationsofsurfaces. Theelectrondensitycriterionforelectron
self-consistency convergence was set at 1×10 , and the first-order Methfessel-Paxton (MP) smearing [51] was
−7
used for the occupation of orbitals. In geometry and transition state (TS) optimizations, the convergence
criterion for the largest force among all atoms was set to 0.05 eV/Å.
DPA-2 and fine-tuning methodology
DPA-2 is a multi-task pre-trained LAM originating from the DP architecture and evolved from the DPA-1
model [52] . The DPA-1 descriptor has introduced an element-type embedding for encoding the elemental
information covering the whole periodic table, and a gated self-attention mechanism [53] excelling in mod-
eling the importance of neighboring atoms and re-weighting the interaction among them, which also makes
the model generalizable and pre-trainable. Inheriting the DPA-1 backbone, the DPA-2 descriptor further en-
hances its resolution and generalizability of atomic representation through stacking multiple transformer [53]
layers called representation transformer, incorporating operators such as convolution, symmetrization, local-
ized self-attention, and gated self-attention, which can be interpreted as an E(3) equivariant graph neural net-
work (GNN) and offers greater capacity compared to conventional GNNs [42] , ensuring the robust capability
of DPA-2 for serving as a LAM assembling comprehensive knowledge from massive pre-training data.
Besides having a more sophisticated model architecture, DPA-2 employs a multi-task training strategy for pre-
training in multiple datasets labeled with different DFT settings to extract multidisciplinary knowledge. In
particular, the multi-task DPA-2 model has multiple heads, and each head is an identical fitting network used
to fit the DFT labels of each pre-training dataset from different downstream domains. During the pre-training
process, the parameters within the DPA-2 descriptor are concurrently optimized through back-propagation
using all pre-training datasets, while the parameters of the fitting network are updated exclusively with the
specific pre-training dataset to which they are associated [42] .
The pre-trained descriptor and fitting networks can be fine-tuned on specific downstream tasks, and the mul-
tidisciplinary knowledge learned from the multiple upstream datasets can help to reduce the consumption in
model training and the amount of training data. In the fine-tuning process, the descriptor of the downstream
model will be initialized with the pre-trained parameters, and the fitting network could also be initialized by
choosing a fitting network in the pre-trained model. The energy bias in the fitting network will be aligned
to the labels of the downstream dataset subsequently, and then the typical model training process will pro-

