Page 30 - Read Online
P. 30
Liu et al. J. Mater. Inf. 2025, 5, 27 I http://dx.doi.org/10.20517/jmi.2024.105 Page 7 of 18
Table 1. A brief overview of the FT DP downstream dataset used in training the model
2
Type of structures 3D bulks 2D surfaces 1D strings 0D clusters All
Numbers of structures with Fe 6,917 13,829 117 37 20,900
Numbers of structures without Fe 7,341 1,114 330 971 9,756
Total numbers 14,258 14,943 447 1,008 30,656
2
FT DP: Fine-tuned Fischer-Tropsch deep potential; 3D: three-dimensional; 2D: two-
dimensional; 1D: one-dimensional; 0D: zero-dimensional.
the bottom layers remain fixed in structural relaxations. Additionally, Δ Fe or Δ C represents the differences
in the number of Fe or C atoms between the reconstructed structure and the clean surface, respectively.
For convenience, we used the electronic energy of an isolated carbon atom ( C) as the reference for the carbon
chemical potential, that is Δ C = C − C. Since the free energies and chemical potentials are relevant to
temperature, pressure, and gas atmosphere. Here we used the results from Liu et al. to simulate a realistic
iron-based FTS condition ( = 523 K, -6.60 eV ≤ Δ C ≤ -7.45 eV) [24] .
RESULTS AND DISCUSSION
FT DP construction and validation
2
Our model, FT DP, is constructed through fine-tuning on our downstream dataset from the upstream DPA-
2
2.2.0 model [68] , which is a pre-trained open LAM (OpenLAM) from the AIS Square website [69] . Thanks to
the multi-task training protocol, this LAM was trained on more than 20 different datasets containing various
physical and chemical systems including organic molecules, clusters, alloys, semiconductors, surfaces, and ad-
sorbates through multi-task training mechanism, gathering multidisciplinary knowledge in one unified DPA-2
descriptor. Apart from the descriptor, the fitting model is a neural network containing three hidden layers with
the typical numbers of neurons being (240, 240, 240) for all heads in the upstream DPA-2.2.0 model and our
fine-tuned model.
There are 30,656 frames in our FT DP downstream dataset, including Fe-C-H-O element combinations and
2
various types of structures, derived from the previous work by Liu etal. [25] , and approximately 8,000 structures
were removed after the data cleaning procedures below to remove outliers and redundancies: (1) Removal of
structures with identical DFT-calculated energy labels to eliminate redundant conformations; (2) Removal of
the structures with fewer than 12 atoms per cell, which often represent isolated molecules or radicals in a cell.
Such configurations are prone to DFT inaccuracies in PBCs or poor MLP generalizability; (3) Exclusion of
structures having the absolute difference between model prediction and DFT results exceeds 80.0 meV/atom
(energy) or 1.00 eV/Å (maximum atomic forces), where the model referenced here was fine-tuned from the
upstream DPA-2 model on the original datasets following the same fine-tuning protocol detailed in the next
paragraph. All DFT energies and forces were calculated by ABACUS following the computational settings
mentioned above. A brief overview of this dataset is given in Table 1, showing the number of structures (with
or without Fe) in different types, including three-dimensional (3D) bulks, two-dimensional (2D) surfaces, one-
dimensional (1D) strings, and zero-dimensional (0D) clusters. Besides, a sketch-map visualization is shown
in Figure 2 for illustrating the wide configuration distribution of our FT DP dataset.
2
Our fine-tuning protocol was initialized by using the parameters of the global descriptor in pre-trained DPA-
2.2.0 model and fitting network from the Domains_OC2M branch. The fine-tuning process on our dataset is
done by following the default training process of the DPA-2 model with some setting modifications. In the
default pre-training process of DPA-2.2.0 LAM, the learning rate starts from 2 × 10 and gradually decreases
−4
to 3.51×10 by an exponential decreasing scheme with each decrease performed at every 1/200 checkpoint of
−8
the total training step. The setting is usually effective for from-scratch training process, but the initial learning

