Page 84 - Read Online
P. 84
Page 6 of 23 Li et al. J. Mater. Inf. 2025, 5, 43 https://dx.doi.org/10.20517/jmi.2025.17
Table 1. Overview of common benchmark datasets for ML-IAPs
Dataset Description Data scale Benchmark tasks URL
QM9 [24] Stable small organic molecules 134 k molecules Molecular property https://figshare.com/collections/
6
(C, H, O, N, F; ≤ 9 heavy atoms) (~1 × 10 atoms) prediction (energies, Quantum_chemistry_structures_and_
HOMO/LUMO, dipoles) properties_of_134_kilo_molecules/
978904
[25]
MD17 MD trajectories for 8 small ~3-4 M configurations Energy and force prediction http://quantum-machine.org/
8
organic molecules (e.g., benzene, (~1 × 10 atoms) datasets/#md-datasets
ethanol)
MD22 [26] Large molecules/biomolecular 0.2 M configurations (~1 Energy and force prediction https://www.openqdc.io/datasets/
7
fragments (42-370 atoms) × 10 atoms) for large systems md22
ANI-1 [27] Small organic molecules [C, H, N, 20 M DFT Molecular potential https://figshare.com/collections/_/
O; (CNO) ≤ 8 atoms] conformations training; energy prediction 3846712
6
(~5 × 10 atoms)
ANI-1x [28] Small organic molecules (C, H, N, 5 M DFT conformations Molecular potential DOI: 10.6084/m9.figshare.c.4712477.
6
O) (~5 × 10 atoms) training; energy prediction v1
ANI-1ccx [29] Subset of ANI-1x recalculated at 0.5 M high-accuracy High-precision energy DOI: 10.6084/m9.figshare.c.4712477.
CCSD(T)/CBS level conformations prediction v1
5
(~5 × 10 atoms)
[30]
ANI-2x Organic molecules including S, F, 9,651,712 conformers Energy and force prediction https://zenodo.org/records/10108942
6
Cl (H, C, N, O, S, F, Cl) (~9.7 × 10 atoms) across extended chemistry
ISO17 [31] C H O isomer dynamics 0.645 M configurations Energy/force generalization DOI: 10.1038/sdata.2018.75
10
2
7
7
(~1 × 10 atoms) over chemical and
conformational changes
[32]
SPICE Bio-relevant molecules and 1.1 M conformers Energy/force prediction for DOI: 10.5281/zenodo.7338495
8
complexes (15 elements) (~1 × 10 atoms) biomolecular interactions
[33]
OC20 Adsorbate-surface combinations 1.28 M DFT relaxations Adsorption energy, reaction https://fair-chem.github.io/catalysts/
8
on transition-metal facets (~2.65 × 10 single-point pathway, structure datasets/oc20.html
evaluations) optimization
[34]
OC22 Oxide electrocatalyst surfaces 62,331 DFT relaxations S2EF, IS2RE, IS2RS for oxide https://fair-chem.github.io/catalysts/
and adsorbates (~9.85 M single-point catalysts datasets/oc22.html
calculations)
Materials Bulk inorganic crystals (all 130 k optimized crystal Formation energy, band https://materialsproject.org/
[35]
Project element combinations) structures gap, elastic moduli
Matbench [36] 13 derived property tasks from 312-132 k samples per Multi-property materials https://github.com/hackingmaterials/
5
Materials Project task (~3 × 10 total) predictions (energy, gap, matbench
moduli)
The listed datasets are widely used public databases that span small organic molecules, large biomolecular fragments, gas-surface adsorbate
systems and bulk inorganic crystals. For each database it reports the scale, typical benchmark tasks, and a URL for download or access. ML-IAPs:
Machine learning interatomic potentials; MD: molecular dynamics; DFT: density functional theory.
Materials Data Facility , the Materials Project, OQMD, NOMAD and 2DMatPedia now host extensive
[44]
collections of crystal structures, simulation outputs and derived properties. Leveraging these repositories
enables the assembly of diversified, high-quality training sets, streamlines data curation and facilitates
benchmarking against standardized datasets. Active learning monitoring model uncertainty and selectively
incorporating high-uncertainty samples into the training set has proven effective in exploring chemical
space more efficiently, reducing labeling effort and improving generalizability in materials modeling [45,46] .
Likewise, pretraining on lower-fidelity structural databases offers a promising route to broaden coverage
[47]
with manageable computational costs .
Defining the chemical and structural breadth of a training dataset establishes its “target space”, and different
coverage regimes give rise to distinct modeling strategies ranging from general purpose zero shot
frameworks to fine-tuned and bespoke potentials. General purpose approaches, however, face two intrinsic
limitations. First, existing materials databases do not uniformly sample the periodic table, resulting in
pronounced element wise imbalances and variable data quality. Second, these models often struggle to
represent or relax complex architectures such as reconstructed surfaces, defect networks or low symmetry

