Page 84 - Read Online
P. 84

Page 6 of 23                          Li et al. J. Mater. Inf. 2025, 5, 43  https://dx.doi.org/10.20517/jmi.2025.17

               Table 1. Overview of common benchmark datasets for ML-IAPs
                Dataset  Description          Data scale      Benchmark tasks    URL
                QM9 [24]  Stable small organic molecules   134 k molecules   Molecular property   https://figshare.com/collections/
                                                   6
                        (C, H, O, N, F; ≤ 9 heavy atoms)  (~1 × 10  atoms)  prediction (energies,   Quantum_chemistry_structures_and_
                                                              HOMO/LUMO, dipoles)  properties_of_134_kilo_molecules/
                                                                                 978904
                   [25]
                MD17    MD trajectories for 8 small   ~3-4 M configurations   Energy and force prediction http://quantum-machine.org/
                                                   8
                        organic molecules (e.g., benzene,  (~1 × 10  atoms)      datasets/#md-datasets
                        ethanol)
                MD22 [26]  Large molecules/biomolecular   0.2 M configurations (~1  Energy and force prediction  https://www.openqdc.io/datasets/
                                                 7
                        fragments (42-370 atoms)  × 10  atoms)  for large systems  md22
                ANI-1 [27]  Small organic molecules [C, H, N,  20 M DFT   Molecular potential   https://figshare.com/collections/_/
                        O; (CNO) ≤ 8 atoms]   conformations   training; energy prediction  3846712
                                                   6
                                              (~5 × 10  atoms)
                ANI-1x [28]  Small organic molecules (C, H, N,  5 M DFT conformations  Molecular potential   DOI: 10.6084/m9.figshare.c.4712477.
                                                   6
                        O)                    (~5 × 10  atoms)  training; energy prediction  v1
                ANI-1ccx [29]  Subset of ANI-1x recalculated at   0.5 M high-accuracy   High-precision energy   DOI: 10.6084/m9.figshare.c.4712477.
                        CCSD(T)/CBS level     conformations   prediction         v1
                                                   5
                                              (~5 × 10  atoms)
                    [30]
                ANI-2x  Organic molecules including S, F,  9,651,712 conformers   Energy and force prediction  https://zenodo.org/records/10108942
                                                    6
                        Cl (H, C, N, O, S, F, Cl)  (~9.7 × 10  atoms)  across extended chemistry
                ISO17 [31]  C H O  isomer dynamics  0.645 M configurations  Energy/force generalization  DOI: 10.1038/sdata.2018.75
                           10
                             2
                         7
                                                   7
                                              (~1 × 10  atoms)  over chemical and
                                                              conformational changes
                   [32]
                SPICE   Bio-relevant molecules and   1.1 M conformers   Energy/force prediction for  DOI: 10.5281/zenodo.7338495
                                                   8
                        complexes (15 elements)  (~1 × 10  atoms)  biomolecular interactions
                   [33]
                OC20    Adsorbate-surface combinations  1.28 M DFT relaxations  Adsorption energy, reaction  https://fair-chem.github.io/catalysts/
                                                     8
                        on transition-metal facets  (~2.65 × 10  single-point  pathway, structure   datasets/oc20.html
                                              evaluations)    optimization
                   [34]
                OC22    Oxide electrocatalyst surfaces   62,331 DFT relaxations   S2EF, IS2RE, IS2RS for oxide  https://fair-chem.github.io/catalysts/
                        and adsorbates        (~9.85 M single-point   catalysts  datasets/oc22.html
                                              calculations)
                Materials   Bulk inorganic crystals (all   130 k optimized crystal   Formation energy, band   https://materialsproject.org/
                    [35]
                Project  element combinations)  structures    gap, elastic moduli
                Matbench [36]  13 derived property tasks from   312-132 k samples per   Multi-property materials   https://github.com/hackingmaterials/
                                                      5
                        Materials Project     task (~3 × 10  total)  predictions (energy, gap,   matbench
                                                              moduli)
               The listed datasets are widely used public databases that span small organic molecules, large biomolecular fragments, gas-surface adsorbate
               systems and bulk inorganic crystals. For each database it reports the scale, typical benchmark tasks, and a URL for download or access. ML-IAPs:
               Machine learning interatomic potentials; MD: molecular dynamics; DFT: density functional theory.
               Materials Data Facility , the Materials Project, OQMD, NOMAD and 2DMatPedia now host extensive
                                   [44]
               collections of crystal structures, simulation outputs and derived properties. Leveraging these repositories
               enables the assembly of diversified, high-quality training sets, streamlines data curation and facilitates
               benchmarking against standardized datasets. Active learning monitoring model uncertainty and selectively
               incorporating high-uncertainty samples into the training set has proven effective in exploring chemical
               space more efficiently, reducing labeling effort and improving generalizability in materials modeling [45,46] .
               Likewise, pretraining on lower-fidelity structural databases offers a promising route to broaden coverage
                                               [47]
               with manageable computational costs .
               Defining the chemical and structural breadth of a training dataset establishes its “target space”, and different
               coverage regimes give rise to distinct modeling strategies ranging from general purpose zero shot
               frameworks to fine-tuned and bespoke potentials. General purpose approaches, however, face two intrinsic
               limitations. First, existing materials databases do not uniformly sample the periodic table, resulting in
               pronounced element wise imbalances and variable data quality. Second, these models often struggle to
               represent or relax complex architectures such as reconstructed surfaces, defect networks or low symmetry
   79   80   81   82   83   84   85   86   87   88   89