Page 83 - Read Online
P. 83

Li et al. J. Mater. Inf. 2025, 5, 43  https://dx.doi.org/10.20517/jmi.2025.17    Page 5 of 23

                                                        [21]
               Deep potential molecular dynamics (DeePMD)  formulates the total potential energy as a sum of atomic
               contributions, each represented by a fully nonlinear function of local - environment descriptors defined
               within a prescribed cutoff radius. The DeePMD framework, implemented in DeePMD-kit, has been trained
                                                     6
               on extensive DFT datasets of the order of 10  water configurations, achieving energy mean absolute errors
               (MAEs) below 1 meV per atom and force MAE under 20 meV/Å. By encoding smooth neighboring density
               functions to characterize atomic surroundings and mapping these descriptors through deep neural
               networks, Deep Potentials attain quantum mechanical accuracy with computational efficiency comparable
               to classical MD, thereby enabling atomistic simulations at spatiotemporal scales hitherto inaccessible.

               Data representation: balance between quality and quantity
               Notwithstanding these advancements, the predictive accuracy of even SOTA ML models remains
               fundamentally limited by the breadth and fidelity of available training data. Publicly accessible experimental
               materials datasets are orders of magnitude smaller than those in image or language domains, impeding the
               construction of universally transferable and highly precise potentials. DFT datasets with meta-generalized
               gradient approximation (meta-GGA) exchange-correlation functionals offer markedly improved
                                                              [22]
               generalizability compared to semi-local approximations , and thus provide a solid foundation for training
               universal models  [Table 1].
                             [23]
                                                                                                       [37]
               However, the majority of current repositories remain at Perdew–Burke–Ernzerhof (PBE) level accuracy ,
               highlighting the need to incorporate more sophisticated electronic treatments (for example, Hubbard U
               corrections or meta-GGA/hybrid functionals) to capture complex many-body interactions. Recent work
               exemplifies this strategy: the high-fidelity data–based M3GNN framework leverages meta-GGA datasets
               alongside an SE(3)-equivariant GNN to resolve subtle structural and electronic features, establishing a new
                                                             [38]
               benchmark for materials-property prediction accuracy  [Figure 1].
               Model efficiency is equally critical for scaling ML-IAPs to large-system simulations. A linearized NequIP
               architecture  reduces the complexity of tensor contractions while preserving core equivariant operations,
                         [39]
               yielding substantial decreases in inference cost with no measurable loss in force-prediction accuracy. This
               demonstration confirms that judicious architectural simplifications can reconcile high fidelity with practical
               throughput. In a complementary advance, the Meta ML-IAP framework  couples comprehensive data
                                                                               [40]
               curation pipelines with a modular network design to enhance both generalizability and computational
               performance. By leveraging curated high-fidelity datasets alongside streamlined equivariant layers, this
               approach extends ML-driven potentials to multicomponent and structurally intricate materials systems
               without sacrificing efficiency.

               Although MD simulations yield atomistic trajectories, the resulting datasets remain constrained by both
                                           [41]
               label noise and limited sampling . Classical or semi-empirical potentials introduce systematic errors in
               energy and force labels, while the high cost of MD restricts simulations to nanosecond-microsecond
               timescales and nanometre-micrometre lengthscales, impeding the observation of rare events and long-range
               phase transitions. Furthermore, trajectories typically originate from a single initial configuration, yielding
                                                                       [42]
               uneven phase-space coverage and restricted structural diversity . To overcome these challenges, data
               quality may be enhanced by adopting higher-fidelity potential models or experimental calibration, and
               dataset size expanded via enhanced sampling techniques, multiscale simulations or integration with
               experimental measurements, thereby producing ML-IAPs that are both robust and generalized.

               Effective data-sharing initiatives are critical for accelerating materials discovery and model development.
                                                                               [43]
               The findable, accessible, interoperable and reusable (FAIR) framework , and platforms such as the
   78   79   80   81   82   83   84   85   86   87   88