Page 80 - Read Online
P. 80
Page 2 of 23 Li et al. J. Mater. Inf. 2025, 5, 43 https://dx.doi.org/10.20517/jmi.2025.17
science, offering rigorous control over thermodynamic variables alongside atomistic spatial and temporal
resolution. In DFT, the electronic structure is obtained by solving the Kohn-Sham self-consistent field
(SCF) equations and diagonalizing the Hamiltonian matrix to extract its eigenvalues. By contrast, both
geometry optimization and MD rely on potentials: geometry optimization locates minima to identify stable
atomic configurations, whereas MD integrates Newton’s equations of motion to simulate the real-time
evolution of atomic positions and velocities. Despite their widespread use, these techniques face inherent
limitations. The cost of DFT scales as O(N ) (or worse) with the number of atoms N due chiefly to
3
[1]
Hamiltonian diagonalization, thereby constraining studies to relatively small quantum systems .
Conversely, classical MD, though orders of magnitude faster, depends on empirical interatomic potentials
(IAPs) [or force fields (FFs)] that often lack the transferability and accuracy required for complex
chemistries.
Bridging the gap between accuracy and scalability has emerged as a central challenge. Machine learning
(ML) offers a transformative pathway by leveraging high-fidelity ab initio data to construct surrogate
models that operate efficiently at extended scales. ML interatomic potentials (ML-IAPs), or ML force fields
(ML-FFs), implicitly encode electronic effects through training on quantum reference datasets, enabling
faithful recreation of the potential energy surface (PES) across diverse chemical environments without
explicitly propagating electronic degrees of freedom. Their robustness hinges on accurately learning the
mapping from atomic coordinates to energies and forces. In parallel, ML Hamiltonian (ML-Ham)
approaches seek to predict electronic potentials using methods such as ML-derived Kohn–Sham
[2]
[4]
potentials , deep Hamiltonian neural networks (DHNNs) , Hamiltonian graph neural networks (GNNs)
[3]
and deep tight-binding models . Different from conventional ML approaches (structure-property), ML-
[5]
Ham methods (structure-physics-property) provide clearer physical pictures and explainability, delivering
near-ab initio accuracy for quantities ranging from band structures and Berry phases to electron–phonon
couplings.
In this Review, we survey recent algorithmic and architectural advances in ML-driven IAPs and
Hamiltonian models, with particular emphasis on symmetry-aware GNNs, data-efficient training strategies
and interpretability techniques and their successful applications. Moreover, we highlight the critical
challenges in this field, including data fidelity, model generalizability and computational scalability. Finally,
we also outline promising future directions poised to extend ML-accelerated simulations from small
molecules to complex, multiscale materials systems.
ML-IAPS
ML-IAPs or MLFFs have emerged as a transformative approach in computational materials science, offering
[6,7]
a data-driven alternative to traditional empirical FFs . By leveraging deep neural network architectures,
ML-IAPs directly learn the PES from extensive, high quality quantum mechanical datasets , thereby
[8]
obviating the need for fixed functional forms, such as conventional Lennard-Jones or bond-order potentials,
and instead optimizing large parameter spaces via automatic differentiation .
[7]
The principal advantage of ML-IAPs lies in their capacity to reproduce atomic interactions, including
energies, forces and dynamical trajectories, with high fidelity across chemically diverse systems . When
[8]
trained on ab initio molecular dynamics (AIMD) trajectories, these models facilitate accurate simulations
over extended temporal and spatial scales , achieving superior accuracy relative to conventional potentials
[9]
while maintaining the computational efficiency required for large-scale materials modeling .
[10]

