Page 89 - Read Online
P. 89
Wang et al. J. Mater. Inf. 2026, 6, 1 Page 5 of 9
solely from elemental attributes. Composition-based descriptors generally fall into two categories: (1) hand-
crafted descriptors derived from isolated elemental properties, such as electronegativity, first ionization
energy, atomic number [26,27] . Mathematical operations can be applied to these features to capture variations;
(2) learned embeddings obtained from pre-trained models (e.g., ElemNet , CrabNet ) that encode
[28]
[29]
chemical compositions information. These representations can then be fine-tuned on task-specific datasets
to improve performance and adaptability to the problem at hand.
Both approaches exhibit distinct advantages and limitations. The first method benefits from its minimal
reliance on extensive feature parameters, as it primarily leverages domain expertise. Furthermore, the
features are inherently interpretable. This transparency is particularly valuable for electronic structure-based
insights. However, the manually designed features may lack comprehensiveness, failing to capture complex
relationships. In contrast, the second method employs features derived from pre-trained models. These
descriptors enable superior generalization to unexplored regions of chemical space [30,31] . It was also found that
this method has at least 30% lower error in formation energy prediction than the first method, even though it
takes 3-5 times longer to train . Thus, in practical applications, the choice between these descriptors should
[28]
be guided by critical factors such as dataset scale, elemental diversity, and the trade-off between
interpretability and predictive power.
The next step involves leveraging AI to predict key materials properties, such as band gap and bulk modulus,
based on combinations of elements at different crystallographic sites . In the design of garnet-type
[32]
electrolytes [Figure 2B], for example, researchers employed the T f to preliminary assess the geometric
stability of 29,008 candidate compositions , narrowing the pool to 7,067 viable structures. During
[11]
validation, 47 out of 49 (95.92%) predicted garnets were successfully optimized, demonstrating the
robustness of the T f-driven workflow. Similarly, in the discovery of quaternary materials , T f helped narrow
[13]
down an initial dataset of 2,700 compounds to 380 promising candidates [Figure 2C]. Subsequent
hierarchical clustering identified materials with superior optoelectronics performance, where Ag 2BaTiSe 4 was
experimentally confirmed and achieved a power conversion efficiency of 20.08% in solar cell applications .
[33]
We note that rigorous AIMD simulations or experimental synthesis remain essential for the practical
deployment of any newly identified materials.
ONGOING CHALLENGES AND POTENTIAL SOLUTIONS
There are still many challenges to improve the versatility and accuracy of T f. In this section, we propose three
key challenges and corresponding solutions [Figure 3] to fully realize the important role of T f in materials
discovery.
(1) Failure in materials systems with dopes and defects: Traditional T f struggles to capture the effects of
doping and defects, such as vacancies, substitutions, or interstitials, which can introduce local distortion and
alter ionic radii. Substitutional doping may shift lattice parameters gradually, while intrinsic defects often
cause more severe local distortions. These changes disrupt the fixed coordination assumptions of T f, leading
to unreliable stability assessments.
Solution: One solution to this challenge is to build an active learning model to mine the structural
information of multiple components. Taking perovskite (ABX 3) as an example, by collecting and establishing
multi-component data (A 1-xA’ xB 1-yB’ yX 3) using open-source databases (e.g., Materials Cloud , Automatic-
[34]
Flow for Materials Discovery , and Novel Materials Discovery ), a weighted T f is constructed according to
[36]
[35]
the chemical ratio of the multiple components, and then a stability prediction model is constructed. In the
unknown chemical space, high-precision DFT calculations are performed based on the uncertainty of the
prediction results and fed back to the data set. This approach strikes a balance between efficiency and
accuracy: instead of exhaustively running DFT for all configurations, active learning (AL) reduces the
number of required DFT calculations, while achieving comparable or even better performance.

