Page 30 - Read Online
P. 30
Wang et al. Energy Mater. 2026, 6, 600064 Page 23 of 34
introduced a Molecular Scaffold Classifier (MSC). This methodology categorizes the limited pool of known
additives into distinct scaffold groups based on their underlying topological structures, utilizing these
classifications as the basis for data segmentation. This diversity-driven partitioning compels the algorithm to
learn the governing physicochemical principles of all core scaffolds during training, thereby reducing the risk
of memorizing specific local features and improving generalization to entirely novel molecular architectures.
Furthermore, to address the limited feature representation caused by insufficient target data, the framework
deployed a transfer learning strategy using a Junction Tree Variational Autoencoder (JTVAE). Rather than
training this autoencoder from scratch on the small target dataset, the algorithm was pre-trained on a
massive external database comprising millions of organic compounds from ZINC. Through this approach,
the framework translates complex additive structures into continuous latent vectors that encode substantial
chemical prior knowledge. This molecular representation overcomes the limitations of traditional manual
descriptors, granting the predictive model strong molecular analytical capabilities despite the limited volume
of available device data. This integrated strategy combining physical scaffold constraints with pre-trained
features demonstrated strong predictive accuracy during high-throughput screening. As the most rigorous
experimental validation, the researchers applied this framework to evaluate a chemical space of 250,000
unknown molecules, successfully isolating a previously unreported additive named Boc-L-threonine
N-hydroxysuccinimide ester (BTN) . Subsequent device fabrication confirmed that PSCs incorporating the
[165]
BTN passivation strategy achieved a PCE of 25.20%. By successfully navigating extreme data scarcity, this
approach establishes a transferable methodological blueprint for future ML-guided additive engineering
under conditions of limited experimental data [Figure 6].
Representative ML-driven optimization strategies for PSCs are summarized in Table 3. In summary, the
integration of ML into perovskite additive engineering has demonstrated a clear evolutionary trajectory,
advancing from data-driven additive analysis to high-throughput screening and ultimately toward predictive
additive design. However, despite these algorithmic advancements, the majority of current predictive
frameworks optimize PCE as the sole target property. This singular focus overlooks the broader
requirements for commercialization, notably long-term operational stability and manufacturing feasibility. A
key barrier to incorporating these critical factors into computational models is the severe lack of
standardization across the research community. Experimental protocols for assessing device stability vary
considerably in duration, temperature, and atmospheric exposure, while scalability metrics are complicated
by diverse substrate architectures and vastly different active area dimensions. Future methodologies must
prioritize unifying these disparate testing conditions to construct multidimensional target matrices, enabling
algorithms to design additives that meet commercialization requirements rather than merely achieving
isolated laboratory efficiency records. Furthermore, progress is also hindered by the fundamental limitations
of current dataset scale, quality, and bias. Because current predictive models predominantly rely on
aggregated historical literature, they are affected by survivorship bias, learning primarily from successful
devices while overlooking the insights hidden within negative results. To move beyond these finite and
biased historical datasets, the field must pivot toward autonomous learning paradigms.
FUTURE OUTLOOK
Future progress in PSCs will increasingly depend on the synergistic integration of ML and additive
engineering. Rather than serving merely as post-hoc analytical tools, ML frameworks are poised to evolve
into prescriptive design platforms capable of generating actionable molecular blueprints. The rapid
advancement of deep generative models, including variational autoencoders, normalizing flows, and
language model-based molecular design pipelines, enables exploration of chemical spaces that extend well
beyond conventional molecular libraries [170] . By integrating such rigorous physical descriptors as formation
energies for point and extended defects, solvation dynamics, and band alignment targeting the valence and
conduction band edges, these models can deliver both high predictive accuracy and mechanistic

