Page 30 - Read Online
P. 30

Wang et al. Energy Mater. 2026, 6, 600064                                        Page 23 of 34





               introduced a Molecular Scaffold Classifier (MSC). This methodology categorizes the limited pool of known
               additives into distinct scaffold groups based on their underlying topological structures, utilizing these
               classifications as the basis for data segmentation. This diversity-driven partitioning compels the algorithm to
               learn the governing physicochemical principles of all core scaffolds during training, thereby reducing the risk
               of memorizing specific local features and improving generalization to entirely novel molecular architectures.
               Furthermore, to address the limited feature representation caused by insufficient target data, the framework
               deployed a transfer learning strategy using a Junction Tree Variational Autoencoder (JTVAE). Rather than
               training this autoencoder from scratch on the small target dataset, the algorithm was pre-trained on a
               massive external database comprising millions of organic compounds from ZINC. Through this approach,
               the framework translates complex additive structures into continuous latent vectors that encode substantial
               chemical prior knowledge. This molecular representation overcomes the limitations of traditional manual
               descriptors, granting the predictive model strong molecular analytical capabilities despite the limited volume
               of available device data. This integrated strategy combining physical scaffold constraints with pre-trained
               features demonstrated strong predictive accuracy during high-throughput screening. As the most rigorous
               experimental validation, the researchers applied this framework to evaluate a chemical space of 250,000
               unknown molecules, successfully isolating a previously unreported additive named Boc-L-threonine
               N-hydroxysuccinimide ester (BTN) . Subsequent device fabrication confirmed that PSCs incorporating the
                                             [165]
               BTN passivation strategy achieved a PCE of 25.20%. By successfully navigating extreme data scarcity, this
               approach establishes a transferable methodological blueprint for future ML-guided additive engineering
               under conditions of limited experimental data [Figure 6].


               Representative ML-driven optimization strategies for PSCs are summarized in Table 3. In summary, the
               integration of ML into perovskite additive engineering has demonstrated a clear evolutionary trajectory,
               advancing from data-driven additive analysis to high-throughput screening and ultimately toward predictive
               additive design. However, despite these algorithmic advancements, the majority of current predictive
               frameworks optimize PCE as the sole target property. This singular focus overlooks the broader
               requirements for commercialization, notably long-term operational stability and manufacturing feasibility. A
               key barrier to incorporating these critical factors into computational models is the severe lack of
               standardization across the research community. Experimental protocols for assessing device stability vary
               considerably in duration, temperature, and atmospheric exposure, while scalability metrics are complicated
               by diverse substrate architectures and vastly different active area dimensions. Future methodologies must
               prioritize unifying these disparate testing conditions to construct multidimensional target matrices, enabling
               algorithms to design additives that meet commercialization requirements rather than merely achieving
               isolated laboratory efficiency records. Furthermore, progress is also hindered by the fundamental limitations
               of current dataset scale, quality, and bias. Because current predictive models predominantly rely on
               aggregated historical literature, they are affected by survivorship bias, learning primarily from successful
               devices while overlooking the insights hidden within negative results. To move beyond these finite and
               biased historical datasets, the field must pivot toward autonomous learning paradigms.


               FUTURE OUTLOOK
               Future progress in PSCs will increasingly depend on the synergistic integration of ML and additive
               engineering. Rather than serving merely as post-hoc analytical tools, ML frameworks are poised to evolve
               into prescriptive design platforms capable of generating actionable molecular blueprints. The rapid
               advancement of deep generative models, including variational autoencoders, normalizing flows, and
               language model-based molecular design pipelines, enables exploration of chemical spaces that extend well
               beyond conventional molecular libraries [170] . By integrating such rigorous physical descriptors as formation
               energies for point and extended defects, solvation dynamics, and band alignment targeting the valence and
               conduction band edges, these models can deliver both high predictive accuracy and mechanistic
   25   26   27   28   29   30   31   32   33   34   35