Page 27 - Read Online
P. 27
Page 20 of 34 Wang et al. Energy Mater. 2026, 6, 600064
forecast the energy-level shifting capabilities and passivating efficacy of candidate molecules , accelerating
[160]
the discovery of optimal passivators and energy alignment agents. Ultimately, this convergence of additive
engineering and data-driven methods not only enhances molecular design precision but also establishes a
predictive paradigm for interfacial energy management.
MACHINE LEARNING FOR ADDITIVE ENGINEERING IN PEROVSKITE SOLAR CELLS
Data-driven additive retrospection
As underscored throughout the preceding discussion, the role of additives in PSCs is fundamentally
multidimensional, where a single molecular intervention often simultaneously influences defect density,
energy-level alignment, and long-term environmental stability. This intricate interconnectedness between
chemical structure and device physics implies that selecting a high-performing additive is no longer a simple
task of addressing an isolated vulnerability but rather a complex multi-objective optimization process. The
massive compositional and processing parameter space created by these multifunctional additives renders
conventional trial-and-error methodologies increasingly inefficient. Consequently, the high dimensionality
and nonlinear relationships between molecular features and device metrics make additive engineering not
only highly amenable to ML approaches but also dependent on their application. Early efforts focused on
uncovering empirical correlations from large experimental datasets. Odabaşı et al. demonstrated through
[161]
analysis of 1,921 PSCs that the implementation of a ternary dopant mixture comprising lithium
bis(trifluoromethylsulfonyl) imide salt (LiTFSI), 4-tert-butylpyridine (TBP), and
tris(2-(1H-pyrazol-1yl)-4-tert-butylpyridine) cobalt(III) tris-(bis(trifluoromethylsulfonyl)imide)) (FK209)
significantly enhanced device performance, yielding a lift ratio of 2.76 for achieving stabilized PCE exceeding
18%. Quantitatively, while photovoltaic devices employing this specific dopant formulation constituted
merely 8% of the entire experimental dataset, they accounted for 21% of all top-tier high-efficiency devices.
In addition to evaluating specific HTL dopants, Odabaşı et al. [161] systematically assessed the impact of
various fabrication parameters, including mixed-cation perovskite compositions, solvent engineering
strategies (e.g., dimethylformamide and dimethyl sulfoxide mixtures (DMF + DMSO)), antisolvent
treatments (e.g., chlorobenzene (CB)), and diverse ETL architectures incorporating tin oxide. This analysis
provides robust support for optimization practices that have historically relied on empirical trial-and-error
methodologies, while simultaneously laying the groundwork for a paradigm shift from heuristic screening to
predictive molecular design. Despite the scale and statistical rigor of this meta-analysis, the selected
parameter space exhibits several critical omissions that limit its predictive capacity for future additive
development. First, continuous processing variables such as precursor concentrations, spin-coating
parameters, and annealing conditions were excluded due to inherent inconsistencies in reporting standards
across different laboratories. Furthermore, the ML models employed PCE as the sole target output, thereby
neglecting long-term operational stability and scalability metrics, which currently represent the most
pressing bottleneck in perovskite commercialization. Most critically, these algorithms were trained on
discrete categorical labels rather than intrinsic physicochemical properties, such as dipole moments, binding
affinities, and steric effects. Consequently, the resulting models are inherently limited to optimizing existing
formulations rather than predicting entirely novel multifunctional additives, highlighting the need for more
sophisticated frameworks that integrate continuous processing parameters, multidimensional stability
assessments, and quantum chemical molecular descriptors.
Data-driven additive screening
To establish a more robust data-driven foundation, Wu et al. extracted 63 experimentally validated data
[41]
points from approximately 26 different molecular additives reported in the literature to construct their
training dataset. Utilizing the RDKit library, they extracted 14 molecular descriptors (including molecular
weight, complexity, oxygen atom count, and hydrogen bond acceptor count) and comparatively evaluated
five ML models, namely linear regression (LR), random forest (RF), gradient boosting (GB), extreme

