Page 28 - Read Online
P. 28
Wang et al. Energy Mater. 2026, 6, 600064 Page 21 of 34
gradient boosting (XGBoost), and support vector regression (SVR). The RF model demonstrated superior
performance with a test set R = 0.741. Subsequently, they screened 141 candidate molecules from the
2
PubChem database based on their predicted PCE enhancement ratios. Three representative molecules were
selected, namely 5-carboxyphthalide (5-CP, predicted ratio = 1.218 ± 0.05) as a high-performance candidate,
4-(hydroxymethyl)benzoic acid (4-HMBA, predicted ratio = 1.101 ± 0.05) for moderate enhancement, and
ethanethiol (EtSH, predicted ratio = 0.939 ± 0.05) as a potential detrimental additive. Experimental validation
revealed that devices modified with 5-CP achieved a PCE of 21.4%, representing a significant enhancement
compared to the 18.4% PCE of control devices, while EtSH led to a performance decline to 17.8% .
[41]
Although stability testing confirmed that 5-CP-modified devices exhibited superior stability relative to
control samples, this observation was incidental rather than a direct prediction output of the initial ML
model. Since the random forest algorithm was trained exclusively on efficiency enhancement ratios, the
screening process may overlook candidate additives that provide deeper synergistic effects between
photovoltaic performance and structural resilience. This limitation underscores the need for future models to
incorporate multiple performance metrics to identify additives with broader benefits.
Data-driven additive prediction
Starting with a summary and preliminary screening of known additives, the ultimate objective of ML is to
achieve precise prediction of device performance and directly guide experimental discovery. Yang et al. [162]
pioneered this transition by constructing a robust predictive framework based on a large-scale historical
dataset of 2,079 experimentally fabricated solar cells. Rather than analyzing parameters in isolation, their
algorithm simultaneously evaluates interconnected descriptors, using experimental compositions as input
features (rather than material physicochemical properties or characterizations), including precise
stoichiometric ratios of chemical components and varying concentrations of specific additives, to forecast the
final PCE. To validate the accuracy of the predictions, 12 independent devices were fabricated, yielding an
average absolute error of merely 1.6% between the predicted and experimentally measured efficiencies, as
illustrated in Figure 5.
While this impressively low error rate demonstrates strong predictive reliability, it is important to
acknowledge that real-world fabrication is strongly influenced by the surrounding environment. To address
the inherent risk of models overfitting to idealized laboratory conditions, Liu et al. [163] expanded their
predictive feature matrix to encompass dynamic ambient variables such as relative humidity and processing
atmosphere. Rather than treating these environmental factors as noise, a RF algorithm was deployed to
capture the nonlinear interplay between precursor chemistry and atmospheric conditions, directly
forecasting the resultant PCE. By incorporating atmospheric metadata into the computational framework,
this methodology transforms ML from an isolated molecular screening tool into a broader predictive
instrument capable of addressing the complexities of industrial-scale perovskite fabrication.
To further ensure that such computational frameworks capture generalizable physical laws rather than
merely memorizing training data noise, recent methodologies have integrated interpretable frameworks with
advanced molecular modeling. Liao et al. [164] leveraged theoretically computable, effective, and reusable
molecular descriptors rather than relying on coarse structural classifications. Their approach incorporated
three distinct descriptor sets, including the ground-state property descriptor set (MDS-GS, containing 23
descriptors), the absorption spectrum property descriptor set (MDS-ABS, containing 19 descriptors), and the
electron transfer property descriptor set (MDS-ET, containing 5 descriptors). Furthermore, by integrating
Shapley Additive exPlanations (SHAP) analysis, the researchers decoded the algorithmic black box, filtering
out redundant variables and identifying the most critical physical descriptors governing device efficiency.
This targeted feature selection compels the algorithms to anchor their predictions on fundamental chemical
principles rather than spurious statistical correlations, thereby establishing a reliable predictive baseline. The

