Page 28 - Read Online
P. 28

Wang et al. Energy Mater. 2026, 6, 600064                                        Page 21 of 34





               gradient boosting (XGBoost), and support vector regression (SVR). The RF model demonstrated superior
               performance with a test set R  = 0.741. Subsequently, they screened 141 candidate molecules from the
                                         2
               PubChem database based on their predicted PCE enhancement ratios. Three representative molecules were
               selected, namely 5-carboxyphthalide (5-CP, predicted ratio = 1.218 ± 0.05) as a high-performance candidate,
               4-(hydroxymethyl)benzoic acid (4-HMBA, predicted ratio = 1.101 ± 0.05) for moderate enhancement, and
               ethanethiol (EtSH, predicted ratio = 0.939 ± 0.05) as a potential detrimental additive. Experimental validation
               revealed that devices modified with 5-CP achieved a PCE of 21.4%, representing a significant enhancement
               compared to the 18.4% PCE of control devices, while EtSH led to a performance decline to 17.8% .
                                                                                                        [41]
               Although stability testing confirmed that 5-CP-modified devices exhibited superior stability relative to
               control samples, this observation was incidental rather than a direct prediction output of the initial ML
               model. Since the random forest algorithm was trained exclusively on efficiency enhancement ratios, the
               screening process may overlook candidate additives that provide deeper synergistic effects between
               photovoltaic performance and structural resilience. This limitation underscores the need for future models to
               incorporate multiple performance metrics to identify additives with broader benefits.


               Data-driven additive prediction
               Starting with a summary and preliminary screening of known additives, the ultimate objective of ML is to
               achieve precise prediction of device performance and directly guide experimental discovery. Yang et al. [162]
               pioneered this transition by constructing a robust predictive framework based on a large-scale historical
               dataset of 2,079 experimentally fabricated solar cells. Rather than analyzing parameters in isolation, their
               algorithm simultaneously evaluates interconnected descriptors, using experimental compositions as input
               features (rather than material physicochemical properties or characterizations), including precise
               stoichiometric ratios of chemical components and varying concentrations of specific additives, to forecast the
               final PCE. To validate the accuracy of the predictions, 12 independent devices were fabricated, yielding an
               average absolute error of merely 1.6% between the predicted and experimentally measured efficiencies, as
               illustrated in Figure 5.

               While this impressively low error rate demonstrates strong predictive reliability, it is important to
               acknowledge that real-world fabrication is strongly influenced by the surrounding environment. To address
               the inherent risk of models overfitting to idealized laboratory conditions, Liu et al. [163]  expanded their
               predictive feature matrix to encompass dynamic ambient variables such as relative humidity and processing
               atmosphere. Rather than treating these environmental factors as noise, a RF algorithm was deployed to
               capture the nonlinear interplay between precursor chemistry and atmospheric conditions, directly
               forecasting the resultant PCE. By incorporating atmospheric metadata into the computational framework,
               this methodology transforms ML from an isolated molecular screening tool into a broader predictive
               instrument capable of addressing the complexities of industrial-scale perovskite fabrication.

               To further ensure that such computational frameworks capture generalizable physical laws rather than
               merely memorizing training data noise, recent methodologies have integrated interpretable frameworks with
               advanced molecular modeling. Liao et al. [164]  leveraged theoretically computable, effective, and reusable
               molecular descriptors rather than relying on coarse structural classifications. Their approach incorporated
               three distinct descriptor sets, including the ground-state property descriptor set (MDS-GS, containing 23
               descriptors), the absorption spectrum property descriptor set (MDS-ABS, containing 19 descriptors), and the
               electron transfer property descriptor set (MDS-ET, containing 5 descriptors). Furthermore, by integrating
               Shapley Additive exPlanations (SHAP) analysis, the researchers decoded the algorithmic black box, filtering
               out redundant variables and identifying the most critical physical descriptors governing device efficiency.
               This targeted feature selection compels the algorithms to anchor their predictions on fundamental chemical
               principles rather than spurious statistical correlations, thereby establishing a reliable predictive baseline. The
   23   24   25   26   27   28   29   30   31   32   33