Page 29 - Read Online
P. 29

Page 22 of 34                                                Wang et al. Energy Mater. 2026, 6, 600064















































               Figure 5. Machine learning model prediction and experimental photovoltaic performance of perovskite solar cells. (A) The fitting graph of
               PCE results by the CatBoost-based algorithm model, where red represents the training set and blue represents the test set. Reprinted with
               permission [162] . Copyright 2024, John Wiley and Sons. (B) The distribution of experimental results and data points from the database in the
               fitting graph. Reprinted with permission [162] . Copyright 2024, John Wiley and Sons. (C) J-V curves of devices with different perovskite
               components via a two-step spin-spin sequential deposition method. Reprinted with permission [162] . Copyright 2024, John Wiley and Sons.
               (D) Statistics of the PCE of PSCs with different groups based on 7 devices. The central line represents the median, the box limits
               correspond to the upper and lower quartiles, and the whiskers extend to the minimum and maximum values. Reprinted with permission [162] .
               Copyright 2024, John Wiley and Sons. (E) J-V curves of the regular devices and their champion photovoltaic performance. Reprinted with
               permission [162] . Copyright 2024, John Wiley and Sons. (F) J-V curves of the inverted devices and their champion photovoltaic performance.
               Reprinted with permission [162] . Copyright 2024, John Wiley and Sons.


               most effective safeguard against overfitting remains rigorous out-of-sample evaluations. To validate the
               generalization capability of this methodology, the team conducted a blind test on 17 newly designed
               photovoltaic molecules. The predictive results demonstrated strong precision, achieving a mean absolute
               error (MAE) of only 1.02% across the test group. Notably, the PCE of 10 novel molecular candidates were
               predicted within a 1% error margin. This validation confirms that the model has captured the underlying
               structure-property relationships rather than merely interpolating within known historical data, thereby
               enabling rapid in silico evaluation of unexplored chemical spaces and providing a transparent and rational
               blueprint for the design of future functional materials.


               However, the most pervasive challenge in ML-based models arises when the available experimental datasets
               are extremely limited. To address this small-sample bottleneck, the Co-Pilot for Perovskite Additive Screener
               (Co-PAS) framework was developed around a fundamental shift in data partitioning and feature
               representation. Pu et al. [165]  abandoned traditional data augmentation strategies and instead focused on the
               physical logic of data partitioning and the transferability of dimensionality reduction techniques. To address
               the severe performance distortion and high variance caused by the random splitting of small datasets, they
   24   25   26   27   28   29   30   31   32   33   34