Page 72 - Read Online
P. 72
Wen et al. J. Mater. Inf. 2025, 5, 30 https://dx.doi.org/10.20517/jmi.2024.102 Page 17 of 21
reliability of property predictions for SM-HTMs. These strategies aim to improve the model’s generalization
capabilities across different molecular architectures and enhance the accuracy of its performance
predictions. On the other hand, as shown in Supplementary Table 18, the RF model and XGBoost model
exhibited the best performance in predicting hydrophobicity and hole reorganization energy, respectively.
On the other hand, the GBDT model performed better in predicting solvation free energy and maximum
light absorption peak. Overall, the RF, GBDT, and XGBoost models, trained on the database obtained from
high-throughput calculations, demonstrate good generalizability in predicting the material properties of
linear organic SM-HTMs.
Methods for the design and development of new SM-HTMs
The discussion above, exemplified by linear SM-HTMs, presents a novel strategy for the design and
development of SM-HTMs. The corresponding workflow is schematically illustrated in Figure 9. This
methodology outlines a systematic and iterative approach designed to expedite the discovery of high-
performance SM-HTMs. The process begins with the construction of a diverse library of candidate organic
small molecules using a MSA. High-throughput computational methods are then applied to
comprehensively evaluate the performance parameters of these molecules. At this stage, high-performing
candidates can be screened and selected for further study. Additionally, the computational data generated
can be used to train ML models, enabling the development of robust structure-property relationship
models. This facilitates the direct prediction of performance parameters based on molecular structures
generated by the splicing algorithm, significantly reducing dependence on resource-intensive simulations. If
the model encounters new molecular structures that lead to inaccuracies in prediction, high-throughput
computational methods can be reintroduced to optimize and refine the ML model. Moreover, the ML
model can be employed for inverse design, allowing the identification of molecular structures predicted to
exhibit superior performance. This integrated strategy effectively combines the computational efficiency of
ML with the precision of high-throughput calculations, fostering iterative improvement in both predictive
accuracy and material discovery. This seamless and adaptive workflow significantly accelerates the
identification, screening, and optimization of next-generation SM-HTMs. While this study primarily
focuses on the application of the proposed methodology to the design and screening of SM-HTMs for PSCs,
the approach can be extended to other photovoltaic materials and device architectures. The MSA, combined
with high-throughput computational screening and ML, provides a versatile framework that can be adapted
for the discovery of new functional materials in various optoelectronic applications. For instance, in organic
photovoltaics (OPVs), donor-acceptor molecular systems play a crucial role in determining device
performance. By applying the MSA, high-throughput computational screening and ML strategies, a diverse
library of donor-acceptor molecules can be systematically generated and screened based on key parameters
such as frontier molecular orbital energies, exciton binding energy, and charge transport properties.
Similarly, in dye-sensitized solar cells (DSSCs), this methodology can facilitate the identification of novel
organic dyes with enhanced light absorption, redox stability, and efficient charge transfer properties. It
offers a powerful framework for advancing material innovation in applications such as PSCs and other
cutting-edge photovoltaic technologies.
CONCLUSIONS
In summary, this work presents a novel design strategy for key HTMs in PSCs, utilizing the combination of
molecular splicing, high-throughput computational screening, and ML techniques to identify candidate
molecular materials with outstanding structures, comprehensive properties, and synthetic feasibility.
Approximately 200,000 π-type molecular structures were generated using a MSA, from which 7,399
molecules were selected for D-π-D-type molecular construction, followed by high-precision DFT
calculations. Ultimately, a database of 7,222 D-π-D HTMs was compiled, containing property data for
molecular structure models, HOMO levels, hole reorganization energy, solvation free energy, maximum

