Page 71 - Read Online
P. 71
Page 16 of 21 Wen et al. J. Mater. Inf. 2025, 5, 30 https://dx.doi.org/10.20517/jmi.2024.102
Figure 8. Comparison of the calculated DFT values for (A) Hole reorganization energy, (B) Solvation free energy, (C) Maximum
absorption, and (D) LogP of common HTMs with the predicted values from RF, GBDT and XGBoost ML models, respectively. DFT:
Density functional theory; HTMs: hole transport materials; RF: random forest; GBDT: gradient boosted decision tree; XGBoost: extreme
gradient boosting; ML: machine learning.
training dataset by incorporating samples with distinct geometric structures (e.g., helical, star-shaped) to
improve the model’s capacity for learning across a wide range of molecular configurations. Additionally,
applying structural transformation simulations to the existing data could generate diverse molecular
datasets, further enriching the distribution of the training dataset. Alternatively, employing universal
molecular descriptors or advanced feature extraction methods (such as neural networks or ensemble
methods) could enable the model to better capture the fundamental characteristics of various molecular
structures. Neural networks offer significant advantages in handling nonlinear molecular structures and
complex features, enabling better capture of intricate relationships between molecules and thereby
improving prediction accuracy. Ensemble methods, on the other hand, enhance prediction robustness by
combining the outputs of multiple models, effectively reducing model bias and increasing the stability and

