Page 23 - Read Online
P. 23
Page 16 of 33 Liu et al. J Mater Inf 2024;4:33 https://dx.doi.org/10.20517/jmi.2024.48
semiconductor properties and photocatalytic performance, facilitating a generalized performance prediction
across various photocatalysts. Deep learning approaches leverage these diverse features to extract patterns
from large datasets, aiding in identifying and optimizing the most promising candidate and enhancing the
efficiency of photocatalyst design and application.
Model construction and training
The construction and training of deep learning models for photocatalyst design involve defining the model
architecture, selecting relevant descriptors as input features, and optimizing the model parameters. Typical
deep learning models with their features and applications are listed in Table 6. Traditional deep learning
models, including deep neural networks (DNNs) , are widely used for classification and regression in
[115]
[116]
predicting photocatalytic properties. Convolutional neural networks (CNNs) analyze spatial data, aiding
in surface and structural characterization. Recurrent neural networks (RNNs) and Long Short-Term
[117]
[118]
Memory (LSTM) handle time-series data, crucial for studying reaction kinetics. State-of-the-art models
[120]
[119]
such as transformers capture complex dependencies using self-attention mechanisms . Shapley
additive explanations (SHAP) and Gradient-weighted class activation mapping (Grad-CAM) improve
[122]
[121]
[123]
model interpretability by highlighting key features. Kolmogorov-Arnold networks (KANs) are adept at
modeling nonlinear systems, revealing insights into complex photocatalytic processes.
Model validation and evaluation
Model validation and evaluation are essential steps in ensuring the robustness and effectiveness of deep
learning models in photocatalyst design. These processes verify predictive accuracy and generalization
ability, ensuring alignment between model predictions and experimental data, while mitigating issues such
as overfitting and underfitting . Key evaluation metrics include mean squared error (MSE) and root MSE
[124]
(RMSE), which quantify prediction errors, and R-squared (R²), which measures the proportion of variance
explained by the model. Cross-validation and hold-out methods are commonly used to partition datasets
[125]
into training and testing subsets, providing unbiased estimates of model performance.
DEEP LEARNING APPROACHES IN PHOTOCATALYST DESIGN
Building on the deep-learning-assisted workflow outlined in the previous section, deep learning techniques
have profoundly transformed multiple facets of photocatalyst design. These approaches encompass six
critical areas: novel photocatalyst discovery, microstructure design, property optimization, innovative
methodologies, application exploration, and mechanistic insights into photocatalytic processes. The
following sections provide a detailed examination of each area, illustrating how deep learning drives
advancements in the design and development of high-performance photocatalysts.
Discovery of novel photocatalysts
Deep learning has revolutionized the discovery of novel photocatalysts by enabling the rapid screening of
vast chemical spaces and predicting materials with superior photocatalytic properties. By leveraging data-
driven models, deep learning accelerates the identification of candidates optimized for light absorption,
charge separation, and photocatalytic efficiency, significantly reducing the time and resources required
compared to traditional experimental approaches.
A recent study utilized machine learning to design perovskite oxide materials for photocatalytic water
[126]
splitting, addressing the inefficiency of traditional trial-and-error methods in discovering new visible-light
photocatalysts, as shown in Figure 7A. The study constructed structural-property models to predict
hydrogen production rates and optimal bandgaps using algorithms such as gradient boosting regression
(GBR), support vector regression (SVR), and backpropagation artificial neural networks (BPANN). The
feature selection process began with an initial set of 24 features, comprising 18 atomic parameters and six

