Page 21 - Read Online
P. 21
Page 14 of 33 Liu et al. J Mater Inf 2024;4:33 https://dx.doi.org/10.20517/jmi.2024.48
Data collection and preprocess
Data forms the essential backbone of any deep learning model. In the photocatalyst design, data is typically
derived from experimental measurements, published literature, and high-throughput computational
simulations. These sources provide extensive information on material properties, performance metrics, and
structural features, offering valuable insights into photocatalytic behavior. The process of data collection
and preprocessing is critical for ensuring dataset quality and consistency. Data cleaning is conducted to
address issues such as duplicate entries, missing values, and errors. To handle missing data, various
imputation strategies were employed, tailored to the characteristics of the dataset and the specific
attributes . A hybrid framework leveraging Markov Logic Networks (MLNs) was implemented to infer
[97]
probable values by supplementing insufficient integrity constraints with learned instantiated rules, thus
establishing a robust data cleaning mechanism . Additionally, an "active label cleaning" strategy was
[98]
[99]
utilized to improve data quality, particularly in datasets with label noise. This method prioritized samples
for re-annotation based on estimated label correctness and labeling difficulty, enabling a more targeted and
efficient correction process compared to random selection, thereby optimizing data preparation and
significantly enhancing the overall quality of the training dataset.
Standardization techniques such as min-max scaling and z-score normalization are employed to ensure
[100]
uniformity across different data sources, particularly when integrating datasets with varying formats. For
feature extraction, methods such as principal component analysis (PCA) and autoencoders [101,102] are
commonly used to reduce dimensionality and isolate the most relevant characteristics, transforming raw
data into a structured format suitable for deep learning model inputs. Table 4 lists the popular databases
available for photocatalyst design. These databases offer a wide array of data, ranging from crystal structures
(ICSD, CSD, COD) to computational predictions and material properties (Materials Project, AFLOW,
NREL). Some databases, including MatWeb and MatNavi, focus on experimental data, while others, such as
the Open Quantum Materials Database (OQMD), provide computationally derived properties and high-
throughput calculations, enabling researchers to explore new material candidates and optimize
photocatalytic performance.
Feature engineering and descriptors
Feature engineering is a fundamental step in constructing deep learning models for the design of
[103]
photocatalysts, where raw data must be transformed into structured and informative features that reflect the
photocatalyst properties. This process directly affects the model capability to capture complex relationships
within the data, ultimately determining its accuracy and predictive power. A crucial aspect of feature
engineering is the use of descriptors, which quantitatively represent the physical and chemical
characteristics of photocatalysts. Selecting the appropriate descriptors is essential for ensuring the model
focuses on the most relevant features, significantly enhancing its ability to predict photocatalytic
performance and optimize photocatalyst design .
[104]
Common descriptors in photocatalyst design include electronic structure, geometric structure, and surface
properties, as listed in Table 5. Electronic structure descriptors, such as band gap, valence band maximum
(VBM), conduction band minimum (CBM), and DOS, are crucial for understanding light absorption and
charge carrier dynamics. Geometric descriptors, including size, morphology, crystal planes, and lattice
symmetry, provide insights into the spatial arrangement of atoms and its effect on photocatalytic activity.
Surface property descriptors, including specific surface area, pore structure, and surface functional groups,
influence reactivity and adsorption behavior. These descriptors are essential for building deep learning
models that can accurately predict and optimize photocatalytic performance across various applications,
such as water splitting, pollutant degradation, and hydrogen production. Furthermore, the descriptors
encompass not only electronic structure parameters but also geometrical and surface characteristics, such as

