Page 21 - Read Online
P. 21

Page 14 of 33                          Liu et al. J Mater Inf 2024;4:33  https://dx.doi.org/10.20517/jmi.2024.48

               Data collection and preprocess
               Data forms the essential backbone of any deep learning model. In the photocatalyst design, data is typically
               derived from experimental measurements, published literature, and high-throughput computational
               simulations. These sources provide extensive information on material properties, performance metrics, and
               structural features, offering valuable insights into photocatalytic behavior. The process of data collection
               and preprocessing is critical for ensuring dataset quality and consistency. Data cleaning is conducted to
               address issues such as duplicate entries, missing values, and errors. To handle missing data, various
               imputation strategies were employed, tailored to the characteristics of the dataset and the specific
               attributes . A hybrid framework leveraging Markov Logic Networks (MLNs) was implemented to infer
                       [97]
               probable values by supplementing insufficient integrity constraints with learned instantiated rules, thus
               establishing a robust data cleaning mechanism . Additionally, an "active label cleaning" strategy  was
                                                        [98]
                                                                                                    [99]
               utilized to improve data quality, particularly in datasets with label noise. This method prioritized samples
               for re-annotation based on estimated label correctness and labeling difficulty, enabling a more targeted and
               efficient correction process compared to random selection, thereby optimizing data preparation and
               significantly enhancing the overall quality of the training dataset.


               Standardization techniques such as min-max scaling and z-score normalization  are employed to ensure
                                                                                   [100]
               uniformity across different data sources, particularly when integrating datasets with varying formats. For
               feature extraction, methods such as principal component analysis (PCA) and autoencoders [101,102]  are
               commonly used to reduce dimensionality and isolate the most relevant characteristics, transforming raw
               data into a structured format suitable for deep learning model inputs. Table 4 lists the popular databases
               available for photocatalyst design. These databases offer a wide array of data, ranging from crystal structures
               (ICSD, CSD, COD) to computational predictions and material properties (Materials Project, AFLOW,
               NREL). Some databases, including MatWeb and MatNavi, focus on experimental data, while others, such as
               the Open Quantum Materials Database (OQMD), provide computationally derived properties and high-
               throughput calculations, enabling researchers to explore new material candidates and optimize
               photocatalytic performance.

               Feature engineering and descriptors
               Feature engineering  is a fundamental step in constructing deep learning models for the design of
                                 [103]
               photocatalysts, where raw data must be transformed into structured and informative features that reflect the
               photocatalyst properties. This process directly affects the model capability to capture complex relationships
               within the data, ultimately determining its accuracy and predictive power. A crucial aspect of feature
               engineering is the use of descriptors, which quantitatively represent the physical and chemical
               characteristics of photocatalysts. Selecting the appropriate descriptors is essential for ensuring the model
               focuses on the most relevant features, significantly enhancing its ability to predict photocatalytic
               performance and optimize photocatalyst design .
                                                       [104]

               Common descriptors in photocatalyst design include electronic structure, geometric structure, and surface
               properties, as listed in Table 5. Electronic structure descriptors, such as band gap, valence band maximum
               (VBM), conduction band minimum (CBM), and DOS, are crucial for understanding light absorption and
               charge carrier dynamics. Geometric descriptors, including size, morphology, crystal planes, and lattice
               symmetry, provide insights into the spatial arrangement of atoms and its effect on photocatalytic activity.
               Surface property descriptors, including specific surface area, pore structure, and surface functional groups,
               influence reactivity and adsorption behavior. These descriptors are essential for building deep learning
               models that can accurately predict and optimize photocatalytic performance across various applications,
               such as water splitting, pollutant degradation, and hydrogen production. Furthermore, the descriptors
               encompass not only electronic structure parameters but also geometrical and surface characteristics, such as
   16   17   18   19   20   21   22   23   24   25   26