Page 104 - Read Online
P. 104

Page 6 of 15                         Niu et al. J. Mater. Inf. 2025, 5, 45  https://dx.doi.org/10.20517/jmi.2025.22

               domain. In contrast, those experts tasked with predicting different photophysical properties are recognized
               as experts from distinct domains. This specialized approach not only tailors the predictive capabilities of
               each model to specific aspects of luminescent properties but also ensures that insights from diverse domains
               can be integrated. To mitigate the impact of experimental errors and data insufficiencies, our elite selection
               strategy integrates the collective decisions of the Union-GCN. To ensure model diversity within the Union-
               GCN ensemble, we use identical network architectures and hyperparameters across all models, while
               varying the random seeds for training. This approach introduces diversity through different data
               permutations and initialization paths, leading to complementary predictions. This strategy shifts the focus
               from precisely predicting the luminescent properties of molecules to comprehensively identifying high-
               quality luminescent molecular sets. Specifically, we randomly select 80% of the luminescent molecule
               samples as a training set and use different random seeds to train Union-GCN. As a result, we obtained
               multiple experts from the same and different domains. Then, Union-GCN predicts the luminescent
               properties of sampled molecules in a hypothetical solvent and evaluates the consistency of experts from the
               same domains. If there is significant divergence among experts within the same domain [e.g., different
               experts in (λ)], it indicates the difficulty of accurately evaluating the luminescent properties based on the
               current dataset. A molecule can only advance to the next round of classification if all experts within the
               same domain agree that it belongs to the same grade in Table 1. Thirdly, Spectral Discriminator channels
               molecules with three “Excellent” luminescent properties, as determined by experts from three different
               domains [e.g., (ε), (σ), and (λ)], into the high-quality MolElite and those with three “Qualified” properties
               into the MolMediocrity.

               In the Sampling Augmentor, we utilize Morgan fingerprints, a type of circular fingerprint for molecules, as a
               descriptor to facilitate the clustering of molecules within the MolElite. These fingerprints capture the
               molecular structure by encoding the presence of specific chemical substructures within a fixed radius
               around each atom, providing a robust basis for comparing molecular similarities. The entire dataset is
               divided into 50 distinct groups using the k-means algorithm depending on the distribution and diversity of
               the molecular structures within the dataset. By dividing the molecules into well-defined groups, the
               generator is exposed to a wide array of structural motifs. By selecting 300 molecules from different clusters
               for the next generation, we maintain a diverse and representative high-quality light-emitting molecule
               subset. This approach fosters a broad exploration of the chemical space, enhancing the likelihood of
               discovering novel, efficient luminescent materials in each iteration of the cycle, from molecular generation
               through to spectral identification. The module selection and hyperparameter optimization procedures are
               illustrated in Supplementary Figures 3 and 4.


               RESULTS AND DISCUSSION
               Baseline performance evaluation of LumiGen
               To comprehensively evaluate the baseline performance of the LumiGen framework, we conducted a
               systematic analysis of its three key components: Molecular Generator, Spectral Discriminator, and Sampling
               Augmentor. The Molecular Generator was assessed in terms of chemical space exploration, molecular
               novelty, and diversity, verifying its capability to reduce human intervention and optimize sampling quality.
               The performance of the Spectral Discriminator was evaluated by examining its prediction accuracy for
               photophysical properties in different solvent environments, assessing its ability to capture the interactions
               between emissive molecules and solvents. Additionally, the iterative optimization process of the Sampling
               Augmentor was monitored, quantifying changes in the proportion of high-quality molecules to determine
               its effectiveness in enhancing screening efficiency and refining molecular generation. The following sections
               provide a detailed performance assessment of each component within LumiGen.
   99   100   101   102   103   104   105   106   107   108   109