Page 104 - Read Online
P. 104
Page 6 of 15 Niu et al. J. Mater. Inf. 2025, 5, 45 https://dx.doi.org/10.20517/jmi.2025.22
domain. In contrast, those experts tasked with predicting different photophysical properties are recognized
as experts from distinct domains. This specialized approach not only tailors the predictive capabilities of
each model to specific aspects of luminescent properties but also ensures that insights from diverse domains
can be integrated. To mitigate the impact of experimental errors and data insufficiencies, our elite selection
strategy integrates the collective decisions of the Union-GCN. To ensure model diversity within the Union-
GCN ensemble, we use identical network architectures and hyperparameters across all models, while
varying the random seeds for training. This approach introduces diversity through different data
permutations and initialization paths, leading to complementary predictions. This strategy shifts the focus
from precisely predicting the luminescent properties of molecules to comprehensively identifying high-
quality luminescent molecular sets. Specifically, we randomly select 80% of the luminescent molecule
samples as a training set and use different random seeds to train Union-GCN. As a result, we obtained
multiple experts from the same and different domains. Then, Union-GCN predicts the luminescent
properties of sampled molecules in a hypothetical solvent and evaluates the consistency of experts from the
same domains. If there is significant divergence among experts within the same domain [e.g., different
experts in (λ)], it indicates the difficulty of accurately evaluating the luminescent properties based on the
current dataset. A molecule can only advance to the next round of classification if all experts within the
same domain agree that it belongs to the same grade in Table 1. Thirdly, Spectral Discriminator channels
molecules with three “Excellent” luminescent properties, as determined by experts from three different
domains [e.g., (ε), (σ), and (λ)], into the high-quality MolElite and those with three “Qualified” properties
into the MolMediocrity.
In the Sampling Augmentor, we utilize Morgan fingerprints, a type of circular fingerprint for molecules, as a
descriptor to facilitate the clustering of molecules within the MolElite. These fingerprints capture the
molecular structure by encoding the presence of specific chemical substructures within a fixed radius
around each atom, providing a robust basis for comparing molecular similarities. The entire dataset is
divided into 50 distinct groups using the k-means algorithm depending on the distribution and diversity of
the molecular structures within the dataset. By dividing the molecules into well-defined groups, the
generator is exposed to a wide array of structural motifs. By selecting 300 molecules from different clusters
for the next generation, we maintain a diverse and representative high-quality light-emitting molecule
subset. This approach fosters a broad exploration of the chemical space, enhancing the likelihood of
discovering novel, efficient luminescent materials in each iteration of the cycle, from molecular generation
through to spectral identification. The module selection and hyperparameter optimization procedures are
illustrated in Supplementary Figures 3 and 4.
RESULTS AND DISCUSSION
Baseline performance evaluation of LumiGen
To comprehensively evaluate the baseline performance of the LumiGen framework, we conducted a
systematic analysis of its three key components: Molecular Generator, Spectral Discriminator, and Sampling
Augmentor. The Molecular Generator was assessed in terms of chemical space exploration, molecular
novelty, and diversity, verifying its capability to reduce human intervention and optimize sampling quality.
The performance of the Spectral Discriminator was evaluated by examining its prediction accuracy for
photophysical properties in different solvent environments, assessing its ability to capture the interactions
between emissive molecules and solvents. Additionally, the iterative optimization process of the Sampling
Augmentor was monitored, quantifying changes in the proportion of high-quality molecules to determine
its effectiveness in enhancing screening efficiency and refining molecular generation. The following sections
provide a detailed performance assessment of each component within LumiGen.

