Page 105 - Read Online
P. 105
Niu et al. J. Mater. Inf. 2025, 5, 45 https://dx.doi.org/10.20517/jmi.2025.22 Page 7 of 15
Table 1. Rules for grade assignment
-1
Grade lg(ε) (M ·cm ) Φ QY σ (nm) λ (nm) a
-1
emi
emi
Excellent > 4.5 > 0.7 < 60 Δλ < 20
Qualified 3.5-4.5 0.4-0.7 60-100 20-40
Bad < 3.5 < 0.4 > 100 Δλ > 40
a
Δλ = λ max - λ . λ max and λ min represent the maximum and minimum emission wavelengths predicted by Union-GCN, respectively. GCN: Graph
min
convolutional neural network.
The pre-trained LSTM model achieved high-performance metrics with a validity index of 97.4%, a
uniqueness index of 97.3%, and a novelty index of 93.8% for the generated SMILES. The high validity
indicates that the model can efficiently translate generated molecules back into molecular structure,
showcasing its capability to generate chemically meaningful SMILES strings. The considerable proportion of
novelty across the dataset highlights the model’s ability to explore a broad chemical space, substantiating its
efficacy in the de novo generation of novel molecules. When the learning chemical space from ChEMBL24
to high-quality luminescent molecule subsets, the validity and novelty of generated data indices slightly
decreased to 81.5% and 81.1%, respectively, while the uniqueness increased to 99.8%. These changes can be
attributed to the increased complexity of the luminescent molecular structures. Subsequently, the Molecular
Generator samples every ten training sessions to monitor the learning process and generate new molecules.
A redundancy rate of 63% by the 20th sampling indicates sufficient exploration within the target space of
luminescent molecules.
Then, uniform manifold approximation and projection (UMAP) technology and the fraction of sp -
3
3
hybridized carbon atoms (Fsp ) are employed to visualize the movement of the chemical space center. In
Figure 3A, the chemical space of ChEMBL is not only in close proximity to but also partially overlaps with
that of the DB . Nevertheless, a clear distinction is evident when comparing the target spaces occupied by
exp
high-quality luminescent molecule subsets. A significant proportion of the chemical space associated with
molecules exhibiting narrow FWHM, high extinction coefficient, and high quantum yield displays
discrepancies from the established DB , indicating the existence of hitherto unexplored regions. These
exp
unknown chemical regions are of critical importance, as they are likely to contain molecules with enhanced
luminescent properties that have not yet been realized experimentally.
Otherwise, molecules from ChEMBL exhibit significantly higher Fsp compared to those in the DB in
3
exp
Figure 3B. A reduction in Fsp typically results in increased molecular rigidity, which can enhance
3
luminescence efficiency by limiting intermolecular rotations and reducing non-radiative transitions. The
3
stability and alignment of Fsp with target space characteristics during the transfer learning process illustrate
the Molecular Generator’s adaptability in learning structural features. Furthermore, Shannon entropy (SSE)
consistently averaged around 0.9 during the transfer learning process, indicating substantial diversity within
sampled molecules . These findings underline the effectiveness of our transfer learning approach in
[32]
generating novel and diverse high-quality molecular structures from high-quality subsets, which is also
shown in Figure 3C and D.
Moreover, we evaluated the performance of the Spectral Discriminator. Firstly, as Union-GCN constructs
and learns from luminescent molecules and solvents as discrete molecular graphs, the model’s performance
exceeds that of GCNs with a single graph as input by approximately 3.15 nm in predicting maximum
emission wavelength in Supplementary Tables 2 and 3. This illustrates the benefit of Union-GCN in
simulating interactions between luminescent molecules and solvents, emphasizing its capacity to capture
accurately the intricate interaction that influences luminescent properties in diverse solvent environments.

