Page 103 - Read Online
P. 103
Niu et al. J. Mater. Inf. 2025, 5, 45 https://dx.doi.org/10.20517/jmi.2025.22 Page 5 of 15
Figure 1. DB dataset distributions for (A) λ , (B) σ , (C) lg(ε ), and (D) Φ .
exp emi emi max QY
Figure 2. LumiGen framework, comprising a Molecular Generator, a Spectral Discriminator, and a Sampling Augmentor. The Molecular
Generator leverages a ChEMBL24 pre-trained LSTM model to produce three types of high-quality luminescent molecules. The Spectral
Discriminator is trained on the DB dataset. The Sampling Augmentor generates new high-quality subsets through clustering,
exp
completing the loop from lab to ML. LSTM: Long short-term memory; ML: machine learning.
Supplementary Table 1]. This feature construction strategy is superior for differentiating luminescent
molecules from solvent molecules, thus accurately simulating changes in luminescent properties when
molecules are in different solvents. Due to the unevenness of the experimental dataset, where different
photophysical properties of molecules may not be reported simultaneously, we have implemented a multi-
expert voting strategy to comprehensively evaluate a luminous molecule, known as Union-GCN. Certain
experts are dedicated to predicting quantum yield, while some focus on predicting FWHM or maximum
emission wavelength. Experts who specialize in the same property are identified as experts within the same

