Page 102 - Read Online
P. 102
Page 4 of 15 Niu et al. J. Mater. Inf. 2025, 5, 45 https://dx.doi.org/10.20517/jmi.2025.22
DB was compiled by Joung et al. in 2020 and is the largest experimental luminescent molecular dataset
exp
[25]
with the most complete luminous properties by now . The samples in the DB dataset are derived from
exp
1,358 papers containing organic luminescent compounds. It includes 19,280 luminescent molecule-solvent
pairs, comprising 6,690 unique luminescent molecules and 374 solvents. The samples from the DB dataset
exp
display a Gaussian-like distribution of λ , σ , and ε , as depicted in Figure 1. Nevertheless, in terms of
emi
emi
max
Φ , the distribution skews towards the lower value region, which could be ascribed to the fact that before
QY
2010, it was experimentally challenging to obtain organic luminescent molecules with a quantum yield
greater than 0.75 . This distribution phenomenon reflects the lengthy trial-and-error process in the
[25]
experimental development of high-quality luminescent molecules. Moreover, the bias of quantum yield also
makes it difficult to train the model for judging photophysical properties.
The limited data scale and low fidelity of experimental data hinder the seamless integration of the
generation and identification. Instead of blindly sampling across the entire chemical space, a more effective
strategy is to focus sampling within regions containing molecules with superior individual properties,
followed by a targeted screening process. This optimization approach could significantly enhance the
efficiency of OLED candidate molecule screening. Recognizing this potential, we selected three high-quality
subsets of molecules from the DB for the exploration of chemical space. Each set consists of 300
exp
molecules, meticulously chosen based on possessing the narrowest σ , the highest Φ , and the highest ε .
QY
max
emi
Detailed construction of LumiGen
In this work, we designed a molecular generation algorithm framework, LumiGen, which could be used for
de novo luminescent molecular generation with modified photophysical properties. LumiGen is composed
of three main parts, which work at once and can form a loop [Figure 2]. Part I is the Molecular Generator
for novel luminescent molecule generation. Part II is a Spectral Discriminator evaluated for comprehensive
spectral properties. Part III is a Sampling Augmentor, which can enhance the sampling of new features and
improve the excellence rate of the molecules generated in the next stage. Through the collaboration of these
three components, LumiGen progressively learns molecular distribution patterns from disjoint labeled
datasets, enabling the direct generation of all-round OLED candidates.
In the construction of the Molecular Generator, we adopted a pre-training and transfer learning strategy to
efficiently generate chemically permissible molecules. We utilize a pre-trained LSTM model to capture the
structural and synthetic preference of molecules on the ChEMBL24 dataset, which is enriched with a diverse
array of both experimentally synthesized and naturally occurring molecular compounds [23,30] . During the
training process, each molecule was encoded as a one-hot vector. The LSTM model is trained to predict the
conditional probability distribution of each token of the encoded simplified molecular input line entry
[23]
system (SMILES) sequence . Based on the transfer learning strategy, the LSTM model is adeptly applied to
high-quality independent-property luminescent molecular sets. The molecular generator is implemented as
a four-layer LSTM network, trained on SMILES strings using categorical cross-entropy loss and the Adam
optimizer. During transfer learning, the first layer’s parameters are frozen, while the second layer is fine-
tuned with a reduced learning rate to adapt to the downstream task in Supplementary Figure 1. By
leveraging its learned representations, the LSTM model efficiently samples the target chemical space. This
capability enables the LSTM to generate a variety of structurally diverse molecules, effectively bridging gaps
within the target chemical space and enhancing our library of potential luminescent candidates.
A Spectral Discriminator is crafted based on a graph convolutional neural network (GCN) . In the feature
[31]
extraction stage, the atomic and bonding information of luminescent molecules and solvent molecules are
encoded separately, and connected to form a composite feature matrix [Supplementary Figure 2 and

