Page 101 - Read Online
P. 101
Niu et al. J. Mater. Inf. 2025, 5, 45 https://dx.doi.org/10.20517/jmi.2025.22 Page 3 of 15
encompasses only a few hundred samples, with the larger luminescent experimental datasets only
3
[18]
containing 10 entries . Insufficient data hampers the ability to impose attribute-conditioned constraints in
5
generative models without leveraging large computational datasets, which often contain approximately 10
entries [16,18] . Long short-term memory (LSTM) models excel in distribution learning for small datasets .
[23]
However, under conditions of low-fidelity and limited data, enforcing property constraints becomes
challenging . The process of selecting potential OLED candidates remains constrained by the limited
[24]
number of known materials, experimental optical measurement errors, and the reliance on human
[25]
expertise . With the rapid emergence of diverse ML approaches, selecting targeted and efficient methods
for OLED candidate screening and enabling an autonomous iterative learning framework are crucial steps
toward bridging the gap between research and application, ultimately enabling the successful synthesis and
validation of OLED candidates.
In this work, we introduce LumiGen, an integrated framework designed for the de novo design of high-
quality OLED candidate molecules. To learn molecular distribution patterns from single-property datasets,
effectively capture independent property domains, and directly generate molecules with well-rounded
performance across multiple properties, LumiGen integrates a Molecular Generator, a Spectral
Discriminator, and a Sampling Augmentor. The Molecular Generator explores the chemical space of high-
quality independent-property luminescent molecules, effectively mitigating issues related to over-reliance
on human intervention and randomness in candidate selection. Given the errors in experimental
measurements and the scarcity of experimental data, we have implemented a multi-expert voting strategy
and elite selection approach within the Spectral Discriminator. This shift in focus moves from precisely
predicting molecular luminescent properties to comprehensively identifying high-quality luminescent
molecular sets (MolElite). Time-dependent density functional theory (TD-DFT) calculations of the
generated candidate molecules validate that LumiGen achieves an accuracy of around 80.2%. During the
operation of the Sampling Augmentor, the generation rate of high-quality (Elite) molecules is enhanced by
more than threefold (from 6.56% to 21.13%), while the proportion of lower-quality molecules (Mediocrity)
molecules decreases by over fivefold (from 0.775% to 0.131%). Furthermore, we have successfully
synthesized a new molecular skeleton from MolElite, whose spectral characteristics fully meet the
expectations set by LumiGen. Experimental measurements revealed an FWHM of 49.8 nm, an extinction
4
coefficient of 5.25 × 10 M ·cm , and a high quantum yield of 88.6% in dichloromethane solution. In the
-1
-1
original dataset, only 0.33% of the molecules outperform our synthesized molecules in terms of overall
optical performance. LumiGen fills the gap in direct generative models for organic optoelectronic functional
molecules, addressing the challenge of imbalanced experimental data, and positioning itself as a powerful
tool for the future development of organic optoelectronic materials.
MATERIALS AND METHODS
Data collection and analysis
To harness the capabilities of artificial intelligence (AI) for screening OLED candidate molecules, it is
essential to evaluate the available data resources. Several open material databases of luminescent molecules
exist, including the DB dataset of chromophore experimental data, the ASBase dataset of aggregation-
exp
induced emission (AIE) molecules, the FORMAD dataset containing computed excited-state properties,
and some smaller experimental TADF datasets [9,25-29] . In this study, we focus primarily on DB and ASBase,
exp
as they directly pertain to OLED performance, providing key luminescence properties such as PLQY, Φ ,
QY
and maximum emission wavelength, λ , FWHM, σ , and extinction coefficient, ε . While FORMAD
emi
max
emi
offers detailed excited-state energy calculations, it may not be as intuitive for our objective compared to
experimental datasets . Additionally, existing experimental TADF datasets are typically limited to a few
[29]
hundred molecules, are not publicly available, or lack critical optical property data, making them unsuitable
for direct use.

