Page 100 - Read Online
P. 100
Page 2 of 15 Niu et al. J. Mater. Inf. 2025, 5, 45 https://dx.doi.org/10.20517/jmi.2025.22
LumiGen demonstrates the ability to learn molecular distribution patterns from disjoint labeled datasets, enabling
the direct generation of all-round OLED candidates, thereby advancing OLED material discovery.
Keywords: Machine learning, luminescent molecules, de novo design, OLED
INTRODUCTION
Chromophores absorb light at specific wavelengths, triggering electronic transitions between molecular
energy states, and subsequently emit light, making them suitable for applications such as organic light-
emitting diodes (OLEDs) . In comparison with liquid crystal display (LCD) and traditional light-emitting
[1-5]
diode (LED) technology, the self-emitting properties of OLED display technology provide outstanding
image quality and energy efficiency, thereby making it the preferred choice for high-end display
solutions [6-10] . Nevertheless, intrinsic constraints associated with fluorescent (first-generation),
phosphorescent (second-generation), and thermally activated delayed fluorescence (TADF, third-
[11]
generation) materials restrict their deployment in OLEDs . For example, anthracene-based OLEDs suffer
from relatively low external quantum efficiency (EQE), which generally remains below 10% in most studies
due to forbidden triplet transitions and ineffective light output coupling . Moreover, the
[12]
photoluminescence quantum yield (PLQY) of pure organic room-temperature phosphorescent materials
(RTP) is notably low, often less than 5%, which is attributed to weak spin-orbit coupling . Due to electron
[11]
and hole separation, TADF molecules typically exhibit a broad full width at half maximum (FWHM), which
compromises color purity in display applications . The current scarcity of pure organic luminescent
[11]
[13]
skeletons presents a significant challenge in meeting the diverse luminous demands of OLED devices .
Furthermore, the ambiguity of structure-activity relationships represents a significant limitation to
traditional molecular design methods, which rely on existing photophysical chemical knowledge. This
approach is inherently inefficient, as it is subject to the constraints of experimental or other human-led
molecular design .
[14]
With advances in high-throughput screening, open material datasets, and machine learning (ML)-driven
property predictors, it has become increasingly feasible to screen materials to identify promising candidates.
For example, Joung et al. trained an ML model using experimental spectral data, successfully predicting
three high-quality luminescent molecules designed by scientists . Similarly, Shi et al. developed a PLQY
[15]
prediction model based on 230 experimental samples of TADF and performed high-throughput screening
[16]
to computationally identify potential candidates for deep-blue OLED applications . Sun et al. further
improved the performance of predictors for maximum absorption and emission wavelengths, FWHM, and
[17]
PLQY . Although ML predictors have greatly enhanced screening efficiency, these approaches primarily
rely on human-led chemical intuition and predefined molecular fragments, limiting their potential for
[15]
discovering novel molecular scaffolds . Furthermore, these methods often fail to generalize beyond
training data, as reflected in their unsatisfactory performance on external datasets . Generative ML models
[17]
have emerged as a promising alternative, offering the ability to design novel molecular scaffolds from
scratch, unrestricted by predefined fragment libraries. Weiss et al. utilized a molecular diffusion model
trained on 475,000 computational molecular data points to generate molecules with specific frontier
molecular orbital gaps . Similarly, Zeni et al. trained a generative model on over 600,000 materials to
[18]
design novel materials with targeted physical and chemical properties, such as bulk modulus and magnetic
density and further conducted experimental validation to demonstrate the feasibility of the generated
materials . Popular generative models, including variational autoencoders (VAEs), generative adversarial
[19]
networks (GANs), and diffusion models, have achieved remarkable breakthroughs in materials science [20-22] .
However, these models typically require hundreds of thousands of training samples, making them less
practical for data-scarce scenarios. Especially, the experimental luminescent molecules dataset typically

