Page 111 - Read Online
P. 111
Niu et al. J. Mater. Inf. 2025, 5, 45 https://dx.doi.org/10.20517/jmi.2025.22 Page 13 of 15
promises to enrich the training datasets available for LumiGen, potentially promoting the development and
testing of luminescent materials.
CONCLUSIONS
In this work, we introduced LumiGen, a novel framework combining a Molecular Generator, a Spectral
Discriminator, and a Sampling Augmentor to achieve the de novo design of luminescent molecular
generation. This tool is particularly adept at populating and filtering the chemical space of high-quality
luminescent molecules, and the Sampling Augmentor is capable of optimizing the model to generate
higher-quality luminescent molecules. In particular, the multi-expert voting and elite selection strategy
effectively address substantial statistical errors in experimental data, thereby ensuring that only high-quality
luminescent molecules are retained. The generated molecules display minimal structural resemblance to the
original dataset and exhibit high SSE, demonstrating their novelty and diversity. The validity of LumiGen in
distinguishing between MolElite and MolMediocrity is also confirmed by TD-DFT calculations, which are
conducted with the aim of tailoring targeted luminescent molecules for advanced optoelectronic
applications. By synthesizing and characterizing high-quality molecular scaffolds, LumiGen can seamlessly
integrate theoretical predictions with experimental validations. What is more, LumiGen is effective in
limited datasets ASBase, which makes it advantageous in resource-constrained situations. With the
emergence of batch literature data extraction tools and large language models, we are set to enhance our
training datasets, boosting the learning potential of our framework. As the pioneering ML framework for de
novo luminescent molecular design, LumiGen strategically bridges gaps in high-quality experimental data to
discover versatile luminescent molecules, poised to transform the search for next-generation display
technologies.
DECLARATIONS
Acknowledgments
The authors gratefully acknowledge the National Supercomputer Center in Tianjin (Tianhe 3F) and the
Scientific Computing Center of CIC, Tianjin University for providing computation facilities.
Authors’ contributions
Data curation, methodology, writing - original draft: Niu, X.
Data curation, formal analysis: Su, Z.; Zhang, H.
Data curation: Wang, L.; Shi, W.
Writing - review and editing: Dang, Y.
Data curation, writing - original draft: Yuan, Y.
Funding acquisition, project administration, supervision, writing - review and editing: Sun, Y.
Funding acquisition, writing - review and editing: Hu, W.
Availability of data and materials
The code is open source at https://github.com/YajingSun-Group/LumiGen. The relevant datasets, model
architecture, and generated data can be found in the links.
Financial support and sponsorship
This work was financially supported by the National Natural Science Foundation of China (22473085,
22003046 and 52121002), the Ministry of Science and Technology of China (2022YFA1204401) and Xiaomi
Young Talents Program.

