Page 102 - Read Online
P. 102

Page 4 of 15                         Niu et al. J. Mater. Inf. 2025, 5, 45  https://dx.doi.org/10.20517/jmi.2025.22

               DB  was compiled by Joung et al. in 2020 and is the largest experimental luminescent molecular dataset
                  exp
                                                            [25]
               with the most complete luminous properties by now . The samples in the DB  dataset are derived from
                                                                                   exp
               1,358 papers containing organic luminescent compounds. It includes 19,280 luminescent molecule-solvent
               pairs, comprising 6,690 unique luminescent molecules and 374 solvents. The samples from the DB  dataset
                                                                                                  exp
               display a Gaussian-like distribution of λ , σ , and ε , as depicted in Figure 1. Nevertheless, in terms of
                                                  emi
                                                     emi
                                                             max
               Φ , the distribution skews towards the lower value region, which could be ascribed to the fact that before
                 QY
               2010, it was experimentally challenging to obtain organic luminescent molecules with a quantum yield
               greater than 0.75 . This distribution phenomenon reflects the lengthy trial-and-error process in the
                              [25]
               experimental development of high-quality luminescent molecules. Moreover, the bias of quantum yield also
               makes it difficult to train the model for judging photophysical properties.
               The limited data scale and low fidelity of experimental data hinder the seamless integration of the
               generation and identification. Instead of blindly sampling across the entire chemical space, a more effective
               strategy is to focus sampling within regions containing molecules with superior individual properties,
               followed by a targeted screening process. This optimization approach could significantly enhance the
               efficiency of OLED candidate molecule screening. Recognizing this potential, we selected three high-quality
               subsets of molecules from the DB  for the exploration of chemical space. Each set consists of 300
                                              exp
               molecules, meticulously chosen based on possessing the narrowest σ , the highest Φ , and the highest ε .
                                                                                      QY
                                                                                                       max
                                                                        emi
               Detailed construction of LumiGen
               In this work, we designed a molecular generation algorithm framework, LumiGen, which could be used for
               de novo luminescent molecular generation with modified photophysical properties. LumiGen is composed
               of three main parts, which work at once and can form a loop [Figure 2]. Part I is the Molecular Generator
               for novel luminescent molecule generation. Part II is a Spectral Discriminator evaluated for comprehensive
               spectral properties. Part III is a Sampling Augmentor, which can enhance the sampling of new features and
               improve the excellence rate of the molecules generated in the next stage. Through the collaboration of these
               three components, LumiGen progressively learns molecular distribution patterns from disjoint labeled
               datasets, enabling the direct generation of all-round OLED candidates.


               In the construction of the Molecular Generator, we adopted a pre-training and transfer learning strategy to
               efficiently generate chemically permissible molecules. We utilize a pre-trained LSTM model to capture the
               structural and synthetic preference of molecules on the ChEMBL24 dataset, which is enriched with a diverse
               array of both experimentally synthesized and naturally occurring molecular compounds [23,30] . During the
               training process, each molecule was encoded as a one-hot vector. The LSTM model is trained to predict the
               conditional probability distribution of each token of the encoded simplified molecular input line entry
                                      [23]
               system (SMILES) sequence . Based on the transfer learning strategy, the LSTM model is adeptly applied to
               high-quality independent-property luminescent molecular sets. The molecular generator is implemented as
               a four-layer LSTM network, trained on SMILES strings using categorical cross-entropy loss and the Adam
               optimizer. During transfer learning, the first layer’s parameters are frozen, while the second layer is fine-
               tuned with a reduced learning rate to adapt to the downstream task in Supplementary Figure 1. By
               leveraging its learned representations, the LSTM model efficiently samples the target chemical space. This
               capability enables the LSTM to generate a variety of structurally diverse molecules, effectively bridging gaps
               within the target chemical space and enhancing our library of potential luminescent candidates.

               A Spectral Discriminator is crafted based on a graph convolutional neural network (GCN) . In the feature
                                                                                            [31]
               extraction stage, the atomic and bonding information of luminescent molecules and solvent molecules are
               encoded separately, and connected to form a composite feature matrix [Supplementary Figure 2 and
   97   98   99   100   101   102   103   104   105   106   107