Page 62 - Read Online
P. 62

Wen et al. J. Mater. Inf. 2025, 5, 30  https://dx.doi.org/10.20517/jmi.2024.102  Page 7 of 21

               positions on the molecule. The potential splicing sites correspond to the positions of hydrogen atoms on the
               molecular fragments. Splicing between molecular fragments involves substituting the hydrogen atoms at
               these specific positions. To identify all possible splicing sites on a molecule, it is necessary to scan through
               all atoms in the molecular fragments, locate each hydrogen atom, and assign a unique identifier to each,
               which will then be used for subsequent fragment splicing. Additionally, the symmetry of the molecular
               structure should be taken into account, as symmetric sites need only be considered once, thus reducing the
               computational cost in the subsequent splicing process; (ii) Combination and Structural Optimization:
               Combine the molecules in library A with the basic fragment library B, and perform batch structural
               optimization using the semi-empirical quantum calculation method AM1 in Gaussian16 software. The
                                                                                 *
               optimized molecular structures are then stored in the molecular library A ; (iii) Structural Screening:
                                                                  *
               Conduct structural screening on the molecules in library A  after each splicing round, retaining those with
               good planarity and storing them in the molecular library A . This step significantly reduces computational
                                                                 **
               load and storage space while ensuring that the molecules in the initial library A of each round exhibit good
               planarity, thereby increasing the likelihood of obtaining molecules with desirable planarity in subsequent
               rounds. The molecules from the molecular library A  generated in each iteration are systematically archived
                                                           **
               in the final π molecular database; (iv) Atom Count Evaluation: Evaluate the atom count of the molecules in
                    **
               the A  library. If the atom count is less than 110, the splicing process is repeated. If the count exceeds 110,
               the π molecular splicing process concludes. After 19 rounds of splicing, 200,000 π molecular structures were
               generated in the π molecular database.


               Molecular splicing design of the D-π-D structures
               Upon obtaining the π molecular structure database, the subsequent step is to attach D fragments to both
               ends of the intermediate π molecules, thereby forming the target D-π-D structures. The D fragments are
               spliced at the most distant positions on the π molecules. The process is carried out in three main steps. (1)
               Planarity Screening: A planarity screening is conducted on the 200,000 structures in the π molecular
               database, resulting in the selection of 7,399 molecules with excellent planarity; (2) Molecular Splicing: D
               fragments are attached to both ends of the selected π molecules to form the complete D-π-D target
               molecular structures. Given that multiple potential splicing sites exist on the π molecule, the two most
               distant sites are selected as the connection points for the D fragment. The RDkit toolkit is used to convert
               the molecular structure into a graph-based representation, from which a distance matrix is computed. By
               examining the elements of this matrix, the two nodes with the largest separation are identified as the
               splicing sites for the D fragment; (3) Structural Optimization: The resulting D-π-D molecules undergo
               structural optimization, and the optimized structures of the 7,399 molecules are stored in the D-π-D
               molecular structure database for subsequent high-throughput property calculations. Four representative
               D-π-D spliced molecules with varying atom counts are presented in Supplementary Figure 1.


               High-throughput computational screening
               High-throughput calculations were performed on 7,399 selected molecules in the D-π-D molecular database
               to determine their HOMO energy levels, hole reorganization energies, solvation free energies, maximum
               absorption peaks, and hydrophobicity (LogP) values. Among these, 7,222 molecules yielded normal results
               from DFT calculations, while 177 computational tasks failed to converge. Finally, the SAScore was
               computed for the 7,222 successfully converged molecules using the RDKit toolkit. Specific information on
               these molecules, including molecular structures and calculated properties, is summarized in Supplementary
               Files 1 and 2, respectively.


               Figure 2 illustrates the high-throughput computational results for various performance parameters of 7,222
               D-π-D HTM molecules. The x-axis represents the ID number of the molecule, with larger numbers
               corresponding to molecules with a greater number of atoms. As shown in Figure 2A and B, no significant
               correlation is observed between the HOMO energy levels or the hole reorganization energies and the
   57   58   59   60   61   62   63   64   65   66   67