Page 62 - Read Online
P. 62
Wen et al. J. Mater. Inf. 2025, 5, 30 https://dx.doi.org/10.20517/jmi.2024.102 Page 7 of 21
positions on the molecule. The potential splicing sites correspond to the positions of hydrogen atoms on the
molecular fragments. Splicing between molecular fragments involves substituting the hydrogen atoms at
these specific positions. To identify all possible splicing sites on a molecule, it is necessary to scan through
all atoms in the molecular fragments, locate each hydrogen atom, and assign a unique identifier to each,
which will then be used for subsequent fragment splicing. Additionally, the symmetry of the molecular
structure should be taken into account, as symmetric sites need only be considered once, thus reducing the
computational cost in the subsequent splicing process; (ii) Combination and Structural Optimization:
Combine the molecules in library A with the basic fragment library B, and perform batch structural
optimization using the semi-empirical quantum calculation method AM1 in Gaussian16 software. The
*
optimized molecular structures are then stored in the molecular library A ; (iii) Structural Screening:
*
Conduct structural screening on the molecules in library A after each splicing round, retaining those with
good planarity and storing them in the molecular library A . This step significantly reduces computational
**
load and storage space while ensuring that the molecules in the initial library A of each round exhibit good
planarity, thereby increasing the likelihood of obtaining molecules with desirable planarity in subsequent
rounds. The molecules from the molecular library A generated in each iteration are systematically archived
**
in the final π molecular database; (iv) Atom Count Evaluation: Evaluate the atom count of the molecules in
**
the A library. If the atom count is less than 110, the splicing process is repeated. If the count exceeds 110,
the π molecular splicing process concludes. After 19 rounds of splicing, 200,000 π molecular structures were
generated in the π molecular database.
Molecular splicing design of the D-π-D structures
Upon obtaining the π molecular structure database, the subsequent step is to attach D fragments to both
ends of the intermediate π molecules, thereby forming the target D-π-D structures. The D fragments are
spliced at the most distant positions on the π molecules. The process is carried out in three main steps. (1)
Planarity Screening: A planarity screening is conducted on the 200,000 structures in the π molecular
database, resulting in the selection of 7,399 molecules with excellent planarity; (2) Molecular Splicing: D
fragments are attached to both ends of the selected π molecules to form the complete D-π-D target
molecular structures. Given that multiple potential splicing sites exist on the π molecule, the two most
distant sites are selected as the connection points for the D fragment. The RDkit toolkit is used to convert
the molecular structure into a graph-based representation, from which a distance matrix is computed. By
examining the elements of this matrix, the two nodes with the largest separation are identified as the
splicing sites for the D fragment; (3) Structural Optimization: The resulting D-π-D molecules undergo
structural optimization, and the optimized structures of the 7,399 molecules are stored in the D-π-D
molecular structure database for subsequent high-throughput property calculations. Four representative
D-π-D spliced molecules with varying atom counts are presented in Supplementary Figure 1.
High-throughput computational screening
High-throughput calculations were performed on 7,399 selected molecules in the D-π-D molecular database
to determine their HOMO energy levels, hole reorganization energies, solvation free energies, maximum
absorption peaks, and hydrophobicity (LogP) values. Among these, 7,222 molecules yielded normal results
from DFT calculations, while 177 computational tasks failed to converge. Finally, the SAScore was
computed for the 7,222 successfully converged molecules using the RDKit toolkit. Specific information on
these molecules, including molecular structures and calculated properties, is summarized in Supplementary
Files 1 and 2, respectively.
Figure 2 illustrates the high-throughput computational results for various performance parameters of 7,222
D-π-D HTM molecules. The x-axis represents the ID number of the molecule, with larger numbers
corresponding to molecules with a greater number of atoms. As shown in Figure 2A and B, no significant
correlation is observed between the HOMO energy levels or the hole reorganization energies and the

