Page 224 - Read Online
P. 224
Page 2 of 13 Yuan et al. J. Mater. Inf. 2026, 6, 17
materials science have brought about profound transformations in traditional materials design, synthesis,
and characterization . Scanning tunneling microscopy (STM), a vital technique for investigating surface
[1-5]
structures and adsorbates, has become widely employed in fields such as catalysis, molecular self-assembly,
on-surface synthesis, and nanomaterials preparation [6-12] , primarily due to its ability to provide atomic-level
resolution [13-18] . However, the high-resolution images obtained from STM measurements typically exhibit
considerable complexity and contain rich microscopic information. Consequently, the analysis and
interpretation of such images depend heavily on the experience and subjective judgment of human experts.
This reliance not only reduces research efficiency but also restricts the broader adoption and application of
the cutting-edge technology.
To overcome the critical dependence on prior knowledge of human experts in high-resolution image
analysis, AI techniques have increasingly been introduced [19-28] . Recent studies have demonstrated that deep
learning and computer vision algorithms can effectively automate the recognition and analysis of STM
images, including the automatic detection and correction of imaging artifacts , classification of similar
[20]
molecules , molecular counting , and automated length measurement of polymer chains. These studies
[19]
[22]
highlight the significant potential of integrating AI with STM technology to improve analytical accuracy and
reduce manual intervention.
Nevertheless, current research remains largely limited to isolated applications tailored to specific scenarios,
typically focusing only on achieving a single functionality or a particular objective. Specifically, most existing
studies employ small-scale, customized datasets for training and testing, lacking systematic data preparation
and a general framework for predictive modeling and result analysis. Therefore, the systematic integration of
the entire AI pipeline, which includes data acquisition and preprocessing, model training, automated
prediction, and multi-scale data analysis, is one of the key challenges for achieving efficient, automated, and
intelligent STM-based research.
In this work, we develop a comprehensive framework that integrates multiple aspects of the machine
learning workflow, including image data labeling, dataset creation, model training, inference, and data
analysis. Based on the You Only Look Once version 9 (YOLOv9) model, a state-of-the-art approach that
combines both object detection and instance segmentation capabilities, our program provides a one-stop
solution for annotating image data, generating high-quality datasets, and training advanced models.
Additionally, the program includes tools for data analysis based on prediction results, providing insights that
can guide decision-making and further refine model performance. A key feature of the program is its support
for incremental learning, allowing models to be continuously updated and improved as new data become
available. This approach avoids the need for complete retraining from scratch and ensures that the models
remain accurate and up to date.
MATERIALS AND METHODS
Data preparation and labeling
To develop a robust model for molecular image analysis, we first curated a dataset of molecular images
containing various molecular assemblies. The images were pre-processed to standardize resolution and
contrast, ensuring uniformity across the dataset. Using an annotation tool, molecular structures were
manually labeled with bounding boxes and segmentation masks for training the object detection and
instance segmentation models. Each molecular species was assigned a unique class identifier (ID), and the
dataset was divided into training, validation, and test sets following an 80:10:10 split. All STM images used in
this study were obtained from our group’s experiments on molecular self-assembly. To enhance dataset
diversity, 2-3 original high-resolution STM images were augmented using operations such as rotation,
flipping, cropping, and contrast adjustment, resulting in approximately 500 images in total. These

