Page 171 - Read Online
P. 171
Wang et al. J. Mater. Inf. 2026, 6, 13 Page 15 of 19
Figure 16. Segmentation results based on labels with different qualities. IoUs: Intersection-over-unions.
training data is a combination of 1,200 physical synthetic and 20 real images. The labels could only shrink up
to 5 pixels, so our experiment stopped there. The results, as presented in Figure 16, show that the labeling
accuracy matters for scratch segmentation. An interesting phenomenon is that the larger coverage is slightly
preferred than smaller. Although the larger coverage introduces redundant and misleading information, it
still covers all the scratch features for extraction, while the smaller labels would miss the boundary features.
Evaluation of data fusion strategy
As the real images are not exactly the same as synthetic images, how to merge them for better performance
also requires investigation. Here, we compare two common strategies - direct data merging and transfer
learning - by combining either 10 or 20 real images with varying numbers of synthetic images (200, 400, 600,
800, or 1,000). The results in Figure 17 show that increasing the number of real images improves
segmentation accuracy. In synthetic images, the effect of different lighting and imaging conditions and
scratch shapes can be modeled by physical laws in computer graphics, while texture differences are the main
cause of discrepancies between synthetic and real images. The continuous increasing trend in Figure 17
shows that the different scratch shapes and imaging settings carried by different synthetic data can bring new
and meaningful information for the segmentation model and improve its generalization ability. Differences
in texture can be compensated for by adding real images in the training process. If real images are limited,
the merging strategy plays an important role. The transfer learning method shows a clear improvement (near
10%) in segmentation accuracy. The result of transfer learning with 10 real images is even better than data
merging with 20 real images. This indicates that mixing a limited number of real images into massive
synthetic images cannot make the segmentation networks fully focus on real features and may lead to
learning too many synthetic features. With knowledge transfer, unique synthetic features are down-weighted,
and real features are reinforced, resulting in better performance on real scratch images.
Segmentation example analysis
Figure 18 presents examples of segmentation results from the proposed method using 1,200 synthetic and 20
real images for an intuitive understanding. Most scratch areas are identified correctly, showing that physical
synthetic data can continuously provide useful information in large volumes for scratch feature extraction.
The challenging areas are the transitions between light and dark regions, where the appearance changes

