Page 168 - Read Online
P. 168
Page 12 of 19 Wang et al. J. Mater. Inf. 2026, 6, 13
Figure 11. Images generated by conditional GANs. GANs: Generative adversarial networks.
Figure 12. Segmentation performances with different synthetic and real datasets. The P_0, P_10, and P_20 use physical synthetic images
with 0, 10, 20 real images respectively. The G_20 uses GANs-based synthetic images with 20 real images. The R_100 uses 100 real
images without any synthetic images. The synthetic and real images are directly mixed into one training set in each training process.
GANs: Generative adversarial networks.
more important for segmentation performance. Meanwhile, because there are fewer dark scratch images in
the real dataset, GANs perform poorly in generating dark scratches (e.g., the third image in Figure 11),
showing that data-driven methods cannot reliably generate new or correct information beyond the
distribution of existing real images. This phenomenon has also been reported in other studies .
[25]
The results for P_20 indicate that physically based synthetic images are effective in improving segmentation
performance when real images are limited. With 1,200 synthetic images, the IoU improves from 0.66 to 0.83.
Compared with R_100, real images remain more efficient for segmentation model training: using only 100
real images, the IoU reaches 0.74, whereas 700 physically based synthetic images are required with 20 real
images to achieve similar performance. In this scenario, approximately nine physical synthetic images are
needed to replace one real image. However, since physically based synthetic data and their labels can be
efficiently generated in batches, this approach remains a useful compensation method when real images are
limited. Labeling for segmentation is time-consuming: if manually labeling one image takes 1.5 min, using
synthetic images can save about 2.5 h per 100 real images.

