Page 168 - Read Online
P. 168

Page 12 of 19                                                    Wang et al. J. Mater. Inf. 2026, 6, 13





















                               Figure 11. Images generated by conditional GANs. GANs: Generative adversarial networks.

































               Figure 12. Segmentation performances with different synthetic and real datasets. The P_0, P_10, and P_20 use physical synthetic images
               with 0, 10, 20 real images respectively. The G_20 uses GANs-based synthetic images with 20 real images. The R_100 uses 100 real
               images without any synthetic images. The synthetic and real images are directly mixed into one training set in each training process.
               GANs: Generative adversarial networks.


               more important for segmentation performance. Meanwhile, because there are fewer dark scratch images in
               the real dataset, GANs perform poorly in generating dark scratches (e.g., the third image in Figure 11),
               showing that data-driven methods cannot reliably generate new or correct information beyond the
               distribution of existing real images. This phenomenon has also been reported in other studies .
                                                                                             [25]
               The results for P_20 indicate that physically based synthetic images are effective in improving segmentation
               performance when real images are limited. With 1,200 synthetic images, the IoU improves from 0.66 to 0.83.
               Compared with R_100, real images remain more efficient for segmentation model training: using only 100
               real images, the IoU reaches 0.74, whereas 700 physically based synthetic images are required with 20 real
               images to achieve similar performance. In this scenario, approximately nine physical synthetic images are
               needed to replace one real image. However, since physically based synthetic data and their labels can be
               efficiently generated in batches, this approach remains a useful compensation method when real images are
               limited. Labeling for segmentation is time-consuming: if manually labeling one image takes 1.5 min, using
               synthetic images can save about 2.5 h per 100 real images.
   163   164   165   166   167   168   169   170   171   172   173