Page 52 - Read Online
P. 52

Page 2 of 17                         Hu et al. J. Mater. Inf. 2025, 5, 44  https://dx.doi.org/10.20517/jmi.2025.21

               Existing computer vision models exhibit persistent limitations in industrial defect detection applications,
               particularly concerning prediction accuracy. Numerous studies highlight fundamental challenges including
               inadequate small-object detection capabilities - especially problematic for discerning subtle metal surface
                                                [5-7]
               defects against complex backgrounds  - current approaches demonstrate significant vulnerability to
               environmental variations (e.g., uneven lighting, surface reflections) . And specialized models frequently
                                                                         [8]
                                                                                          [9]
               demonstrate material-specific bias that compromises cross-domain generalizability . Model-specific
               constraints further exacerbate these limitations: The “You Only Look Once” (YOLO) series (e.g., YOLOv5
               and YOLOv7) demonstrate compromised small-object sensitivity despite their efficiency advantages [10,11]
               while the region-based convolutional neural network (R-CNN) series (e.g., Faster R-CNN and Mask R-
               CNN) incur prohibitive computational costs that hinder industrial adoption despite high accuracy . Even
                                                                                                   [12]
                                                                                      [13]
               specialized solutions such as improved random forests (91% steel defect accuracy) , and Mask R-CNN-
               based rail defect identification network enhancing rail safety through precise defect localization , while
                                                                                                  [14]
               methods specifically optimized for NEU-DET - including enhanced Faster R-CNN variants [15,16] , Region of
               Interest (ROI)-pooling-based steel defect detectors , and the YOLO-DSC algorithm, which significantly
                                                           [17]
                                                           [18]
               improves detection speed on the NEU-DET dataset  - demonstrate limited cross-dataset generalizability
               due to their calibration to particular defect distributions and imaging conditions. Despite these
               advancements in accuracy and efficiency, supervised learning models face critical limitations in real-world
               industrial applications due to their computational demands and reliance on extensive labeled data, which
               refers to information that has been annotated with one or more labels to provide context or meaning. Their
                                                                           [19]
               performance degrades significantly under variable imaging conditions , while the need for large volumes
               of labeled training data poses significant challenges , particularly in industrial contexts. Acquiring
                                                              [20]
               sufficient labeled data - especially for specialized tasks such as manufacturing defect detection - is often
               impractical due to constraints of time, cost, and expertise. Moreover, the labor-intensive nature of data
               labeling complicates the deployment of supervised learning models in practical scenarios.

               To overcome these limitations, the scientific community has increasingly focused on leveraging unlabeled
                                                             [21]
               data without any annotations. Self-supervised learning  has emerged as a promising paradigm by utilizing
               unlabeled data to learn intrinsic features and patterns, thereby reducing dependency on manual annotation.
               The efficacy of self-supervised learning has been demonstrated across diverse fields [22-24] , notably in
               AlphaFold2’s Nobel Prize-winning application of predicting protein structures with high accuracy . Given
                                                                                                  [25]
               the complexities and data constraints in industrial defect detection, self-supervised learning presents a
               highly promising solution. In industries such as aerospace, automotive, and manufacturing, the ability to
               detect surface defects in metallic materials is vital for ensuring product reliability and safety .
                                                                                            [26]

               Researchers have increasingly investigated the potential of self-supervised learning frameworks for detecting
               surface defects in unlabeled image datasets [27-29] . To develop a more practical defect inspection system, the
                                                        [30]
               self-supervised efficient defect detector (SEDD)  was introduced. This detector combines self-supervised
               learning with image segmentation, utilizing an enhanced single-responsiveness self-supervised strategy to
               achieve competitive performance without requiring annotated defective samples. The model has been
               evaluated on three representative datasets, consistently demonstrating superior average precision (AP)
               performance. Compared to traditional CNN models, integrating self-supervised learning principles with
               adaptive learning rates - adjusted based on loss and weight - can significantly improve detection
               outcomes . However, this approach was validated only on a small dataset, highlighting the need for further
                       [31]
               research on larger datasets to substantiate the efficacy of self-supervised learning methods. Importantly,
               leveraging extensive collections of unlabeled images from relevant scenarios allows for pre-training self-
               supervised models on upstream tasks, followed by fine-tuning for downstream target domains or tasks,
               aligning with conventional transfer learning practices . Combining self-supervised learning with transfer
                                                             [32]
   47   48   49   50   51   52   53   54   55   56   57