Page 45 - Read Online
P. 45

Page 4 of 18                         Li et al. J. Mater. Inf. 2025, 5, 29  https://dx.doi.org/10.20517/jmi.2024.103

               relationships through expert knowledge and black-box models. We aim to verify the feasibility of SR
               methods on aging data and apply them to the discovery of quantitative relationships in rubber aging.

               Moreover, in terms of evaluating the specific scientific discovery potential of SR, there are also deficiencies.
               Currently, the performance evaluation of SR algorithms is mostly based on artificial datasets, which do not
               conform to the characteristics of experimental data. Although Matsubara et al. proposed a more realistic SR
                                                                                                   [43]
               evaluation framework, it still did not fully describe the characteristics of real experimental data . This
               mismatch between evaluation datasets and real experimental data restricts a more accurate and
               comprehensive understanding of the capabilities and limitations of SR in the context of material aging
               research. As a result, the full exploration of its role in uncovering the quantitative relationships of rubber
               material aging is hindered, further emphasizing the need for a more suitable evaluation framework and in-
               depth investigation in this area.

               This study focuses on the evaluation of SR algorithms and their application to experimental data of rubber
               material aging, with the aim of obtaining more robust quantitative relationships between microscopic
               structure characterization and macroscopic properties. The major contributions of this study are as follows:

               1. A Comprehensive Evaluation Framework for SR: Unlike existing algorithm evaluations centered on
               benchmark datasets lacking physical meaning, this paper assesses SR across diverse real-world scenarios,
               including data paucity, noise, and extraneous variables, to gauge its practical viability in real-world
               applications.
               2. SR Application in Rubber Aging: This study is the pioneer attempt to utilize SR for unveiling and
               modeling the quantitative relationships of micro-macro aging mechanisms in polymer materials based on
               experimental data. It offers potential revelations of relationships eluding traditional methods.
               3. Quantitative Relationship Discovery: By applying SR to aging experimental data, this research presents
               quantitative connections between the microscopic traits and macroscopic performance of aging materials,
               providing rational explanations that enhance the understanding of aging phenomena in particular polymer
               material systems.


               MATERIALS AND METHODS
               To address the issues regarding SR evaluation, this paper proposes the SR4Real evaluation framework, a
               more comprehensive SR evaluation framework. It aims to screen out superior SR methods for aging
               experimental data. The overall workflow is illustrated in Figure 1. The following are the components of the
               workflow.

               SR4Real benchmark
               We follow the dataset partitioning given by Matsubara et al., selecting ten simple formulas, six formulas
                                                                          [43]
               with more operations, and six formulas with large data value ranges . Each formula contains 1,000 data
               points corresponding to the fully fitted equations. Additionally, subsets with fifty data points per formula,
               noise levels of 0.1, and unrelated variables are generated based on ten equations in simple formulas. For
               detailed information on the establishment process of the SR4Real dataset, refer to the “Data construction
               details” section in the Supplementary Materials. This results in six types of datasets, each testing SR under
               various conditions. They are denoted by the following symbols.


               • Base: represents basic complexity, containing ten formulas.
               • Noise: includes datasets characterized by the presence of noise.
               • Num: involves datasets with sparse data.
   40   41   42   43   44   45   46   47   48   49   50