Page 177 - Read Online
P. 177

Page 2 of 22                                                       Hei et al. J. Mater. Inf. 2026, 6, 15






               the   density-related   performance   decline   mainly   to   semantic   overlap   and   syntactic   complexity,   with   upstream
               extraction   errors   more   prominent   under   sparse   supervision   and   allocation   errors   concentrated   in   structurally
               complex templates. This approach delivers precise, structured outputs suitable for downstream analysis and offers a
               domain-adaptable alternative to prompt-based large models when strict correctness is required.



               INTRODUCTION
               The traditional “trial-and-error” paradigm in the development of new materials is time-consuming and
               expensive. However, with emerging technologies such as artificial intelligence, a dual-driven scientific
               artificial intelligence strategy that combines both models and data has been widely applied in the design and
               development of novel materials , as well as in elucidating the interrelationships between structure and
                                          [1,2]
               performance . This paradigm has unveiled considerable potential for the efficient design and optimization
                          [3-5]
               of materials [6-9] . Undoubtedly, data remain a foundational and critical element in achieving the
               aforementioned objectives. However, for certain material properties, particularly those pertaining to
               performance under service conditions, available data remain scarce. The acquisition of high-quality data by
               labor incurs significant costs, making it difficult to meet the requirements for training models that demand
               high accuracy and robust performance.


               The scientific literature encompasses a vast array of peer-reviewed, high-quality, and relatively reliable data,
               serving as a crucial resource. However, much of this information exists in unstructured text within diverse
               sources, such as textbooks, material handbooks, patents, and research articles, rendering it unsuitable for
               direct use as structured data. Consequently, techniques for the automatic extraction of materials-related data
               and information present a promising opportunity to construct extensive databases for machine learning and
               data-driven methodologies.


               Natural language processing (NLP) is an interdisciplinary field that integrates linguistics and artificial
               intelligence. One of its most prevalent applications is the automated extraction of structured information,
               facilitating the efficient mining of data from semi-structured tables and unstructured texts . Numerous
                                                                                              [10]
               studies have reported various methodologies for mining structured information within specific domains of
               materials science [11,12] , with a general trend moving from generic approaches to specialized techniques.
               Initially, these processes relied on traditional processing pipelines, generally including named entity
               recognition (NER) and relation extraction , such as ChemDataExtractor  for chemical information
                                                     [13]
                                                                                 [14]
               extraction. Specific techniques in these workflows include look-ups, rule-based methods , and machine
                                                                                            [13]
               learning approaches, which have gradually shifted toward more sophisticated deep neural models , such as
                                                                                                 [15]
               long short-term memory (LSTM) networks . Due to the excellent performance of the pre-training and
                                                     [13]
               fine-tuning paradigm, many works now rely on pre-trained models for information extraction. By
               pre-training on a substantial corpus of materials-related literature, these models enable a deeper analysis of
               relationships among chemical compositions, structural features, and corresponding properties. Many
               variations of bidirectional encoder representations from transformers (BERT) , such as MatSciBERT ,
                                                                                                        [17]
                                                                                   [16]
               MaterialsBERT , BioBERT , and BatteryBERT , have shown improved performance on downstream
                                                          [20]
                                       [19]
                            [18]
               tasks. Large language models (LLMs), such as Generative Pre-trained Transformer (GPT)-4 and GPT-5 [21,22] ,
               Large Language Model Meta AI (LLaMA) , Pathways Language Model (PaLM) , Google’s Gemini
                                                                                        [24]
                                                     [23]
               model , and Fine-tuned LAnguage Net (FLAN) , also exhibit great understanding of textual information.
                    [25]
                                                        [26]
               When prompted appropriately, LLMs can extract targeted information and data from literature. For
               example, Dagdelen et al. used multiple LLMs to extract structured information .
                                                                                 [27]
   172   173   174   175   176   177   178   179   180   181   182