Page 177 - Read Online
P. 177
Page 2 of 22 Hei et al. J. Mater. Inf. 2026, 6, 15
the density-related performance decline mainly to semantic overlap and syntactic complexity, with upstream
extraction errors more prominent under sparse supervision and allocation errors concentrated in structurally
complex templates. This approach delivers precise, structured outputs suitable for downstream analysis and offers a
domain-adaptable alternative to prompt-based large models when strict correctness is required.
INTRODUCTION
The traditional “trial-and-error” paradigm in the development of new materials is time-consuming and
expensive. However, with emerging technologies such as artificial intelligence, a dual-driven scientific
artificial intelligence strategy that combines both models and data has been widely applied in the design and
development of novel materials , as well as in elucidating the interrelationships between structure and
[1,2]
performance . This paradigm has unveiled considerable potential for the efficient design and optimization
[3-5]
of materials [6-9] . Undoubtedly, data remain a foundational and critical element in achieving the
aforementioned objectives. However, for certain material properties, particularly those pertaining to
performance under service conditions, available data remain scarce. The acquisition of high-quality data by
labor incurs significant costs, making it difficult to meet the requirements for training models that demand
high accuracy and robust performance.
The scientific literature encompasses a vast array of peer-reviewed, high-quality, and relatively reliable data,
serving as a crucial resource. However, much of this information exists in unstructured text within diverse
sources, such as textbooks, material handbooks, patents, and research articles, rendering it unsuitable for
direct use as structured data. Consequently, techniques for the automatic extraction of materials-related data
and information present a promising opportunity to construct extensive databases for machine learning and
data-driven methodologies.
Natural language processing (NLP) is an interdisciplinary field that integrates linguistics and artificial
intelligence. One of its most prevalent applications is the automated extraction of structured information,
facilitating the efficient mining of data from semi-structured tables and unstructured texts . Numerous
[10]
studies have reported various methodologies for mining structured information within specific domains of
materials science [11,12] , with a general trend moving from generic approaches to specialized techniques.
Initially, these processes relied on traditional processing pipelines, generally including named entity
recognition (NER) and relation extraction , such as ChemDataExtractor for chemical information
[13]
[14]
extraction. Specific techniques in these workflows include look-ups, rule-based methods , and machine
[13]
learning approaches, which have gradually shifted toward more sophisticated deep neural models , such as
[15]
long short-term memory (LSTM) networks . Due to the excellent performance of the pre-training and
[13]
fine-tuning paradigm, many works now rely on pre-trained models for information extraction. By
pre-training on a substantial corpus of materials-related literature, these models enable a deeper analysis of
relationships among chemical compositions, structural features, and corresponding properties. Many
variations of bidirectional encoder representations from transformers (BERT) , such as MatSciBERT ,
[17]
[16]
MaterialsBERT , BioBERT , and BatteryBERT , have shown improved performance on downstream
[20]
[19]
[18]
tasks. Large language models (LLMs), such as Generative Pre-trained Transformer (GPT)-4 and GPT-5 [21,22] ,
Large Language Model Meta AI (LLaMA) , Pathways Language Model (PaLM) , Google’s Gemini
[24]
[23]
model , and Fine-tuned LAnguage Net (FLAN) , also exhibit great understanding of textual information.
[25]
[26]
When prompted appropriately, LLMs can extract targeted information and data from literature. For
example, Dagdelen et al. used multiple LLMs to extract structured information .
[27]

