Page 112 - Read Online
P. 112
Page 4 of 31 Li et al. J. Mater. Inf. 2026, 6, 10
Figure 1. Framework of ChemMiner [22] . Data extraction systems typically adopt multi-agent systems, where different functional agents are
responsible for specific forms of data extraction to ensure the comprehensiveness and accuracy of data extraction. Reproduced from
“ChemMiner: A Large Language Model Agent System For Chemical Literature Data Mining”, arXiv:2402.12993, with permission of the
authors [22] . LLM: Large language model; OCR: optical character recognition; JSON: JavaScript Object Notation.
multimodal analysis. This framework enables highly accurate end-to-end extraction of nanomaterial
structural properties (e.g., chemical formulas, crystal systems, and surface characteristics) as well as kinetic
parameters of nanozymes . Multicrossmodal LLM-agent coordinates a team of specialized agents to extract
[23]
and integrate multimodal data from sources such as literature, simulation videos, and microscopy images
into a shared embedding space. Through cross-modal fusion and visual information interpretation, it
achieves unified reasoning of multimodal data, thereby improving extraction accuracy . Eunomia
[24]
autonomously plans and executes the creation of structured material datasets in research fields such as
solid-state electrolytes (SSEs) and metal-organic frameworks (MOFs) from scientific literature, and extracts
design guidelines for materials with specific properties . MechGPT employs a fine-tuned LLM to retrieve
[25]
and connect cross-scale knowledge from literature, constructs knowledge graphs, and extracts structured
insights to provide support for knowledge retrieval and hypothesis generation .
[26]
To overcome the limitations of single-modal approaches, the Descriptive Interpretation of Visual Expression
(DIVE) multi-agent workflow focuses on extracting experimental data from figures and tables in scientific
literature. When parsing data related to solid-state hydrogen storage materials, its MAS significantly
improves both extraction accuracy and coverage compared to directly using multimodal models. Based on
4,000 research articles, DIVE successfully constructed a database containing over 30,000 data entries and
achieved reverse design of materials within two minutes . In terms of enhancing the reliability and
[27]
generalization capability of agents, SLM-MATRIX introduces a multi-path collaborative reasoning and
verification framework based on small language models (SLMs). Through three pathways, multi-agent
collaboration, generator-discriminator mechanism, and cross-validation, it achieves high-precision
extraction in material names, numerical values, and physical units, providing an efficient solution for
resource-constrained scenarios .
[28]
These studies collectively demonstrate that MAS, by emulating the role specialization of research teams,

