Page 112 - Read Online
P. 112

Page 4 of 31                                                        Li et al. J. Mater. Inf. 2026, 6, 10





































               Figure 1. Framework of ChemMiner [22] . Data extraction systems typically adopt multi-agent systems, where different functional agents are
               responsible for specific forms of data extraction to ensure the comprehensiveness and accuracy of data extraction. Reproduced from
               “ChemMiner: A Large Language Model Agent System For Chemical Literature Data Mining”, arXiv:2402.12993, with permission of the
               authors [22] . LLM: Large language model; OCR: optical character recognition; JSON: JavaScript Object Notation.


               multimodal analysis. This framework enables highly accurate end-to-end extraction of nanomaterial
               structural properties (e.g., chemical formulas, crystal systems, and surface characteristics) as well as kinetic
               parameters of nanozymes . Multicrossmodal LLM-agent coordinates a team of specialized agents to extract
                                    [23]
               and integrate multimodal data from sources such as literature, simulation videos, and microscopy images
               into a shared embedding space. Through cross-modal fusion and visual information interpretation, it
               achieves unified reasoning of multimodal data, thereby improving extraction accuracy . Eunomia
                                                                                               [24]
               autonomously plans and executes the creation of structured material datasets in research fields such as
               solid-state electrolytes (SSEs) and metal-organic frameworks (MOFs) from scientific literature, and extracts
               design guidelines for materials with specific properties . MechGPT employs a fine-tuned LLM to retrieve
                                                              [25]
               and connect cross-scale knowledge from literature, constructs knowledge graphs, and extracts structured
               insights to provide support for knowledge retrieval and hypothesis generation .
                                                                                [26]

               To overcome the limitations of single-modal approaches, the Descriptive Interpretation of Visual Expression
               (DIVE) multi-agent workflow focuses on extracting experimental data from figures and tables in scientific
               literature. When parsing data related to solid-state hydrogen storage materials, its MAS significantly
               improves both extraction accuracy and coverage compared to directly using multimodal models. Based on
               4,000 research articles, DIVE successfully constructed a database containing over 30,000 data entries and
               achieved reverse design of materials within two minutes . In terms of enhancing the reliability and
                                                                  [27]
               generalization capability of agents, SLM-MATRIX introduces a multi-path collaborative reasoning and
               verification framework based on small language models (SLMs). Through three pathways, multi-agent
               collaboration, generator-discriminator mechanism, and cross-validation, it achieves high-precision
               extraction in material names, numerical values, and physical units, providing an efficient solution for
               resource-constrained scenarios .
                                         [28]
               These studies collectively demonstrate that MAS, by emulating the role specialization of research teams,
   107   108   109   110   111   112   113   114   115   116   117