Page 197 - Read Online
P. 197
Page 22 of 22 Hei et al. J. Mater. Inf. 2026, 6, 15
16. Devlin, J.; Chang, M. W.; Lee, K.; Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In
Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language
Technologies, Minneapolis, USA, June 2, 2019; Association for Computational Linguistics; Vol. 1. pp. 4171-86. DOI
17. Gupta, T.; Zaki, M.; Krishnan, N. M. A.; Mausam, . MatSciBERT: a materials domain language model for text mining and information
extraction. npj. Comput. Mater. 2022, 8, 102. DOI
18. Shetty, P.; Rajan, A. C.; Kuenneth, C.; et al. A general-purpose material property data extraction pipeline from large polymer corpora
using natural language processing. npj. Comput. Mater. 2023, 9, 52. DOI
19. Lee, J.; Yoon, W.; Kim, S.; et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining.
Bioinformatics 2020, 36, 1234-40. DOI PubMed PMC
20. Huang, S.; Cole, J. M. BatteryBERT: a pretrained language model for battery database enhancement. J. Chem. Inf. Model. 2022, 62,
6365-77. DOI PubMed PMC
21. Brown, T. B.; Mann, B.; Ryder, N.; et al. Language models are few-shot learners. arXiv 2020, arXiv:2005.14165. Available online: https
://doi.org/10.48550/arXiv.2005.14165. (accessed 23 Mar 2026).
22. OpenAI, Achiam, J.; Adler, S.; et al. GPT-4 technical report. arXiv 2023, arXiv:2303.08774. Available online: https://doi.org/10.48550/a
rXiv.2303.08774. (accessed 23 Mar 2026).
23. Touvron, H.; Lavril, T.; Izacard, G.; et al. LLaMA: open and efficient foundation language models. arXiv 2023, arXiv:2302.13971.
Available online: https://doi.org/10.48550/arXiv.2302.13971. (accessed 23 Mar 2026).
24. Chowdhery, A.; Narang, S.; Devlin, J.; et al. PaLM: scaling language modeling with pathways. arXiv 2022, arXiv:2204.02311. Available
online: https://doi.org/10.48550/arXiv.2204.02311. (accessed 23 Mar 2026).
25. Team, G.; Anil, R.; Borgeaud, S.; et al. Gemini: a family of highly capable multimodal models. arXiv 2023, arXiv:2312.11805.
Available online: https://doi.org/10.48550/arXiv.2312.11805. (accessed 23 Mar 2026).
26. Wei, J.; Bosma, M.; Zhao, V. Y.; et al. Finetuned language models are zero-shot learners. arXiv 2021, arXiv:2109.01652. Available
online: https://doi.org/10.48550/arXiv.2109.01652. (accessed 23 Mar 2026).
27. Dagdelen, J.; Dunn, A.; Lee, S.; et al. Structured information extraction from scientific text with large language models. Nat. Commun.
2024, 15, 1418. DOI PubMed PMC
28. Vinyals, O.; Fortunato, M.; Jaitly, N. Pointer networks. In NIPS'15: Proceedings of the 29th International Conference on Neural
Information Processing Systems, Montreal, Canada, December 7-12, 2015; MIT Press: Cambridge, Massachusetts, United States, 2015;
Vol. 28. pp. 2692-700. DOI
29. Vaswani, A.; Shazeer, N.; Parmar, N.; et al. Attention is all you need. arXiv 2017, arXiv:1706.03762. Available online: https://doi.org/1
.48550/arXiv.1706.03762. (accessed 23 Mar 2026).
30. Sun, F.; Jiang, P.; Sun, H.; Pei, C.; Ou, W.; Wang, X. Multi-source pointer network for product title summarization. arXiv 2018,
arXiv:1808.06885. Available online: https://doi.org/10.48550/arXiv.1808.06885. (accessed 23 Mar 2026).
31. Anthropic. The Claude 3 model family: opus, sonnet, haiku. https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc61885
7627/Model_Card_Claude_3.pdf. (accessed 2026-03-23).
32. Gemini Team Google. Gemini 1.5: unlocking multimodal understanding across millions of tokens of context. arXiv 2024,
arXiv:2403.05530. Available online: https://doi.org/10.48550/arXiv.2403.05530. (accessed 23 Mar 2026).
33. Grattafiori, A.; Dubey, A.; Jauhri, A.; et al. The Llama 3 herd of models. arXiv 2024, arXiv:2407.21783 Available online: https://doi.org/
10.48550/arXiv.2407.21783. (accessed 23 Mar 2026).
34. Ji, Z.; Lee, N.; Frieske, R.; et al. Survey of hallucination in natural language generation. ACM. Comput. Surv. 2023, 55, 1-38. DOI
35. Farquhar, S.; Kossen, J.; Kuhn, L.; Gal, Y. Detecting hallucinations in large language models using semantic entropy. Nature 2024, 630,
625-30. DOI PubMed PMC
36. Hei, M.; Liu, Q.; Zhang, X. Enhancing Information Extraction from Low-sample Materials Science Literature by Transfer Learning. In
2024 10th International Conference on Big Data and Information Analytics (BigDIA), Chiang Mai, Thailand, Oct 25-28, 2024; IEEE,
2024; pp. 736-41. DOI
Disclaimer/Publisher’s Note: All statements, opinions, and data contained in this publication are solely those of the individual author(s) and
contributor(s) and do not necessarily reflect those of OAE and/or the editor(s). OAE and/or the editor(s) disclaim any responsibility for harm to
persons or property resulting from the use of any ideas, methods, instructions, or products mentioned in the content.
© The Author(s) 2026. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License
(https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, sharing, adaptation, distribution and
reproduction in any medium or format, for any purpose, even commercially, as long as you give appropriate credit to the original author(s) and
the source, provide a link to the Creative Commons license, and indicate if changes were made.

