Page 106 - Read Online
P. 106
Kobayashi et al. J. Mater. Inf. 2025, 5, 50 https://dx.doi.org/10.20517/jmi.2025.44 Page 5 of 7
databases. As Trunschke found from her experience in the field of catalysis, there is currently not enough
[16]
data for standard ML techniques to be effectively used . Such data are hard to find, often only available
through papers, sometimes in the form of figures, and negative results often go unreported. She showed
examples of how the SISSO method can work well with sparse data provided it is “clean”, i.e., well-
characterized data . Consequently, her recent efforts have been concentrated on strategies for data
[17]
acquisition, storage and use . The importance of data sharing was found to be paramount, though there
[18]
was recognition of possible limitations due to proprietary concerns, especially in industry. However, even
with freely shared data there were still questions of trustworthiness and making sense of it. This
immediately raised the thorny issue of setting standards, especially for interoperability, and brought out a
lot of differing opinions, essentially who, what and why. There is an inherent belief that for data to be
shareable it needs to be in a certain format following certain rules, such as in the philosophy of FAIR
data . However, who decides what rules to set? In reality, it is hard to get people to agree to which standard
[19]
to adopt. A dominant publisher such as Materials Project or the Protein Data Bank has enough driving
[20]
[21]
force for people to follow their lead, but, on the whole, most people want to do their own thing. The analogy
was given of electrical plugs around the world. However, as with electrical plugs, why should people need to
be forced to adhere to the one standard when it may be less troublesome to just work with converters.
Building community
As part of the discussion of setting standards, the primary question was “who sets the rules?” and it was
agreed that it has to be done by community consensus rather than be imposed by some governing authority
as happens all too often. It became clear that the materials science community needed to talk, for which the
Workshop provided a good forum, but not how to start the conversation. There are the beginnings of
communities being built in materials science through data and tools platforms, such as Materials Project ,
[20]
Materials Cloud and DP Technology . Also, there are community efforts to establish ontology for
[22]
[23]
sharing data, such as NFDI4Cat . However, these are not as established as the Molecular Sciences Software
[24]
Institute (MolSSI) nexus for science, education and co-operation for the global computational molecular
[25]
sciences community . The success of MolSSI is rooted in community engagement. Its origins were in
[26]
identifying common needs and letting standards grow naturally. Notably, unlike the aforementioned
materials science platforms, it encompasses a range of software packages and tools and works with the
developers as being the drivers of what people will end up using.
CONCLUSION
The International Workshop on DCTMD demonstrated that work in the area of AI/ML in materials science
is still going strong and producing new insights. Though there has yet to be an equivalent AlphaFold
breakthrough moment there have been many small successes or achievements. AI/ML has improved greatly
the success rate, saving time and reducing cost, by guiding iterative high-throughput experiments along the
whole process of materials development, though humans are still needed in the loop. And there is still a lot
of work for humans to do. There remains a strong feeling that AI/ML is not yet at a stage to be trusted in
isolation and theory and modelling are still the way forward. Even so, AI/ML is proving a useful addition to
the toolkit augmenting fundamental theory and experiment. AI can accelerate existing computational
prediction, bridge the gaps across multiple length/time scales, and even map the relationships between
structure-property relationships without an explicitly defined underlying theory. The rational design of
physically meaningful ML features to do this is essential to materials science. More materials science-
adapted ML algorithms need to be developed to tackle the challenge of scarce materials data. For AI/ML to
work well, high-quality data that are well-curated and fully characterized by metadata - so that they are
accessible, shareable, and reusable - are essential. The community still needs to come together to achieve
this, and conferences such as these provide a good way to move forward.

