Page 12 - Read Online
P. 12
Page 6 of 17 Schertzer et al. J. Mater. Inf. 2025, 5, 5 https://dx.doi.org/10.20517/jmi.2024.69
previous work, we explored a large space of PEMs using ML models for proton conductivity, WU, and
[21]
several other properties .
In the present study, we leverage ML to develop predictive models based on curated AEM data from the
literature, aiming to design robust polymers with high anion conductivity, low WU, and low SR. The
informatics-based approach, outlined in Figure 1C, comprises four distinct steps. First, we identify critical
properties and their optimal desired values based on design of experiments (DOE) and literature guidelines.
Our objective is to identify innovative copolymers that exhibit hydroxide conductivity greater than 100 mS/
cm, equilibrium WU less than 35 wt%, and equilibrium SR below 50%, with SR and WU serving as proxies
for degradation and indicators of anion conductivity. We use WU and SR as proxies for stability because of
data scarcity and because polymers with lower water adsorption are less likely to experience degradation
due to unwanted swelling. Next, we train ML models to predict these key properties using measured
property data curated from the literature. Then, a large set of candidate copolymers built using known co-
monomers was generated to define the search space. Finally, we use the ML models to predict the key
properties of these candidates and screen them to find the optimal candidates for further exploration.
Future work will address mechanical and chemical stability directly by incorporating time-dependent
conductivity and mechanical property retention data into a holistic AEM dataset.
MATERIALS AND METHODS
Dataset
The AEM dataset comprises approximately 1,100 data points collected from the literature, involving 108
unique monomers, including information on hydroxide conductivity, WU, and SR for polymer
electrolytes [5,22-60] . The chemical space includes copolymers; thus, we record both the chemistry information
for each monomer and the relative composition of each monomer in the copolymer. For each data point,
we collect an associated measurement of temperature, RH, and experimental IEC. Figure 2 shows the
distribution of various properties in the dataset, the correlation between each pair of properties, and the
distribution of variables used as descriptors in the ML models. Table 1 shows a fragment of the dataset,
including its structure. The full data set and training splits are available on the polyVERSE GitHub.
Feature engineering
In previous work, polymers were characterized using the Co-Polymer Genome fingerprinting scheme,
which has been demonstrated to accurately model the properties of copolymers . Each monomer repeat
[61]
unit is decomposed into a hierarchical feature vector, which serves as a numerical representation of the
chemical building blocks. We then perform a linear combination of the fingerprint vectors of each co-
monomer weighted by their respective composition within the copolymer. Each fingerprint vector is
concatenated with the sample temperature, RH, and IEC, as defined by Equations (1-3), respectively.
ML algorithm
The fingerprint vectors are fed through a Gaussian process regression (GPR) framework for parameter
optimization which allows us to map the chemical structure and environmental conditions to the desired
properties. This allows us to predict the properties of candidate polymers later. We use a radial basis
function (RBF) kernel multiplied by white noise and linear kernels to train the property prediction
models [61-63] .
We train both single-task (ST) and multi-task (MT) models to predict multiple properties of interest. In
[61]
ST learning, each model is trained to predict only one property at a time, while in MT learning, a single
model is trained to predict several properties simultaneously. To assess the predictive capabilities of both
approaches, we divide the data into training and test sets in three distinct ways:

