Page 15 - Read Online
P. 15
Schertzer et al. J. Mater. Inf. 2025, 5, 5 https://dx.doi.org/10.20517/jmi.2024.69 Page 9 of 17
Figure 6. Comparison of predicted vs. true values of hydroxide conductivity (mS/cm) for PS and CS datasets using ST and MT models
for a single instance of 80% training data and 20% testing data. (A) and (B) represent ST and MT predictions for the PS dataset, while
(C) and (D) represent ST and MT predictions for the CS dataset. The diagonal dashed line indicates the ideal case in which predicted
2
values perfectly match the true values. The RMSE and R values for each model are provided, demonstrating improved predictive
performance with the MT models in both the PS and CS splits. The error bars represent prediction uncertainty, highlighting the
increased confidence of the MT model. PS: Polymer split; CS: composition split; ST: single-task; MT: multi-task; RMSE: root mean
2
squared error; R : coefficient of determination.
properties of novel polymers without prior examples of the corresponding chemistry (e.g., extrapolation to
novel polymer classes). The MT model outperformed the ST model for all split types and test-train ratios,
indicating its ability to generalize across shared structural features despite the absence of direct training
examples. This suggests that augmenting the dataset with additional monomer chemistries could
significantly enhance the model’s extrapolative capabilities. The superior performance of the MT model
implies that capturing deeper polymer structure-property relationships leads to better generalization with
more comprehensive data.
Compared with the PS split, the CS, where the test set included new combinations or compositions of
known monomers, represented an easier challenge because the model could rely on learned knowledge
about individual monomers and their interactions. The MT model outperformed the ST model in this case
as well. This is due to the ability of the MT model to leverage the shared information across tasks, allowing
it to make more accurate predictions for unseen copolymer compositions by drawing on correlations
between different properties. The strong performance of the model in this split demonstrates its proficiency

