Do Erroneous Metal-Organic Frameworks Impact the Predictive Abilities of Machine Learning Models for Different Gas Isotherms?

Sydney L. O'Shaughnessy, Marco Gibaldi, Jun Luo, and Tom K. Woo

Department of Chemistry and Biomolecular Sciences, University of Ottawa, Ottawa, Ontario, Canada K1N 6N5

Metal-organic frameworks (MOFs) have consistently attracted the attention of material scientists due to their diverse and tunable pore properties, which have made them viable candidates for many industrial chemical processes. To accelerate the identification and development of high-performing MOFs, “computationally ready” databases containing experimental and computer-generated structures have been curated, providing the raw data for high-throughput and machine-learning (ML) driven studies. Most of these studies have focused on screening or predicting various gas adsorption properties. Recent analysis found that CoRE-2019, the most-cited MOF database, has a structural error rate of \(~50%^{ 1 }\). In other words, about half of the structures contain some sort of structural error, such as missing atoms and metal centres with impossible oxidation states. Unfortunately, hundreds of computational MOF studies have used this database “as is” to build and deploy machine learning models.
The goal of this work was to assess how machine learning models trained and validated on “good” and “bad” structures would predict the adsorption properties of \(CO_{ 2 }\), Xe, \(H_{ 2 }\), and \(CH_{ 4 }\). To do so, six supervised regression models were trained, validated and tested on datasets consisting exclusively of chemically valid (“good”) or chemically invalid MOFs (“bad”), or a 50:50 mix of these two types of structures (“mixed”).
A reduction of up to 15% in the Pearson Correlation Coefficient (PCC) was observed when comparing the performance of models trained and tested on “mixed” datasets with those trained and tested exclusively on “good” sets. This effect was most pronounced for models predicting \(H_{ 2 }\) selectivity and \(CO_{ 2 }\) uptake, with an average reduction of of PCC=0.07 to PCC=0.08. The loss was least stark for models predicting Xe uptake and \(CH_{ 4 }\) working capacity, with the decline in accuracy amounting to only 0.4% and 0.3%, respectively. To this effect, it was noted that the performance of models trained on “mixed” sets could be improved by up to 7% if they were tested on sets of “good” structures instead of “mixed” ones. Nonetheless, this improvement depended heavily on the target.
To summarize, chemically invalid metal-organic frameworks can negatively impact the performance of ML models. In fact, a noticeable decline in performance is observed when training and testing on “mixed” datasets relative to “good” datasets. This means that the hundreds of previous studies that utilized the CoRE-2019 database could be underreporting the accuracy for “good” top performers and structures.

\(^{ 1}\)White, A. J. et al. High Structural Error Rates in “Computation-Ready” MOF Databases Discovered by Checking Metal Oxidation States. J. Am. Chem. Soc. 2025, 147 (21), 17579–17583. https://doi.org/10.1021/jacs.5c04914

Back to List of Abstracts