Surrogate Reliability and Applicability Domain in ML-Driven CO₂ Adsorption Screening of Metal–Organic Frameworks

Kimia Shahbazi and Mark Thachuk

Department of Chemistry, University of British Columbia

Machine learning surrogates enable rapid screening of metal–organic frameworks (MOFs) for CO2 capture, yet their reliability on candidates outside the training distribution remains poorly characterized. We train Tabular Transformer regressors on 324,107 ToBaCCo-generated hypothetical MOFs to predict CO2 uptake and heat of adsorption at P = 0.15 bar and T = 298 K. Each MOF is described by 15 mixed descriptors: geometric properties from Zeo++ (accessible surface area, pore volume, void fraction, pore diameters, channel/pocket counts) and categorical identifiers (metal node, organic linker, topology). The uptake model achieves R2 = 0.8765 ± 0.0026 (RMSE = 0.1872 ± 0.0019 mmol/g) by 5-fold stratified cross-validation, outperforming linear regression, random forest, XGBoost, MLP, and CatBoost baselines.
SHAP feature attribution reveals a physically interpretable relationship between pore geometry and predicted uptake: MOFs with narrow micropores (low void fraction) adsorb more CO2 per gram than those with large, open pores. This reflects confinement-dominated adsorption at low partial pressure, where gas molecules interact with multiple pore walls simultaneously, producing stronger binding than in wide pores far from saturation. Linker identity modulates this effect, producing substantial variation at equivalent void fractions — indicating that linker chemistry is a major design lever alongside pore geometry.
We deploy the surrogate on 34 generated candidates and validate against RASPA2 Grand Canonical Monte Carlo (GCMC) at 298 K, revealing two independent failure modes. First, 76% of candidates lie outside the training distribution by Mahalanobis distance (a measure of how far descriptor values fall from the training population). Error severity depends on topology: well-represented topologies (nbo, sra) show modest errors (~0.25 mmol/g), while under-represented topologies (bcu) show much larger errors (~1.27 mmol/g). Second, some statistically in-domain candidates still show large errors due to a descriptor–simulation mismatch. The Zeo++ analysis uses a hard-sphere probe to assess CO2 accessibility and reports zero accessible surface area and volume for these structures, indicating no CO2-accessible porosity. However, the GCMC force-field simulation finds that CO2 can energetically penetrate these borderline micropores, yielding non-zero adsorption. This mismatch is invisible to distribution-based screening and produces errors of 1.0–1.7 mmol/g.
A key question is whether the surrogate can estimate its own uncertainty on novel candidates. We test this with conformal prediction, which constructs intervals calibrated to contain the true value 90% of the time. On held-out data these intervals achieve 90% coverage, but on generated candidates coverage collapses to 40%. Reweighting to correct for the distribution shift recovers coverage to 64% but falls short of the target, showing that post-hoc uncertainty estimates alone are insufficient. Two-stage pre-screening — first verifying that Zeo++ reports meaningful CO2-accessible porosity, then checking that the candidate lies within the training distribution — provides a more robust reliability guarantee. The only validated-reliable design rule is a Cu-paddlewheel pcu framework with ethynyl-type linkers, consistent with the confinement physics identified by SHAP.

Back to List of Abstracts