MOF-HydrA: A lightweight and robust machine learning model for reconstructing missing protons in crystal structures of metal-organic frameworks
University of Ottawa
Experimental 3D crystal structures are often missing proton positions due to the low electron density of hydrogen atoms. The positions of missing protons can be easily determined for typical organic molecules with heuristic algorithms; however, no robust methods exist for reconstructing missing protons near metal centers in organometallic compounds or crystalline solids such as metal-organic frameworks (MOFs). This is problematic when preparing structures for atomistic simulations since missing protons would lead to incorrect total charge or metal oxidation states of the compounds. Recent studies on structure errors of common MOF databases indicated that missing protons is a common error category. This work presents MOF-HydrA (HYDRrogen Assigner), a machine learning (ML) model based on the graph attention network that simultaneously predicts the number of missing protons for each atomic sites via ordinal classification and reconstructs the positions of the missing protons via edge regression coupled with post hoc trilateration. Most proton assignment tools can restore protons on organic linkers and terminal ligands but often fail to recognize missing protons near metal centers, particularly in bridging sites (e.g., µ-OH and µ-OH₂), whereas MOF-HydrA aims to restore missing protons for all atoms including the bridging sites between metal centers. Preliminary results achieved >99% accuracy for identifying the correct number of missing protons on each atomic site; and the model was capable of fully restoring all missing protons for >96% of MOF structures, with an average RMSD of <0.2 Å vs reference proton coordinates determined from DFT geometry optimization. When tested on structures with missing protons only on metal bridging sites, the model performed better reaching an accuracy of >98%. Currently, the model is still under active development to improve proton position precisions without relying on compute-intensive architectures (e.g., diffusion or large language models) or post hoc force field optimizations. Moreover, our proton assignment workflow offers flexible routines, such as fully autonomous restoration (replying solely on model predictions) or restoration on request (add designated number of protons based on a provided empirical formula).