Sparse and interpretable neural network architecture for scientific data
University of Toronto Scarborough
In deep learning, "grokking" describes a delayed phase transition where a model, after a long period of memorizing training data, suddenly discovers a generalizable solution and test accuracy spikes. To date, this phenomenon has been treated as a curiosity observed primarily in clean, algorithmic tasks (e.g., modular arithmetic). The prevailing dogma in the physical sciences is that grokking is inaccessible for scientific datasets, which are characteristically small, high-dimensional, and inherently noisy. Consequently, most models in computational chemistry—from property predictors to Machine Learning Interatomic Potentials (MLIPs)—operate in the "memorization" regime, resulting in brittle performance when extrapolating to new out-of-distribution chemical spaces.
Here, we challenge this assumption. We demonstrate a method to accelerate the onset of grokking on arbitrary, noisy datasets. We show that models are compelled to abandon "shortcut learning" (memorizing noise) in favor of "structural understanding" (learning the underlying physics). We demonstrate the implications of this approach for materials science: enabling the training of robust MLIPs that maintain stability outside their training domain and extracting physical laws from sparse, noisy experimental data.
This work suggests that the "Small Data" problem in science is not a lack of information, but a failure of the training dynamics to access the generalized solution.
