AI That Learns to Generalize Later: The Mystery of Grokking
Research note on Xu, Vardi & Safran, “To Grok Grokking: Provable Grokking in Ridge Regression” , ICML 2026. A neural network is trained on a task. On the training data it performs well almost immediately. On new, unseen data it fails. Then, after a long period during which nothing appears to change, it begins to produce correct answers on data it has never seen. This phenomenon is known in machine learning as grokking . The core finding: a team of researchers has rigorously proven all three stages of grokking in a simple linear model and shown how the effect can be controlled through training parameters. The problem The common assumption is that longer training leads to better generalization. Grokking contradicts this assumption. The process typically unfolds in three stages: Overfitting. The model memorizes the training data. Plateau. Performance on new data remains poor for a long period, with little visible change. Generalization. The model begins to pro...