2026-07-18 20:04:20
First, let’s talk about “grokking”.
In 2022, OpenAI published a paper showing that if you train a model on a simple dataset
(for instance, a simple mathematical operation like division),
and keep training it long after the training looks like it’s stalled out,
the model will suddenly make a massive jump in capability.
Why does this work?
The first stage of training is like rote memorization:
the model has to compress as much of the training data as…













