Thinking Fast and Slow Enables Stochastic Guidance
In the last post, I covered how a flat recurrent architecture outperformed fast/slow hierarchy at the 300K parameter scale. Here we explore how restoring the two-timescale carry converts GRAM-style hidden-space stochasticity from redundant with input augmentation into orthogonal to it, compounding +5 to 9 points at pass@1000. However, GRAM's training tax on the backbone eats the win, so the practical levers are augmentation and embedding width.
Read more