SOTAVerified

Blog

Thinking Fast and Slow Enables Stochastic Guidance

In the last post, I covered how a flat recurrent architecture outperformed fast/slow hierarchy at the 300K parameter scale. Here we explore how restoring the two-timescale carry converts GRAM-style hidden-space stochasticity from redundant with input augmentation into orthogonal to it, compounding +5 to 9 points at pass@1000. However, GRAM's training tax on the backbone eats the win, so the practical levers are augmentation and embedding width.

Read more

Hierarchical Reasoning Doesn't Beat Recurrence at 300K Parameter Scale

A replication study of HRM/TRM-style fast/slow hierarchical recurrence on small-scale ARC. The two-timescale split looked like a clean win and then lost every control, tying a flat transformer on single-shot accuracy and losing on pass@K once compute is matched. Here's what carried over from the papers and what didn't at 300K parameters on 10x10 grids, and the open question that sets up where I'm going next.

Read more