← All posts

Thinking Twice Can Make You Dumber: The Test-Time Compute Playbook Behind September's Smartest Models

2026-09-27 · AI infrastructure / ML architectures (test-time compute scaling)

Diagram 1

For a decade, making AI smarter meant one thing: train bigger. More data, more GPUs, more parameters — that was the entire playbook. Then researchers found a third axis nobody was fully exploiting: make the model think longer after the question arrives.

The payoff is absurd. A model 4x smaller, given the right thinking strategy, matched or beat a model up to 14x its size on hard reasoning tasks (Snell et al., 2024). You read that right: no extra training, no extra weights — just thinking.

This is test-time compute. It's the reason every flagship model released this month ships a "thinking" mode, and the reason September's smartest models are defined less by their parameter counts than by how they spend their time.

ELI5: studying vs. thinking during the exam

There are three ways to make someone smarter:

  1. Study longer before the test — pre-training. Years of school, trillions of words.
  2. Bring a bigger brain — parameters. More capacity to hold what was learned.
  3. Think longer during the test — test-time compute. Re-read the question, try two approaches, check your work.

For years, only (1) and (2) got funding. Then OpenAI's o1 showed that on math competitions, accuracy climbs roughly log-linearly with thinking tokens — and the whole industry pivoted. In 2026, reasoning isn't a special mode anymore. It's standard equipment: Claude Opus 5.5, GPT-5.6 Sol, Grok 4.7, DeepSeek-V4-Pro, Qwen3.8, GLM-5.3 — all "think first, then answer" by default.

How it works: four ways to spend compute after the question arrives

Diagram 2

1. Sequential thinking (the long chain of thought). The model just… keeps writing before answering. Propose, reflect, backtrack, try again. Simple, and it's the foundation everything else builds on. The catch: a single reasoning request can now generate 16,000–128,000 internal thinking tokens before the final answer — 10–100x more tokens than the old one-shot models.

2. Self-consistency (majority vote). Ask N times, take the most common answer. On grade-school math (GSM8K), this lifted PaLM 540B from 56.5% to 74.4% — a 17.9-point gain from voting alone (Wang et al., 2023). It works when correct answers cluster together in output space. It costs N× compute for sublinear gains.

3. Best-of-N + verifiers. Sample N solutions, then have a verifier model score them and pick the best. Two flavors: outcome reward models (ORMs) score the final answer; process reward models (PRMs) score every intermediate step. PRMs beat outcome-only verification by 10–20% on math — because catching a wrong step is more useful than grading a wrong answer. Training them is expensive (PRM800K used ~800,000 human step-level labels), so automated PRM training is now an active research front.

4. Search (tree of thought / MCTS). Explore multiple reasoning paths, evaluate partial solutions, backtrack when stuck — the model as both explorer and judge. Most powerful for hard combinatorial problems, and the most complex to run.

The scaling law: a new axis, same diminishing returns

The empirical rule: quality ∝ (inference compute)^β, roughly log-linear. Doubling thinking tokens buys a reliable but shrinking accuracy gain. Typical best-of-N sweeps run N=4 to N=256, with most of the value captured by N=64.

Diagram 3

The compute-optimal finding (Snell et al., 2024) is the one that reshaped budgets: on problems where the small model has a non-trivial baseline, a smaller model with the right test-time strategy beats a 14x larger model with a naive one. Two mechanisms drive it: searching against dense process verifiers, and adapting the strategy to the prompt's difficulty — easy problems get iterative refinement of one solution, hard problems get broad search.

The plot twist: thinking longer can make you dumber

Here's where September 2026 got interesting. New research on large reasoning models reached the opposite of the "more thinking is better" story:

Diagram 4

There's theory behind the damage. Bay & Yearick's "sampling ceilings" work shows test-time sampling is cluster sampling: attempts on one problem are correlated draws, so n samples buy only n/(1+(n−1)ρ) effective samples — capped at 1/ρ. And majority voting has a modal ceiling: where the model's most common answer is wrong, more samples sharpen the confident error. Coverage (getting one right answer for a verifier to find) keeps scaling; majority vote plateaus and can anti-scale.

The companion notebook reproduces all of this with exact binomial math: self-consistency 56.5% → 74.4%, coverage that never saturates, anti-scaling on wrong-mode problems, and a toy trace where taking the first answer beats taking the last.

SOTA, September 2026: the thinking leaderboard

The labs have turned these insights into product:

The infra strain is real: reasoning turned serving clusters into memory-bandwidth machines (single-digit MFU), a single 64K-token trace can eat 18 GB of VRAM in KV cache, and long traces risk infinite backtracking loops. That's why the KV-cache compression race (DeepSeek's ~90%) matters as much as the models.

The practitioner's rule: spend thinking like money

The research consensus for 2026 is a budget, not a slogan:

Test-time compute didn't replace training scale — it added a third axis. The labs that win from here are the ones that don't just think longer, but think smarter about when to stop.

Takeaways

Companion notebook

test_time_compute.ipynb — the runnable tutorial for this post (download, or open it in Colab/Jupyter).

← All posts