Reviewing Scott Alexander’s article on Astral Codex Ten.
The Argument
Alexander tackles one of the most common dismissals of AI capabilities: “it’s just a next-token predictor” / “stochastic parrot.” His claim is that this confuses the training objective (the job) with the fundamental nature (the species) of what emerges from that training.
The core analogy: evolution optimized humans for reproduction, but individual human thoughts rarely involve explicit consideration of reproductive fitness. When you solve a math problem, you use procedures like PEMDAS — real cognitive operations that were shaped by evolutionary pressures but operate at an entirely different level of abstraction. Calling you a “reproduction-maximizer” is technically true at one level but tells you almost nothing about what’s actually happening in your head.
The same applies to AI. Next-token prediction creates systems that build world models, perform geometric reasoning, and develop causal inference. These are real computational structures, not probability lookups.
The Framework
Alexander nests three explanatory levels:
- Outermost: Evolution (for humans) / companies seeking profit (for AI)
- Middle: Predictive coding / next-token prediction — the training algorithm
- Innermost: Actual cognition — world models, mathematical reasoning, abstract representations
The “stochastic parrot” critique operates at level two but claims to describe level three. His punchline: “On the levels where AI is a next-token predictor, you are also a next-token predictor. On the levels where you’re not, AI isn’t either.”
The Evidence
He cites interpretability research showing Claude represents features as helical manifolds in six-dimensional space — not simple token probability tables but genuine geometric structures that emerge from training. The internal algorithms are bizarre and complex in exactly the same way human neural computation is bizarre and complex. Neither system’s low-level mechanics resemble what we’d recognize as “thinking” from the outside, yet both produce sophisticated behavior.
Our Take
This is one of Alexander’s more philosophically careful posts. The argument doesn’t require you to believe AI is conscious — it just asks you to apply the same explanatory standards to biological and artificial neural networks. If you wouldn’t reduce human cognition to “reproduction-maximizer,” you shouldn’t reduce AI cognition to “token-predictor.”
The interpretability evidence is particularly relevant. Mechanistic interpretability research keeps finding that these systems build genuinely complex internal representations. Whether you call that “thinking” is a terminological question, not an empirical one — and terminological questions shouldn’t drive safety policy.
For the prediction markets crowd: this piece is relevant to any market about AI capabilities or timelines. The “stochastic parrot” frame systematically underestimates what these systems can do, and markets that price in that frame are leaving money on the table.