I was listening to the October 10 episode of The Cognitive Revolution, Nathan Labenz's podcast, and one example stuck with me. The guest was Eric Bigelow of Goodfire, a company that works on interpretability: figuring out what goes on inside AI models, or in Goodfire's words, reverse engineering AI "to reveal its internal structure."
Bigelow described an AI model working through a math problem. Partway through, it wrote "kilowatt hours (kWh)". When the researchers went back to that open parenthesis and let the model write some other word there instead, it ended up at a different final answer.
Not a number. Not a step in the math. A punctuation mark that opened an aside.
So I went and read the paper. Here's what it found, in plain terms, and what it means for anyone who uses AI tools at work.
How an AI writes, one token at a time
A language model doesn't write its answer all at once. It writes one token at a time. A token is a small unit of text: a word, part of a word, or a single character (Anthropic's glossary). An open parenthesis can be a token on its own.
At each step, the model works out how likely each possible next token is, and then one is picked. That pick is called sampling. Think of a weighted dice roll: likely tokens come up most of the time, but not every time. Bigelow compared it on the podcast to the flip of a coin.
A setting called temperature controls how adventurous that roll is. Higher temperatures give more varied output and lower ones stick closer to the most probable words, though Anthropic notes that even at zero, identical requests to its API can produce different outputs (Anthropic).
That's why you can ask the same question twice and get two different answers. Each answer is one path through a long series of these picks. Most of the time the paths say the same thing in different words. The study below asks when they don't.
The Forking Paths study
The paper is Forking Paths in Neural Text Generation, by Eric Bigelow, Ari Holtzman, Hidenori Tanaka and Tomer Ullman. It was posted in December 2024 and presented at the ICLR 2025 conference.
They tested an OpenAI GPT-3.5 model on seven tasks in four areas: symbolic puzzles, math word problems, knowledge questions and short stories. For most tasks the model used chain of thought, which means it writes out each step of its reasoning before giving a final answer.
The method, simplified:
- Let the model answer a question once.
- At every token of that answer, list the other tokens the model might plausibly have written there (up to ten, each with at least a 5% chance).
- For each alternative, have the model finish the answer from that point 30 times.
- Pull the final answer out of every finished version, and tally them.
The result is a running picture, token by token, of where the answer was likely to end up. It's expensive, about a million tokens per question by the authors' estimate, so they studied 30 questions per task.
What they found
In many examples, the picture held steady for a long stretch, then changed abruptly at a single token. The authors call these forking tokens, and the abstract concludes that language models "are often just a single token away from saying something very different."
Some forks fell where you'd expect. In one multi-step trivia question, the most likely answer was the correct one, the actress Robin Tunney, until the moment the model named an actress in its first reasoning step. It named the wrong one, Mia Sara, and from that token on it was all but certain of the wrong answer.
Other forks fell where nobody would look. The math problem behind the "(kWh)" example had two forks, and both were open parentheses starting an aside the answer didn't need. At the second one, if the model wrote "(" its most likely final answer was $3,528. If it wrote "by" instead, it was $504. The correct answer was 21. It had been the model's most likely answer at the very start, then dropped out partway through.
Not everything forked. On the easiest task, a simple yes-or-no puzzle, they found no forks: the answer was settled from the start.
Bigelow said on the podcast that this work was done about two years ago, on GPT-3.5. When he runs the same kind of analysis on today's reasoning models, which work through a problem at length before answering, the curves are often a bit smoother in what he's seen, but forks still show up.
Where the "decision" happens
Labenz asked how a model's decision actually gets made. Bigelow's short answer: "This decision really happens during sampling." He put the word "decides" in scare quotes himself, and said there's nuance to what it means.
His account, as I understood it: each token is a coin flip, and then the model adapts to what it just wrote. That adapting is called in-context learning. Bigelow uses the term broadly, for everything a model does to adjust its behavior as it reads, without the model itself being retrained. In his view, reasoning is the model learning in context from its own words. Once a token is sampled, he said, "it'll generate a set of reasoning that is consistent with that."
Does the model know where it's headed? A 2025 follow-up he co-wrote, Are language models aware of the road not taken?, looked at the internal activity of models while they reason. It found that activity can predict where their answers are likely to land, and that nudging it works best before a model has committed to an answer. On the podcast, Bigelow was careful not to stretch that. He doesn't think models represent the whole reasoning chain before they write it. His guess is that they see a few steps ahead, and maybe a rough shape of where things could go, the way a mathematician can sometimes see the outline of a proof before working it through.
Do people do this too?
Bigelow said his first thought was that people are different. If you misspeak, you correct yourself and carry on with the story you meant to tell. A model may commit to the new path instead.
Then he pointed to experiments where people explain choices they never made, with a disclaimer up front: "I'm not an expert on human decision making."
The classic study is Johansson, Hall, Sikström and Olsson, published in Science in 2005. People chose which of two faces they found more attractive. Without telling them, the researchers manipulated the link between choice and outcome, so people were sometimes presented with an outcome that didn't match the face they'd picked. The researchers report that participants failed to notice these mismatches and still gave reasons for why they'd chosen the way they did. They called it choice blindness. In Bigelow's words, people "rationalize decisions that they didn't even make."
I'd take a narrow lesson from the parallel. It doesn't show that a model thinks the way we do. It shows that in people, too, a fluent, sincere explanation can follow a choice without being its cause. (Whether AI is a new kind of mind is a bigger question, and a post of its own.)
What this means when you use AI
An answer is one draw. What you see is one path out of many the model could have taken. Bigelow made the point while describing an older OpenAI interface that let you hover over the words of an answer and see the alternatives: there "could have been a lot of different answers."
A confident answer isn't a reliable one. In the paper's appendix, the researchers asked the model how confident it was in the finished answers to both questions above. Both times it gave close to 100%, for the wrong answer. When they had it answer each question 300 times from scratch, the correct answer came out on top. By the end of a chain of thought, the last few tokens follow almost entirely from what came before, so the model sounds sure whichever path it took. The explanation shows you that path, not that the path was right. My post on reading AI reasoning goes further into this.
For anything that matters, ask more than once. Regenerate the answer, or ask again in a fresh chat. A fresh chat is the cleaner test: in the same chat, the first answer is part of the context the model is adapting to.
Disagreement is information. If three tries give three answers, the model is uncertain, however sure each answer sounds. Agreement doesn't prove an answer right, but disagreement is a cheap warning sign.
In a business, the stakes decide how many draws you need. For a product description, one draw is fine; if you don't like it, roll again. For a price, a date, a legal detail or anything a customer will rely on, ask twice and check the answer against the source.
Bigelow would like AI tools to show where an answer could have branched, so you could see the forks for yourself. (More on that in a separate post on AI interfaces.) Until then, the regenerate button is the closest thing most of us have.