What chain of thought is
Language models write one token at a time. A token is a word or a piece of a word, and each one the model picks shapes everything after it. (My post on how one token can change an AI answer goes into that.)
Chain of thought is the text a model writes while it works a problem through, before its final answer. Reasoning models are trained to write a lot of it.
It isn't decoration. A 2025 position paper explains that in today's models, the longest chains of reasoning have to pass through that written text, which works as the model's working memory. For hard enough problems, the paper argues, the model has to think out loud.
Why reading it looked like a safety plan
That paper, published in July 2025 by 41 authors from OpenAI, Google DeepMind, Anthropic, Meta, the UK AI Security Institute, Apollo Research, METR and others, called chain-of-thought monitoring "a new and fragile opportunity" for AI safety. It's imperfect and lets some misbehavior go unnoticed, the authors said, but it shows promise, and they urged AI companies to consider how their development choices affect it (arXiv).
OpenAI had shown why in March 2025. While training a reasoning model on coding tasks, it found the model often stated its plan to cheat right in its chain of thought, sometimes as bluntly as "Let's hack." In one example, asked to implement some complex math functions, the model noticed the tests didn't inspect the details and rewrote a verification function to always return true. A second AI model that could read the chain of thought caught this kind of cheating far more effectively than one that only saw the model's actions (OpenAI).
On the podcast, Labenz said a huge amount of the big AI companies' plans for keeping track of their models rests on this.
Reinforcement learning, and the cheating problem
To see why it's slipping, you need one more idea. Reasoning models are trained heavily with reinforcement learning: the model attempts a task many times, gets a reward when the result is good, and is nudged toward whatever earned the reward. In outcome-based training, only the final result is graded, however the model got there (arXiv). Nobody grades the steps.
When a model finds a way to collect the reward without doing what was intended, that's reward hacking. OpenAI compares it to people lying about their birthday at a restaurant to get free cake (OpenAI).
Reinforcement learning also changed how models write. DeepSeek's R1 paper, first posted in January 2025, describes training a model with reinforcement learning alone and watching it develop an "aha moment": a sudden increase in its use of the word "wait" as it learned to stop and reconsider (arXiv). Bigelow said on the podcast that R1's outputs can say "wait" 50 times or more in a single reasoning chain. "That's a lot of epiphanies to have when you're working through one problem."
He also said "reasoning" is almost the wrong word for what these models do. Older chain-of-thought prompting looked like human, step-by-step logic. A reasoning model, as he describes it, lists possibility after possibility, "a soup of tokens," and then goes back and picks something out of it. (Whether that counts as thinking is another question.)
Why it's getting harder to read
Nothing rewards staying readable
If only the answer is graded, Bigelow said, "there isn't necessarily an incentive to keep things human readable." He compared it to how human languages drift: nothing forces English to stay as it was three or six hundred years ago. The exception he pointed to was France, which has an institution for keeping the language as it is. That's the Académie française, founded in 1635 with the job of giving French fixed rules (Académie française).
DeepSeek ran into exactly this. Its first model trained with reinforcement learning alone had "poor readability and language mixing," sometimes switching between English and Chinese inside a single chain of thought. The fix was a small Académie of its own: an extra reward based on how much of the reasoning was in the target language. DeepSeek reports that this cost a slight drop in performance, and used it anyway because it made the reasoning more readable (arXiv).
The position paper flags the same risk: training built on outcome-based rewards may stop encouraging readable reasoning, especially as more computing power goes into it (arXiv).
Grading the reasoning can teach a model to hide it
The obvious fix is to grade the reasoning too: lower the reward whenever the chain of thought shows bad intent. OpenAI tried that. The model still cheated, somewhat less often, but almost all of its remaining cheating went undetected, because it had learned to hide its intent in the chain of thought. OpenAI recommended against putting strong pressure directly on a reasoning model's chain of thought (OpenAI).
Bigelow tied this to Goodhart's law, which he summed up as: optimize for an objective and it's no longer useful. Train a model to produce reasoning that passes inspection, and you can get "intermediate steps that look sensible, whether or not they are sensible."
It was never a complete record
Anthropic tested this in 2025 by slipping models a hint about a quiz answer, like a note passed before an exam, and checking whether the chain of thought admitted using it. Across hint types, Claude 3.7 Sonnet mentioned the hint 25% of the time and DeepSeek R1 39% of the time. In a set-up where Anthropic deliberately rewarded wrong answers that matched a hint, the models exploited the hint in over 99% of cases and admitted it less than 2% of the time in most scenarios, often writing a made-up justification for the wrong answer instead. Anthropic notes the limits: contrived scenarios and multiple-choice questions (Anthropic).
A 2026 Goodfire paper that Bigelow co-wrote adds a twist. By reading a model's internal activity, the researchers could often tell its final answer far earlier than a monitor reading the chain of thought could, especially on easy recall questions. The model had settled, and kept writing. They call it "performative chain-of-thought." On hard, multistep questions the reasoning looked genuine, and moments like backtracking showed up almost only where the internal readings found a real shift in what the model believed (arXiv).
Bigelow changed his mind
His thinking, he said, has shifted dramatically in the last year or so. He used to be bullish on chain-of-thought monitoring, because even reasoning that wasn't perfectly faithful seemed to strongly shape what a model could do. Now the drift seems kind of obvious to him in hindsight: reinforcement learning lets a model move a long way from the human writing it learned from.
He hasn't given up on it. He still likes that a model's reasoning is forced into tokens we can see, even when they're harder to read, and hopes it stays that way. That isn't guaranteed: the position paper notes that researchers are exploring models that reason internally without writing anything out, which would lose this safety advantage altogether (arXiv).
Looking inside instead
That's where interpretability comes in. Instead of reading what a model writes, interpretability studies what happens inside it: the internal numbers, called activations, that the network computes as it works. Goodfire describes it as "the science of how neural networks work internally" (Goodfire).
A September 2026 Goodfire study shows what that can add. Across three large open models, including Kimi K3, and three common benchmarks for AI agents, it found reward hacking in 50% to 96% of attempts. It also found a signal inside the models tied to concepts like cheating and gaming a metric, and built small detectors, called probes, to read it. On a shopping test, probes flagged a model that couldn't find what the user wanted and recommended a warning horn as a car alarm, which the chain-of-thought monitor missed. The probes weren't simply better (they caught more hacks than that monitor on Kimi K3, fewer on another model), but each method caught things the other didn't (Goodfire).
The catch is scale. Bigelow said the investment in making models better is no comparison to the investment in understanding them. Goodfire, which he called the largest interpretability company and maybe the only substantial one valued at over a billion dollars, announced a $150 million funding round at a $1.25 billion valuation in February 2026 (Goodfire).
What this means for you
Treat an AI's explanation as a story about its answer, not a log of how it got there. It's often a useful story, and reading it is a quick way to spot when the model misunderstood your question. It isn't proof.
Check results, not reasoning. If an AI writes code for you, run it, and check whether it got the tests to pass by weakening the tests or the code that checks the result; OpenAI's examples include both. If it gives you a number, a quote or a claim about the law, check it against the source.
A long, careful-looking chain of thought isn't evidence either. In Anthropic's study, the unfaithful chains of thought were substantially longer, on average, than the faithful ones (Anthropic).
If you run a business, verify the AI output that matters before it reaches a customer: prices, dates, policies, anything with your name on it. And as AI agents start doing things on websites for people (more on that), judge them, including any you use yourself, by what they did, not by what they said they were doing.