Skip to content
Popproxx brand logo in stylized cursive font.
Close-up of an octopus's eye, a dark slit pupil in a speckled orange iris, set in pale, bumpy skin

Is AI Showing Us a New Kind of Mind?

Jamin Giersbach

I listened to the October 10 episode of The Cognitive Revolution, an interview with Eric Bigelow of Goodfire, then went and read his papers. What stayed with me was his method: he studies AI language models the way a psychologist studies a mind.

A large language model (LLM) is the kind of AI behind chat assistants like Claude (Anthropic). It's trained on huge amounts of text to predict the next token, a word or piece of a word, over and over.

Popular takes call these systems either minds like ours or just autocomplete. The research I read fits neither.

A psychology PhD about language models

Bigelow's Harvard dissertation, dated May 2026, is titled Towards a Cognitive Science of Large Language Models (Harvard DASH). Cognitive science studies the mind and the hidden processes behind behavior. Bigelow proposes studying AI the same way and calls it cognitive interpretability, interpretability being the field that tries to see inside a model. He now works at Goodfire (his site), an AI interpretability research lab (Goodfire).

On the podcast, he said that when LLMs arrived he dropped everything to study them, and got a lot of pushback from scientists around him. Many cognitive scientists found the models boring until recently, he said; some of that has softened as the models have proved so successful.

One of his studies shows why this pays off. Bigelow and colleagues borrowed a task from human psychology: a fictional game show where two contestants, each with a strength and an effort level, try to lift a box together. They asked three GPT-4-family models how much blame or credit each contestant deserved. The models mostly judged by force, meaning how much each contestant actually did. People in earlier studies also weighed effort: how hard someone tried, and how hard they could have tried (Xiang, Bigelow et al.). His dissertation sums it up: LLMs use a different mechanism from humans when assigning responsibility.

Models learn while they read

Training sets a model's weights, the billions of numbers that store what it learned, and then they stay fixed. Yet the model still adapts to what you give it. Show it three examples of a format and it follows the pattern. That adapting on the fly, with no change to the weights, is in-context learning. The context is everything in front of the model right now: your prompt, the conversation, any document you pasted in.

Bigelow takes a broad view of it and credits a paper by Andrew Lampinen and colleagues, The broader spectrum of in-context learning. It suggests that whenever earlier text helps a model predict later text, as when it follows instructions, something like in-context learning is going on. The authors trace possible roots to basic reading skills like working out who a pronoun refers to.

He likes the example of a story. As a model reads one, he said, you'd want it updating what it believes: who the good guys and bad guys are, who's happy and who's sad. His 2026 paper with seven co-authors, Stories in Space, found that a model's changing beliefs about a story trace a path through a small, structured space of concepts, a path the team could read from the model's internals and nudge in predictable ways.

Prompts, nudges and sudden flips

You can change what a running model does in two ways: change what it reads (prompting), or reach inside and change what it computes (activation steering).

Activations are the internal numbers a model produces as it processes text. Researchers can find a pattern in those numbers that goes with a concept, then turn it up or down while the model runs. In 2024, Anthropic found a feature tied to the Golden Gate Bridge inside its Claude 3 Sonnet model and turned it up. With no prompt asking it to play-act, the model brought up the bridge in almost every answer, and when asked about its physical form, said it was the bridge (Anthropic, Anthropic).

Bigelow's belief dynamics paper, a poster at ICML 2026, argues that prompting and steering are two handles on the same thing: the model's belief in a hidden concept. It borrows Bayesian inference from cognitive science: start with a prior, how likely something seems before any evidence, and update it as evidence arrives. In the paper's account, steering shifts the prior, and examples in the prompt add evidence.

The team tested this on personas, mostly with a model called Llama 3.1 8B, feeding it more and more example answers written in a character, such as a narcissist or a moral nihilist. How likely it was to answer in character followed an S-curve: little change for a while, then a fast flip. Steering moved the flip point, so they could predict how many examples it would take to tip the model into a persona. Because the two effects add up, a small change to either can trigger a sudden, dramatic shift in behavior, which the authors flag as a safety concern. The paper notes its limits: only simple two-way concepts, one steering method, and at least one model where steering had no clear effect.

That's the appeal for a cognitive scientist, Bigelow said: it's hard to check whether brains really work this way, but LLMs can be opened up and experimented on.

What about the stochastic parrot?

The phrase comes from a 2021 paper by Emily M. Bender, Timnit Gebru and two co-authors, On the Dangers of Stochastic Parrots. Most of it is about the costs and risks of ever-larger models: environmental and financial cost, biased and poorly documented web-scraped training data, and research chasing leaderboard scores. The parrot is one section. It argues that a language model learns the form of language without access to its meaning, stitching together patterns from its training data by probability, and that the coherence readers see is supplied by the readers. The worry was the harm that follows when people take that text at face value.

Asked whether the stochastic parrot still has a place, Bigelow said he thinks the metaphor should probably be put to rest. He called it "a fairly shallow argument that doesn't really say very much" and compared it to the Chinese room.

In that 1980 thought experiment, the philosopher John Searle imagined himself in a room, following a program for answering Chinese characters slipped under the door. He knows no Chinese, yet the people outside assume a Chinese speaker is inside. His conclusion: a program can make a computer appear to understand language without real understanding (Stanford Encyclopedia of Philosophy).

Bigelow's answer: if the room's rulebook can hold a conversation, the interesting question is what's in it. Pointing to work that finds structured models of the world inside LLMs, he argued that past some point, learning a world model is more efficient than memorizing a giant lookup table.

He kept "stochastic," though, which means involving chance. A model picks each next token by sampling, drawing from a spread of likely options like a weighted coin flip. His Forking Paths paper found that models are often one token away from saying something very different, sometimes at a token as small as a punctuation mark. He argued that this chance is essential if you want models that give varied answers and stay consistent with themselves. More on that in one token can change an AI answer.

My read: the parrot paper's warnings about cost, training data and people trusting fluent text haven't expired. What Bigelow's work pushes back on is the narrower claim about what's going on inside the model.

Talking about AI as if it were a person

Bigelow described two healthy ways to anthropomorphize, meaning to talk about something as if it were human. The first is loose and just for you: if your car starts clunking, you might say it's trying to tell you something. That's fine, he said, as long as it helps and you don't take the words too seriously. The second is serious: treat words like belief, decision and intention as things to study in a model. He invited critics in psychology, linguistics and philosophy to help with that rather than just object.

He also pointed out that people aren't as tidy as we assume. In a 2005 study, participants chose the more attractive of two faces while the researchers covertly switched the outcome they were shown. Many failed to notice, and gave reasons for a choice they hadn't made. The researchers called it choice blindness (Science).

What this means if you use AI tools

I'd put it this way: these tools are neither a person nor a calculator, and neither picture predicts how they behave.

What the research suggests, in practice:

  • Your context does the steering. Everything you give a model is evidence it updates on, stray details and mistakes included. Bigelow's advice for working with AI research agents applies: give them context and explain your reasoning.
  • Ask twice when it matters. The same question can come out differently, sometimes because of one early word. For anything important, ask again and look at where the answers split.
  • Expect tipping points. In the persona experiments, behavior didn't drift; it flipped once enough examples piled up. If a conversation seems stuck on one track, starting a fresh one is a cheap test.
  • Check its judgments about people. As the blame-and-credit study shows, a model can reach a human-sounding verdict by a different rule than you'd use.

Bigelow would like chat tools to show where an answer could have branched (beyond the chat box). Whether its written reasoning can be trusted is another question (why AI reasoning is getting harder to read).

The bigger lesson for me is that intelligence is turning out to be less one thing than we assumed. A model can follow a story, hold a conversation and work through a problem, yet assign blame by a different rule than we do. People, for our part, can explain choices we never made. Bigelow sees this research as a chance to learn something about ourselves, and after reading his work, I agree.

Jamin Giersbach, who runs Popproxx
Written by

Jamin Giersbach

Jamin Giersbach is one of the three designers who started Popproxx in New York City in 2000. Early clients included WebMD, MSD Capital and Bookmans. In 2007 he took Popproxx to Oregon, and he has run it on his own ever since.

Read the full story →