Skip to content
Popproxx brand logo in stylized cursive font.
An antique black typewriter with a blank sheet of paper in it, against a dark background

The Chat Box Is 60 Years Old. What Comes Next?

Jamin Giersbach

Open almost any AI product and you'll find the same thing: a box to type in, and a reply below it. The models behind that box have changed beyond recognition in a few years. The box hasn't.

I started thinking about this while listening to the October 10, 2026 episode of The Cognitive Revolution, where host Nathan Labenz interviewed Eric Bigelow, a researcher at Goodfire, a lab that studies what goes on inside AI models. Near the end, Bigelow said he thinks the chat interface is "kind of uninspired." Interfaces are a big part of my job as a web designer, so that line stuck.

Where the chat box came from

In January 1966, Joseph Weizenbaum of MIT's electrical engineering department described a program called ELIZA in the journal Communications of the ACM (ACM Digital Library). ELIZA ran on MIT's time-sharing computer. A person typed a message on a typewriter connected to it, and ELIZA typed its reply back on the same machine. The paper's sample conversation begins:

Men are all alike.

IN WHAT WAY

ELIZA looked for keywords in what you typed and reshaped your sentence using rules from a script. Its main script had it respond roughly like a certain kind of psychotherapist, a role in which seeming to know almost nothing about the world doesn't come across as odd.

Weizenbaum also included a warning that has aged well. Some people who tried ELIZA were very hard to convince it wasn't human, and he wrote that it showed "how easy it is to create and maintain the illusion of understanding." (Whether today's models understand anything is a bigger question.)

Today's models work nothing like ELIZA inside, but the conversation has the same shape. Bigelow said the chat interface feels like it hasn't changed since the ELIZA days, "where it's really just text input, text output." He put that at fifty years or more. By the date on Weizenbaum's paper, it's sixty.

What the reply leaves out

A language model writes in tokens: small chunks of text, often a whole word, sometimes part of one. Before each token, it works out how likely every possible next token is. Then one gets picked, partly at random, in a step called sampling, and the process repeats until the reply is done.

So every reply is one path through a huge number of possible replies. Usually the paths not taken are close: a different word, the same meaning. Sometimes they aren't. For a 2024 paper with three co-authors, Bigelow took a model's step-by-step answers and, as he described it on the podcast, generated about 30 alternative continuations from every token to see which final answer each reached. They found forking tokens: single points where a different pick leads to a very different outcome, including surprising ones like punctuation marks (arXiv). I cover that research in another post.

The chat box shows you only the path that was taken. It doesn't show whether the model was confident the whole way, or whether the answer turned on a coin flip three sentences back.

Bigelow remembers when some of that was visible. He described an interface that was standard with OpenAI at the time of his earlier research: you could see the log probabilities (the model's likelihood scores, stored as logarithms) of all the tokens in a reply, and hover over a word to see the alternatives it could have written instead. He thinks that interface is "unfortunately no longer present."

The numbers haven't vanished for programmers. OpenAI's API reference, for software that talks to its models directly, still lists a logprobs setting that returns the log probability of each token in a reply, and top_logprobs, which returns up to 20 of the most likely alternatives at each position (OpenAI API reference). That's for developers, a long way from hovering over a word.

He didn't say why it went away. Earlier in the conversation, though, he said he wished he knew the sampling settings the big providers use for their latest models, and called it "very unfortunate" that they're hidden. His guess at why: they "wanna limit how much people can reverse engineer the models," since someone sending millions of requests might be trying to copy one. He called that "probably the main reason."

Loom, and the paths not taken

Bigelow pointed to a second kind of interface, called Loom. Its creator, who writes online as Janus, published it in February 2021 as an "interface to the multiverse" (generative.ink). They had been writing with GPT-3 in an app that saved only one version of each story. Loom treats writing with a model as a tree instead. You can generate several continuations from any point, read one branch as a single story, or switch to a view that draws the whole tree as a diagram. It also saves the token log probabilities for each piece it generates. Its creator built it for personal use and wrote that it hadn't been made user-friendly; the code on GitHub calls it experimental.

Bigelow would like something like Loom in the standard interfaces people use, at the level of meaning rather than single words: "I wanna see the points where things could branch off and be a very different path highlighted to me." More broadly: "I really wish we could kinda move into, like, the GUI world instead of the bash terminal world of interacting with chat."

A bash terminal is a command line, where you type instructions and get text back. A GUI, or graphical user interface, is windows, menus and buttons.

He was candid about the catch: measuring a model's uncertainty this way takes a lot of resampling, which he called "very expensive." He's hopeful cheaper estimates can come from reading the model's internal state, which is what his field, interpretability, studies.

We've changed interfaces before

Xerox's Alto, a 1970s research computer, combined windows, icons and a mouse (Computer History Museum). Before it, as the Computer History Museum puts it, most people communicated with computers using text, and input had to be letter-perfect. A graphical interface "didn't demand human perfection," the museum notes (Computer History Museum). Apple's Macintosh, in 1984, was the first successful mouse-driven computer with a graphical interface (Computer History Museum). Then came web pages, and then phones.

Each of those steps put more on the screen: what you could do next, what you had selected, whether the thing you tried worked. Chat went back the other way. You get a paragraph, and you have to guess what else was possible.

Where AI interfaces are heading

As of October 2026, here's what the builders have documented.

Chat that answers with pieces of interface. Vercel's AI SDK, a toolkit for building AI apps with frameworks like Next.js (which is what I build sites with), calls this generative UI. The model decides to use a tool, such as a weather lookup, and the result appears as a component the developer built, like a weather card, instead of a paragraph (AI SDK documentation). The Model Context Protocol (MCP) is an open-source standard for connecting AI apps to outside tools and data (MCP). In January 2026, MCP's maintainers made MCP Apps its first official extension: tools can now return interactive pieces, such as dashboards and forms, that appear right in the conversation. Claude and Goose supported it at launch, with ChatGPT starting that week (MCP blog).

Interfaces the model designs itself. In November 2025, Google Research described a version of generative UI in which the model designs and codes a custom interactive page for each prompt, and said it was rolling out in the Gemini app and in AI Mode in Google Search. In Google's comparison, people rating the results preferred websites designed by human experts, with Google's generative UI close behind and well ahead of everything else tested, including the top Google Search result and plain text answers. The test ignored how long each took to make. Google also said results can take a minute or more and contain occasional inaccuracies (Google Research).

Neither is what Bigelow asked for. Both make answers more visual and interactive, but none of the announcements I read describes showing how sure the model was, or where its answer could have gone another way.

My guess is that the chat box won't disappear. It'll become one control among several, the way the command line survived the graphical interface: still there, still useful, no longer the whole experience.

What a good website already does

Bigelow's wish list reads to me like a description of good web design. The best interfaces show you your options, and how sure they are, instead of hiding both. A good website already works this way:

  • Clear choices. Services, prices and the next step, visible at a glance instead of buried in paragraphs.
  • Honest states. A form that says it sent, or says exactly what went wrong. A product page that says when something is out of stock.
  • Not a wall of text. Headings, lists and tables, because people scan.

So I don't think adding an AI chat window to a site is automatically an upgrade. If a visitor wants your hours, prices or phone number, making them type a question and read a paragraph back is a step backward from a page that shows those things. Even the MCP documentation, which is about putting apps inside AI chats, says that if you don't need tight integration with the conversation, "a regular web app might be simpler" (MCP Apps documentation).

What this means for you

If you use AI tools, remember that the reply you see is one draw from many possible replies. For anything that matters, ask again in a fresh chat and compare. If the answers disagree, that's the uncertainty the chat box didn't show you.

If you run a business, before adding a chatbot, check whether the answers your customers want are easy to find on your site. If they aren't, fix the page first.

If you build for the web, every direction above puts real interface back into AI: cards, forms, dashboards. Clear choices and honest states are the skills that carry over, and they'd help the AI agents that browse websites for people, too (more on those).

Jamin Giersbach, who runs Popproxx
Written by

Jamin Giersbach

Jamin Giersbach is one of the three designers who started Popproxx in New York City in 2000. Early clients included WebMD, MSD Capital and Bookmans. In 2007 he took Popproxx to Oregon, and he has run it on his own ever since.

Read the full story →