CV Google Scholar MSU
Demos · How these models work

One word, then one sentence

The same model, now asked a question instead of handed a sentence to finish. First it gets one word to answer in, then a sentence. It is still predicting what comes next. Guide: Why fluency tells you nothing →

One word

In one word, what is the most important limitation of prospect theory?

We ran it 20 times.

What the model thought could come next

The probability it gave each word, before choosing.

rationality
25.1%
predictability
22.7%
descriptiveness
21.1%
complexity
4.1%
oversimplification
3.6%
ambiguity
2.5%
consistency
2.4%
prediction
1.4%
scope
1.3%
everything else
15.8%

What it actually said, 20 times

Each box is one run. Same sentence, same model, every time.

rationalityrationalityrationalitypredictabilitypredictabilitypredictabilitypredictabilitypredictabilitydescriptivenessdescriptivenessdescriptivenessdescriptivenessnarrownessinconsistencyprobabilitycontext-dependenceapplicabilitycomplexityobservabilitydescriptive
gpt-4o via the API · temperature 1.0 · September 2026 · 20 runs

Asked as a question, the probability spreads out. The favorite takes about a quarter, and 20 runs give 11 different words. The word the model said most often wasn’t even its favorite.

One sentence

I am writing a paper that builds on prospect theory. In one sentence, what is the single most important gap it leaves?

Two runs, same prompt, same settings. Both began “Prospect theory addresses decision-making under risk but leaves a significant gap in…” and then had to pick the next word.

What the model thought could come next

“Prospect theory addresses decision-making under risk but leaves a significant gap in ___”

explaining
47.9%
accounting
18.8%
understanding
14.2%
fully
11.7%
its
4.2%
comprehensively
1.2%
adequately
0.8%
not
0.3%
how
0.2%
everything else
0.7%
gpt-4o via the API · temperature 1.0 · September 2026
Run 1
…leaves a significant gap in accounting for how individuals consistently integrate emotions and long-term consequences into their evaluations.
Run 2
…leaves a significant gap in explaining how people evaluate outcomes that have both elements of ambiguity and uncertainty beyond quantifiable risks.

What to notice

  1. It’s still predicting what comes next, not answering your question. The question is just more text to continue from.
  2. Two runs, two different research gaps. Emotion and long-term consequences in one, ambiguity in the other. Neither was chosen because it was the better answer.
  3. Both read like a confident claim. Each word is predicted from everything before it, so the sentence always hangs together, whichever way it went.

Limits

One prompt of each kind, one model. Your own runs will come out differently, draw by draw.