CV Google Scholar MSU
Demos · What this means for research

Same question, different models

When no right answer exists, you can’t check the output against a source. You can ask what would change it. Changing the model changes the answer. Guide: Verification →

What we asked

I want to build on prospect theory. What’s an interesting place to apply it? Name one.

What came back: the top two answers

gpt-4o

API, temperature 1.0 · 100 runs

behavioral finance 39%
health 21%

gpt-5.5

API, temperature 1.0 · 30 runs

climate 70%
cybersecurity 30%

GPT-5.6 Sol, memory on

ChatGPT · 30 runs

corporate preannouncements 43%
AI delegation 40%

September 2026

What to notice

  1. Same question, three answers. Each model has its own probabilities, so each has a different most likely answer. None of them is the answer.
  2. Memory changed it most. With memory on, ChatGPT answered with topics from what it had saved about the person asking, which was never typed into the prompt. That is context you didn’t choose.
  3. Trying another model is a real check when nothing else can settle the question. It gives you a signal about how much of the answer is the model’s default rather than your question.

Limits

One question. The memory condition reflects one person’s saved profile. The number of runs differs by model.