Demos
Most of these demos ask a model the same thing many times and show what came back. The probability demos run on gpt-4o through the API, because that is the only way to see the probabilities the model gives each word. How the demos were run →
How these models work
- Fill in the blank”Complexity” had about half the probability and came out 13 times in 20.
- Fill in the blank, narrowSame machine, a sentence the text pins down. “Risk,” 20 times in 20.
- One word, then one sentenceAs a question, the answer spreads out. Two runs of one sentence land on two different research gaps.
- A fabricated citationReal authors, the right journal, a real volume, and a paper that doesn’t exist.
What this means for research
- “Are you sure?”It apologized 20 times in 20, and 19 moved from 80 to 92. The paper says 83.
- Attaching the sourceWithout the paper it refused or invented a quote. With it, the real sentence every time.
- Thirty citationsIn a thin literature, 17 of 30 DOIs existed and 8 were the paper named.
- Where would you apply it?Asked a hundred times: 49 wordings, nine ideas, and the most common one is the convention.
- Same question, different modelsThree models, three different favorite answers. Memory changed it most.