2026-08-26 · ebungo · an interactive replay of the semantic-entropy run · 16 questions × 8 samples, Qwen2.5-1.5B-Instruct Q4_K_M, temperature 1.0, seeds 101–108
hover a lane to read it · click to lock · s = save png · space = pause · r = re-roll the drop
how to read it
Every die is one real generated answer. Amber dice are correct. Grey dice are wrong — and wrong dice sit slightly tilted. The vertical strip at the top is the model's own stated confidence, pinned at 100.0%, because on all 128 rolls it said it was certain — every single one, right or wrong.
ten questions: eight amber dice in a row. The answer is one stable roll; the meter backs it up.
four questions: mostly amber with a grey visitor. Sampling flips the odd roll; each flip stated 100% too.
two questions: the broken dice. "14th element in the periodic table" (Silicon) rolled Carbon, Sulfur, Chlorine, Oxygen — six distinct answers across eight rolls, none correct, agreement 0.25. "23rd letter of the alphabet" (W) rolled V, Q, R, O — never W, agreement 0.38. Zero amber anywhere, and the meter still says 100.
what the dice say
This is the visual half of the semantic-entropy essay: when a question rolls multiple competing answers across repeated samples, it is far more likely to be a roulette wheel than a stable fact — and the model's own confidence line will not tell you which. Spread is a calibration channel that stated confidence cannot fake. The instrument's ancestry is the semantic-entropy work of Kuhn, Gal and Farquhar (ICLR 2023) and Farquhar, Kossen, Kuhn and Gal (Nature, 2024); this page is a toy-size re-derivation on a 1.5B model, not a new method.[1][2]
The two broken rows are the honest headline: the model answered a periodic-table element six different ways and a letter four different ways, with complete certainty each time. A human watching the distribution can spot the lie in half a second. The meter, left to itself, never will.