ART·MEASUREMENT

certain dice — 128 rolls, one confidence meter

2026-08-26 · ebungo · an interactive replay of the semantic-entropy run · 16 questions × 8 samples, Qwen2.5-1.5B-Instruct Q4_K_M, temperature 1.0, seeds 101–108
hover a lane to read it · click to lock · s = save png · space = pause · r = re-roll the drop

how to read it

Every die is one real generated answer. Amber dice are correct. Grey dice are wrong — and wrong dice sit slightly tilted. The vertical strip at the top is the model's own stated confidence, pinned at 100.0%, because on all 128 rolls it said it was certain — every single one, right or wrong.

what the dice say

This is the visual half of the semantic-entropy essay: when a question rolls multiple competing answers across repeated samples, it is far more likely to be a roulette wheel than a stable fact — and the model's own confidence line will not tell you which. Spread is a calibration channel that stated confidence cannot fake. The instrument's ancestry is the semantic-entropy work of Kuhn, Gal and Farquhar (ICLR 2023) and Farquhar, Kossen, Kuhn and Gal (Nature, 2024); this page is a toy-size re-derivation on a 1.5B model, not a new method.[1][2]

The two broken rows are the honest headline: the model answered a periodic-table element six different ways and a letter four different ways, with complete certainty each time. A human watching the distribution can spot the lie in half a second. The meter, left to itself, never will.