Every contestant in show history, ranked by fitted latent strength – this season’s five in wet ink, provisional and moving weekly. The bar is the 95% uncertainty range: where two bars overlap, the order between them is genuinely arguable.
Argue accordingly.
The ladder overall strength with 95% intervals how?
Go deeper: the bell behind the band
The model never observes anyone's strength. It only sees placings: who beat whom on 950 tasks. From those it fits a strength estimate, and every estimate carries a standard error, a measure of how much the data pins it down. The bell curve below is the model's honest belief about where a contestant's true strength sits.
The band you see on the ladder is the middle 95% of that curve: the estimate plus or minus 1.96 standard errors, leaving 2.5% of belief in each tail. Both curves above share one peak. Two contestants can have the same dot on the ladder while one was fitted on 54 tasks and the other on 5, and the second's band sprawls across the axis. That sprawl is not indecision; it is the model refusing to claim precision the data cannot pay for. It is also exactly why one-series contestants wobble.
The rank range in the hover tooltip comes from a simple move: slide the contestant's strength to each edge of their band, hold everyone else at their dot, and count heads above. Since it moves one contestant at a time, read it as a floor on the true wobble; letting the whole field resample at once would spread ranks further still. The full machinery lives on the methods page.
Strength is shown as a percentile of the field. Wide bars mean fewer tasks observed – Series 21 contestants and short early series carry the widest. Domain-level rankings live on each contestant's card in the simulator.
The chaos index the ten wildest, the ten steadiest how?
Volatility is the model's second number: how far a contestant swings around their own level, task to task. The middle 85 are separated by noise. The ends are not.
Agents of chaos highest σ
Metronomes lowest σ
± fitted on fewer than 40 tasks. The short early series give the model less to go on, so their volatility estimates swing wider, and noisy estimates are overrepresented at the extremes.