Re-simulate all twenty-one series, 3,000 times each, using only the fitted skill of the people in the room. Then answer one question per series: did the right person win?

The receipts are below.

The ledger title probability the model retro-assigns each real champion

Coronation ≥40% Deserved ≥28% Fair fight ≥20% Upset ≥12% Daylight robbery below Pure chance would give every champion 20%.

Most robbed actual points vs model expectation, worst gaps of all time

Biggest overachievers the players who beat their own talent

The State of Greg has the Taskmaster gone soft? 21 series of judging, measured

Punitive zeroes non-DQ 0-point awards, % of scores

Tie generosity tasks with shared scores, %

The fan theory is half right. Early Greg reached for the zero as a punchline – one in every eight Series 2 scores was a nought – and that habit has clearly faded: punitive zeroes run a third lower in the modern era. But tie generosity tells a stranger story. Shared scores climbed steadily into a soft-touch golden age around Series 8–15, peaking at two tasks in three, then fell back through Series 19–21. If anything, the newest Greg is re-toughening. Disqualifications, for the record, show no trend at all: DQ rate belongs to the cast, not the judge.

Caveats, stated plainly: twenty-one data points, a task mix that changes across eras, and some zeroes are earned by the attempt rather than issued by the mood. Treat this as evidence, not verdict. how?

The Knappett Question is Greg harder on women? An on-air theory, tested

Jessica Knappett raised it on the show itself, and the fandom has debated it since: does Greg judge women more harshly? The naive version of this question – do women score lower? – is unanswerable, because lineups, eras, and task mixes all differ. So we asked the honest version instead. Prize tasks are the one arena of pure Greg judgment: no stopwatch, no measurement, only his taste. For every contestant we compared their prize-task strength against their own overall baseline, so ability cancels out and only the judging remains.

The result: nothing. Women sit a hair below their baselines on prize tasks and men a hair above, a gap of 0.03 on a scale where task-to-task noise is 1.0 – a difference chance produces 45% of the time. Live tasks, Greg's other judgment arena, are dead even. And rule enforcement leans mildly the other way: men collect more punitive zeroes (5.4% vs 4.8%) and slightly more disqualifications. Mae Martin, the roster's one nonbinary contestant, is reported separately rather than binned, and is a single data point, not a pattern.

Stated plainly: with 43 women and 61 men, this test can only rule out moderate-to-large bias; an effect below that floor would be invisible to us. And scoring is all the numbers can see – the banter is a different question for a different instrument. But on the evidence available, Greg's gavel does not know who is holding the prize. how?

The Strongest Series the fandom's favourite argument, measured – and why you can't have it

Every series' mean fitted strength, with 95% intervals. Series 20 sits nominally on top and Series 5 nominally at the bottom – and the honest reading is that it does not matter, because every interval overlaps every other. The gap between the "strongest" and "weakest" series ever is a quarter the width of the uncertainty on any one of them. Within a series, contestants are ranked against each other, so a series' absolute level is only pinned down by Champion of Champions bridges – thin evidence, honestly reported as such. The most-argued question in the fandom is, statistically, a tie. Argue anyway. how?

The prize problem the only round where the skill is having taste