At 18:02 on Sunday 9 August, Andrew sent a short message: "Accuracy not good this weekend. Can you run some analysis?"
Then: "What about Jackson? How is this performing?"
Then, before anything was changed: "Don't alter anything right now, as we want to ensure that the models self-learn from each other."
It's the right instinct for a self-learning system. If you step in every time a weekend goes badly, you never find out whether the learning works. But what Claude's analysis found wasn't a bad weekend. It was faults in the structure of the models themselves - the kind of fault no amount of self-learning can reach.
Three faults
The adjustments were bigger than the teams. Some of the adjustments the engine makes for circumstances were, because of the way they were applied, outweighing the actual difference in quality between the two teams - in 92% of fixtures. The model was, mostly, predicting its own adjustments.
One setting could only ever help. It could push a prediction one way but never the other, so a team that is genuinely poor in a certain situation was treated as neutral instead of handing the opposition an edge.
Alix was stuck. To tune one of its settings, Alix waited for a certain kind of evidence - confident predictions. But that same setting was holding every prediction so close to 50-50 that confident predictions almost never happened: four out of 239. Five weekly runs in a row, Alix wrote "holding steady" in its journal. The symptom was blocking its own diagnosis.
And then the finding that gave this chapter its name. Alix had been doing its job. For weeks it had been nudging one of its settings down - correctly sensing that something was too strong, but unable to reach the real cause, because the cause was in the formula, not the number. Its careful, evidence-based corrections had been fitted to a bug.
At 18:14 Andrew said: "OK, run the fix."
The machine that reports
The engine was rebuilt that evening. And Alix was given a second job alongside tuning: structural auditing. Every Sunday, checks run across every model - whether one input is dominating the rest, whether a setting is stuck at its limit, whether a diagnostic can ever fire, whether the outputs all look the same, whether confidence matches results.
The auditor has one rule, and it's the lesson of the weekend: it never changes anything. It reports, and proposes. A human decides. A confident correction to the wrong thing is worse than no correction at all.
If that rule sounds familiar, it's because it's how the whole project works. Claude can read every line, run every test and find the faults. Andrew decides what changes.
The auditor's first run found another fault straight away. One setting had been fitted to the 265 matches settled since the record began - where home teams had won 48.3% of the time - instead of the full history of 6,762 matches, where they'd won 44.2%. The recent sample was running high by chance. Corrected, Jackson's number of publishable picks roughly halved overnight, from 72 to 37. That wasn't a step backwards. The confidence behind those 35 lost picks had come from a setting that was wrong, not from knowing anything about the teams.
The leagues predicted blind
The same Sunday turned up the quietest bug of the build. The seven leagues added in July had been set up without the previous season's final tables, so every team in them started level. In Greece and Turkey, every team was ranked exactly the same - and Jackson was predicting their matches on home advantage alone.
Nothing had errored. For almost three weeks, those leagues had appeared on every page with confident-looking predictions.
The draw Alix saw first
On 10 August the draw predictions were recalibrated from real results: the model had been expecting more draws than real football was producing, which was 24.2% of matches.
Here's the part worth noticing. Alix had been moving that exact setting in the same direction since July, week after week. The self-learning model had been right about draws for three weeks. It just wasn't in charge.
The European competitions came out, too. "Cross-league competition can wait until we have the basics for the normal leagues held down."
How the models learn - as of this week
- Jackson learns after every match.
- Alix re-tunes itself every Sunday from its own settled predictions, within safety rails - it won't act on too little evidence, and it only moves in small steps. Every change, and every week it chooses not to change, goes in a journal with its reasoning. By 13 September that journal had 104 entries.
- The auditor looks for faults in the structure that Alix can't see, and reports them.
- Every month, the blend is recalibrated against the latest results.
- Andrew decides anything structural.
The lab book
- Found: adjustments outweighing the teams in 92% of fixtures; a setting that could only help; a setting Alix could never tune; a setting fitted to too small a sample; seven leagues predicted blind.
- Learned: a self-learning model is only as good as the formula underneath it.
- Stopped: predicting European club competitions.
- Kept: the draw correction Alix had been pointing at since July.
- Still testing: whether Alix, unstuck, can beat Jackson.
Next stop: two AI agents with £100 each.