In the last week of July, the new leaderboard digest named its best model of the week.
"Good morning! Last week's top performing model was Form - 58.5% accuracy over 41 fixtures."
Form - the simplest model in the building, which does nothing more than look at recent results - had won the week. Avenger was second, 56.0% from 25 fixtures. Elo, the old chess rating, was down in fourth at 47.1%.
Eleven days later, Form was gone.
It's worth sitting with that, because it's one of the most important things this project has learned and it's easy to miss. Forty-one matches is a week of football. It is not a verdict. A model can have a great week by luck - so can a tipster, so can anyone. The only thing that separates skill from a good week is more weeks.
On 6 August Andrew removed Form from the predictions altogether. Its Discord channels were handed to Kestrel, a model added on 2 August that had been built from backtested evidence rather than a theory. Avenger was rebuilt around Alix.
The picks that were worse than a coin toss
The bigger decision came the day before, and it's a good example of how Andrew and Claude work. Andrew had a hunch. Claude measured it.
Every prediction carries a confidence tier - HIGH, MEDIUM or LOW. By early August there was enough settled data to ask whether the tiers meant anything. They did. Across 1,583 settled model predictions:
- HIGH-confidence picks were right 66.2% of the time (397 picks).
- MEDIUM: 51.5% (658).
- LOW: 40.3% (528).
LOW picks were right less often than you'd manage guessing between two outcomes. "I want to remove all low confidence picks from our system for all models," Andrew wrote. "As suspected, these come back with the lowest accuracy."
From 5 August, a LOW-confidence pick counts as no pick at all, everywhere. If Jackson is LOW on a match, Jackson doesn't predict that match - even when every other model is confident. The tiers are still measured on the performance pages, because otherwise the evidence behind the rule would disappear. But the system no longer acts on them.
Specialists
The same week, xG - expected goals, one of the original five - was taken off win-draw-lose predictions and given one job: both teams to score. And a new model was built for the goals markets: Osprey, the project's own goals model. Keep an eye on Osprey. It becomes important.
Andrew also asked for something to run every month: compare all the models, find the best combination, recalibrate the blend. "I want to drive as much accuracy out of this as possible."
Meanwhile, things broke
On 2 August Andrew noticed Wolves and Burnley climbing the power rankings - two teams that had just been relegated. They were still being ranked in their old league. Fixed.
On 5 August the dashboard stopped working altogether. Andrew pasted the error logs into the conversation. The cause was a library the app depends on, which had released an update overnight that broke it. The fix was to pin the old version.
On 31 July the system started posting to X automatically, on the same schedule as Discord. On 3 August the alerts were changed to go out only on match day: "The data gets lost if it's too far away from the event."
The lab book
- Retired: Form. Top of the first weekly leaderboard; gone eleven days later.
- Removed: every LOW-confidence pick, on the evidence - right 40.3% of the time.
- Demoted: xG, from results to both-teams-to-score only.
- Born: Kestrel (built from backtests) and Osprey (goals).
- Didn't work: relegated teams keeping their old league's rankings.
- Still testing: the monthly calibration - whether a blend that re-tunes itself to recent results actually beats the models it's made from.
Next stop: the weekend it all went wrong.