BEHAVIOURAL INTELLIGENCE · THE LICHESS.ORG STUDY
Your AI spots something unusual.
Is that enough to act?
A payment looks suspicious. An account behaves differently. A customer suddenly changes course. An alert gets your attention. The harder question is whether the evidence justifies what happens next.
We used public chess games from Lichess (lichess.org) to put that question under pressure. Chess records decisions in order, with a clear view of the choices available. You don’t need to know the rules to understand what we learned.
739.6m
source games scanned across three experiments
8,914
games analysed in depth across those experiments
2 no-go results
kept on the record when the requirements weren’t met
THE IDEA WE TESTED
An average can hide the story.
Two people can make the same number of unusual decisions. One spreads them over time. The other changes abruptly. A single average treats them alike; the sequence gives you another question to investigate.
CADi tested whether those patterns, measured against a player’s own history, added useful information, and whether that information held up.
Same count. Different sequence.
WHAT THE TESTS ACTUALLY SAID
A promising signal. Two reasons to keep testing.
EXPERIMENT 01 · FIRST TEST PASSED
Does it still work on data it hasn’t seen?
A promising start survived a fresh month.
We fixed the model before opening the next month’s data. It could score 389 of 401 eligible games, left 12 unscored and kept its alert volume and measured patterns within the agreed limits.
That established one month of stability. It did not establish whether any alert was right.
EXPERIMENT 02 · NO-GO
Does a more complex model earn its place?
The order of decisions revealed something the average missed.
We compared four ways of reading the same behaviour. The sequence of decisions carried information that a whole-game average missed. But none of the four approaches passed every requirement, including the agreed scoring-cost limit.
A later cost study found sequence scoring was very fast. The original failed result still stands; a new test must earn a new decision.
MEASURED · EXPERIMENT 02
Did the same people stand out next month?
Dashed line: required persistence of 0.30. Higher means more consistent rankings.
EXPERIMENT 03 · NO-GO
Does the result hold beyond a selected group?
The wider population changed the answer.
A broader sampling design exposed thin individual histories and weak consistency in player rankings between months. Some groups also fell below the minimum coverage requirement, even though overall coverage exceeded 95%.
A healthy headline average can conceal the very people your model struggles to serve.
MEASURED · EXPERIMENT 03 · MARCH
The headline looked healthy. One group was being missed.
Required: 90%. Passed.
Required: 80%. Failed.
PLANNED EXPERIMENT 04 · ISAMBARD-AI
What happens when you stop sampling?
We have already secured 20,000 GPU-hours on Isambard-AI and plan to use them to analyse every game in the planned corpus: 1,550,748,649 inventoried public games, with one frozen CADi model applied to every later eligible decision.
1.55bn
public games inventoried for the planned corpus
20,000
secured Isambard-AI GPU-hours
Every move
planned eligible-decision analysis without player, game or move caps
THE QUESTION
Was the signal weak, or did we see too little history?
Experiment 3 could only observe a thin slice of each person’s play. That may have buried a stable behavioural pattern in measurement noise. The larger experiment is designed to separate those two explanations.
THE DECISION RULE
Make the pattern earn its credibility.
We will measure whether a person’s behavioural signature becomes more stable when it is built from 2, 5, 10, 20, 50 and every available game. A rising curve would support the sampling explanation, but the result must also clear the pre-agreed reliability threshold with sufficient confidence. Otherwise the experiment can fail or remain inconclusive.
PLANNED TEST · NO RESULTS YET
Give the same instrument more history.
2
games / player-month
5
games / player-month
10
games / player-month
20
games / player-month
50
games / player-month
All
games / player-month
Planning status: the protocol candidate has passed independent review and 20,000 GPU-hours are secured. Execution planning remains open. No population-scale run has begun, and the experiment will not create cheating labels that the public data does not contain.
WHAT THIS MEANS FOR YOUR BUSINESS
Before you act on the score,
ask what earned it.
Ask to see performance on untouched data, the people the model cannot assess, and the rule that would stop it going live. Then test whether its alerts improve real decisions at an acceptable cost.
These experiments established parts of a repeatable evidence process. They did not establish cheating detection, causal effects or commercial returns. The public games have no verified assistance labels, so an unusual pattern cannot tell us whether someone cheated.
This research informs CADi Player’s Integrity work: building evidence that helps a human reviewer investigate unusual behaviour. Isambard tests whether longer histories make that evidence more reliable. A separate blinded partner evaluation will test whether it improves real review decisions. Neither is a completed result.
01 / MEASURE
Does the signal hold up?
Planned Isambard research tests reliability across longer histories.
02 / REVIEW
Does it help someone decide?
Seal predictions before revealing partner findings. Measure useful cases, disagreements and the cost of review.
03 / LEARN
What changed after action?
A later operator evaluation must follow interventions and outcomes. A case closure alone does not prove assistance or benefit.
Chess explores one part of Player: Integrity. Care and game protection need their own evidence and evaluation.
Explore CADi Player