Skip to main content
Case studies

BEHAVIOURAL INTELLIGENCE · THE LICHESS.ORG STUDY

Your AI spots something unusual.
Is that enough to act?

A payment looks suspicious. An account behaves differently. A customer suddenly changes course. An alert gets your attention. The harder question is whether the evidence justifies what happens next.

We used public chess games from Lichess (lichess.org) to put that question under pressure. Chess records decisions in order, with a clear view of the choices available. You don’t need to know the rules to understand what we learned.

739.6m

source games scanned across three experiments

8,914

games analysed in depth across those experiments

2 no-go results

kept on the record when the requirements weren’t met

THE IDEA WE TESTED

An average can hide the story.

Two people can make the same number of unusual decisions. One spreads them over time. The other changes abruptly. A single average treats them alike; the sequence gives you another question to investigate.

CADi tested whether those patterns, measured against a player’s own history, added useful information, and whether that information held up.

Same count. Different sequence.

Illustration: four highlighted decisions out of twelve in each row, spread out above and clustered at the end below. This explains the idea; it is not a study result or evidence of wrongdoing.

WHAT THE TESTS ACTUALLY SAID

A promising signal. Two reasons to keep testing.

EXPERIMENT 01 · FIRST TEST PASSED

Does it still work on data it hasn’t seen?

A promising start survived a fresh month.

We fixed the model before opening the next month’s data. It could score 389 of 401 eligible games, left 12 unscored and kept its alert volume and measured patterns within the agreed limits.

That established one month of stability. It did not establish whether any alert was right.

EXPERIMENT 02 · NO-GO

Does a more complex model earn its place?

The order of decisions revealed something the average missed.

We compared four ways of reading the same behaviour. The sequence of decisions carried information that a whole-game average missed. But none of the four approaches passed every requirement, including the agreed scoring-cost limit.

A later cost study found sequence scoring was very fast. The original failed result still stands; a new test must earn a new decision.

MEASURED · EXPERIMENT 02

Did the same people stand out next month?

Whole-game average0.218
Ordered decisions0.408

Dashed line: required persistence of 0.30. Higher means more consistent rankings.

Spearman correlation across October and November, 68 recurring players. Ordered decisions cleared this requirement; the overall experiment still failed its compute gate. These are consistency measures, not detection accuracy.

EXPERIMENT 03 · NO-GO

Does the result hold beyond a selected group?

The wider population changed the answer.

A broader sampling design exposed thin individual histories and weak consistency in player rankings between months. Some groups also fell below the minimum coverage requirement, even though overall coverage exceeded 95%.

A healthy headline average can conceal the very people your model struggles to serve.

MEASURED · EXPERIMENT 03 · MARCH

The headline looked healthy. One group was being missed.

Overall coverage95.97%

Required: 90%. Passed.

Lowest required subgroup73.88%

Required: 80%. Failed.

Weighted scoring coverage, March 2025. Each marker shows that row’s required minimum. Coverage measures who could be assessed; it says nothing about whether an alert was correct.

PLANNED EXPERIMENT 04 · ISAMBARD-AI

What happens when you stop sampling?

We have already secured 20,000 GPU-hours on Isambard-AI and plan to use them to analyse every game in the planned corpus: 1,550,748,649 inventoried public games, with one frozen CADi model applied to every later eligible decision.

1.55bn

public games inventoried for the planned corpus

20,000

secured Isambard-AI GPU-hours

Every move

planned eligible-decision analysis without player, game or move caps

THE QUESTION

Was the signal weak, or did we see too little history?

Experiment 3 could only observe a thin slice of each person’s play. That may have buried a stable behavioural pattern in measurement noise. The larger experiment is designed to separate those two explanations.

THE DECISION RULE

Make the pattern earn its credibility.

We will measure whether a person’s behavioural signature becomes more stable when it is built from 2, 5, 10, 20, 50 and every available game. A rising curve would support the sampling explanation, but the result must also clear the pre-agreed reliability threshold with sufficient confidence. Otherwise the experiment can fail or remain inconclusive.

PLANNED TEST · NO RESULTS YET

Give the same instrument more history.

2

games / player-month

5

games / player-month

10

games / player-month

20

games / player-month

50

games / player-month

All

games / player-month

Study design, not a forecast: the blocks illustrate increasing history, not proportional game counts or predicted reliability. We will measure the answer at each step using one frozen model.

Planning status: the protocol candidate has passed independent review and 20,000 GPU-hours are secured. Execution planning remains open. No population-scale run has begun, and the experiment will not create cheating labels that the public data does not contain.

WHAT THIS MEANS FOR YOUR BUSINESS

Before you act on the score,
ask what earned it.

Ask to see performance on untouched data, the people the model cannot assess, and the rule that would stop it going live. Then test whether its alerts improve real decisions at an acceptable cost.

These experiments established parts of a repeatable evidence process. They did not establish cheating detection, causal effects or commercial returns. The public games have no verified assistance labels, so an unusual pattern cannot tell us whether someone cheated.

This research informs CADi Player’s Integrity work: building evidence that helps a human reviewer investigate unusual behaviour. Isambard tests whether longer histories make that evidence more reliable. A separate blinded partner evaluation will test whether it improves real review decisions. Neither is a completed result.

01 / MEASURE

Does the signal hold up?

Planned Isambard research tests reliability across longer histories.

02 / REVIEW

Does it help someone decide?

Seal predictions before revealing partner findings. Measure useful cases, disagreements and the cost of review.

03 / LEARN

What changed after action?

A later operator evaluation must follow interventions and outcomes. A case closure alone does not prove assistance or benefit.

Chess explores one part of Player: Integrity. Care and game protection need their own evidence and evaluation.

Explore CADi Player
Discuss a decision you need to trust

Research summary · 5 September 2026. Based on CADi’s Experiment 1 and 2 wrap-ups, trajectory-cost study and Experiment 3 readout, consolidated in the 3 September programme report. Scan totals describe archive processing; only the stated samples received detailed engine analysis.

Source: Lichess open database (CC0). Independent CADi research using public data; no Lichess partnership or endorsement is implied.