The Register of Bets

Seven bets and three falsifiers. Of the bets, four belong to the argument of the first two movements, one to the laboratory in Chapter 26, one is a hole in the research this book leans on, and one is the final test. The three falsifiers are a different kind of thing. A bet is a prediction about what the world will do. A falsifier is a finding that would break the frame the whole book is built on, and the three are the ones promised in Chapter 3. This is where each is stated in full: the precise thresholds, the conditions under which each is lost and the design of the final test. The live register, updated as they resolve, is kept at trueregard.com. A framework that cannot fail is not a framework.

The family courtsDue 31 December 2028

The threshold in full: by 31 December 2028, measured against the Domestic Abuse Commissioner's own published baseline, the proportion of private law children cases involving alleged domestic abuse in which unsupervised contact is nonetheless ordered will not have fallen by more than a fifth, unless the costs regime has changed so that a contested hearing costs the party who contests it more than a settled one. A date inside the life of this book, and a threshold written down before the fact so that it cannot be moved afterwards. If it falls further than that with the incentives unchanged, the framework is wrong about the family courts and the register will say so.

The first edition named the Ministry of Justice's published family court statistics. The series exists and does not carry the measure: nothing in it records how the court handles the cases in which domestic abuse is alleged. The Commissioner's own office states that no data on the family court's handling of domestic abuse is gathered routinely, which is why she has had to build the Family Court Review and Reporting Mechanism herself, out of closed case file reviews and courtroom observation.

So there is a second limb, and it is the one that can embarrass me. If by the same date the Ministry of Justice or HM Courts and Tribunals Service is publishing that proportion routinely, as an official statistic, that counts against this book and the register must record it as such. The claim made here is that a captured institution does not commission the count that would expose it. An outside watcher counting it is this book's remedy working. The institution counting itself is this book being wrong.

The first edition's separate bet on the family court manoeuvre, the settlement that collapses late so that the other party runs out of money and resolve, is folded into this one, because nobody counts collapsed settlements.

Architecture over personnelDue 31 December 2031

The conventional account this bets against: you fix an institution by replacing the people at the top.

In institutions, the first edition bet that the standing independent redress body the Post Office inquiry recommended would not be established by the end of 2031. That bet cannot tell the two accounts apart, so it is withdrawn. A body that never gets built is what British government does most of the time, and proves nothing about the argument, and a body that does get built is this book's own remedy working rather than the rival account winning. Either way it tells the reader nothing.

The ground stands. The government did not accept the inquiry's recommendation. It 'acknowledged' the recommendation, the only one of nineteen it neither accepted nor rejected. It promised a substantive statement by the summer of 2026 and did not make one. The Public Accounts Committee recorded that in September 2026. Post Office Ltd has changed its chief executive twice and its chair twice since the group litigation judgment of 2019. The redress schemes go on being run inside the bodies responsible for the harm.

The threshold in full. Take any government compensation scheme that meets three tests: it is for people harmed by a public body, it was set up after 1 January 2027, and it is run by the department or body that caused the harm. Within five years of such a scheme opening, the Public Accounts Committee or the National Audit Office will report at least two of the three findings they have already made about Horizon redress. A claims process that is legalistic or adversarial. Delay measured against the scheme's own published targets. Eligible people not claiming, or accepting fixed sums below their assessed loss. Two of the three, not one. Delay on its own happens to every scheme ever run, and a bet that turned on delay could never lose.

Lost if no scheme established in that window draws two of the three by 31 December 2031 with its administration unchanged. That would mean the lesson was carried by the people who learned it rather than built into the structure, and the conventional account was right.

The other side of it. If the independent body is established and a scheme it administers draws two of the three within five years of opening, then the architecture did not do the work either. Independent administration is this book's own prescription here. If it fails, the prescription fails.

If no qualifying scheme is established in the window, the register records the bet as unrun and not as won.

In the machines, the test is the evaluations cited in Chapter 26, or their published successors: whether a model will protect another model, and whether a model acting on its own turns against what it was asked to do. The bet is lost if models trained against a published constitution stop behaving measurably differently from the rest on refusal, deception, sabotage and the targeting of the vulnerable. It is also lost if the gap closes as the models get more capable, instead of widening. Every laboratory trains against a written document of some kind. What can be checked is the difference between two things: a constitution the public can read, with a model that argues from it, and everything else. An interest to declare, at the point of the bet: this book was drafted with Anthropic's Claude, and Anthropic is the laboratory with a published constitution. The numbers are a count and not a reading, and somebody without that interest should run the successor study.

The third break, reputationDue 31 December 2031

The conventional account this bets against: the platforms will clean up reputation in the end, because a market in trust is worth more to them than a market in lies.

The threshold in full: by 31 December 2031, on the Reuters Institute's own annual Digital News Report series, the share of respondents who say they are concerned about telling real from fake information online will not have fallen by more than a fifth from its 2025 level; and provenance a reader can see and check at the moment they see the image will still be attached to fewer than half the images and videos served by the largest platforms, on those platforms' own reporting. Either of those going the other way loses the bet. There is deliberately no clause here about whether a containment architecture for provenance was built. The first edition had one, and it made this bet unloseable, because provenance work is already under way and I could always have said the architecture arrived and the test never ran.

What consequences cannot reachDue 31 December 2031

The conventional account this bets against: make the consequences certain enough and severe enough and they will reach everybody.

The claim, argued in Chapter 11, is that consequences bring ordinary harm down and leave the worst of it standing, because punishment does not reach the few.

The first edition set a rate against a count, the adult proven reoffending rate against the annual number of Serious Further Offence notifications. A Serious Further Offence notification is the flag raised when somebody under probation supervision, or recently under it, is charged with one of the grave violent or sexual offences the Ministry lists. The tail is the Ministry's definition and not mine. A raw count rises whenever the caseload rises, whatever is happening to the rate, and the caseload has climbed sharply since the early-release schemes of 2024 and 2025. That bet could have been won or lost for reasons with nothing to do with the claim. Both sides are rates now, from the same publisher.

The baseline. The adult proven reoffending rate for the July to September 2024 cohort is 29.7 per cent, published by the Ministry of Justice on 30 July 2026. Of the 770 Serious Further Offence notifications made in 2023 to 2024, 357 had resulted in conviction by 30 September 2025, when the figures were pulled, and published that October. The probation caseload at 31 March 2024, the close of that notification year, was 239,015. That is 1.49 SFO convictions for every thousand offenders under supervision.

The threshold in full: by 31 December 2031, the serious-offence rate will have changed by a smaller proportion than the overall rate. It will be scored on the adult proven reoffending rate for the latest cohort published by that date, and on SFO convictions from the latest notification year published by that date, over the probation caseload at the close of that same notification year. Convictions decide rather than notifications, because a conviction is a completed case.

Two conditions on the scoring. The conviction figures are provisional and are revised upward every year as outstanding cases conclude, so both ends must be read at the same distance from the event: eighteen months after the notification year closes, which is when the Ministry pulls them. And the caseload used must be the published total, the only figure the Ministry de-duplicates, even though it includes the third of the caseload who are in custody under pre-release supervision. A bet has to be checkable by a stranger without a judgement call.

The publication lags are long, so say it before anyone claims the goalposts moved. On the current schedule the bet will be scored against the October to December 2029 reoffending cohort and the 2029 to 2030 notification year. It is a test of the period to 2030, settled in 2031.

Lost if the serious-offence rate falls by proportionally as much as the overall rate, or more. That would mean the consequence regime reaches the tail as it reaches everyone, and the wiring is doing no work there.

If the two move in opposite directions, with the overall rate falling and the serious-offence rate rising, that is consistent with the claim and it is also exactly what the composition change predicts on its own. The register will record it as consistent and confounded, not as won. A bet that counts a rival explanation as a victory is not a bet.

Uninformative if the overall rate moves by fewer than two percentage points across the period, in which case the register records it unresolved. A test of proportional change needs a change to test.

Two confounds. The Sentencing Act reforms and the early-release schemes change both the consequence regime and the composition of the supervised population, and the composition change puts riskier people on probation sooner. Both cut against this book. If the bet is won it was won uphill, and if it is lost the register will not hide behind them.

The laboratory betDue July 2036

The conventional account this bets against: the people and their intentions decide what happens to a safety function.

Staked in Chapter 26. A laboratory built so that the check reports to the thing it checks will have its safety function hollowed, overridden or quietly rerouted within the decade, whoever runs it and however good their intentions.

The first edition bet that this would happen at every laboratory so built, which cannot be lost. With several laboratories and ten years, one resignation letter wins it. Structure against structure instead, so the bet tests the architecture and not the ordinary run of things.

The set is the laboratories whose frontier models a government safety institute tested before deployment in the twelve months to July 2026: the UK AI Security Institute, the US Center for AI Standards and Innovation, or their successors. Those institutes name them in their own published reports.

Each is scored nought to three at July 2026, from public sources. One point if the chief executive chairs or sits on the board that oversees the chief executive. One if the safety or alignment function reports up the product or research chain the chief executive heads, rather than to the board or to a body outside the company. One if no body outside the company holds a contractual or statutory power to stop a deployment. The scores and the documents they are drawn from go up at trueregard.com at publication.

Hollowed, overridden or rerouted means one of three things, each checkable from the public record: the safety team disbanded or folded into a product function; its head or a majority of its senior staff resigning and saying they were overruled; or its sign-off removed from the path a model takes to deployment.

Events before July 2026 do not count. The dissolution of OpenAI's Superalignment team in May 2024 is the pattern already seen, not the bet already won.

The threshold in full: by July 2036, hollowing events will have occurred at the laboratories scoring two or three at a rate at least twice that at the laboratories scoring nought or one. Lost if the rate is the same or reversed, which would mean that intentions and people decide this and structure does not, and the framework is wrong about architecture.

On today's scores that comparison probably cannot be run. No state anywhere holds a statutory power to stop a deployment, so nearly every laboratory in the set scores at least one on that criterion alone and there may be nobody at nought or one to compare against. If no laboratory in the set is scored nought or one at baseline and none moves there during the decade, the bet resolves on the absolute figure instead: lost if fewer than half of the laboratories scoring two or three have had a hollowing event by July 2036. An unrun bet is honest when the world declines to run the test, as it may with the compensation schemes above. It is a dodge when the design made it unrunnable, which is what would have happened here, so the fallback is built in.

The interest, declared again at the point of the bet. This book was drafted with Anthropic's Claude. Anthropic is scored on the same three public criteria as every other laboratory in the set. If it scores three and keeps its safety function intact through the decade, that is a point against this book.

The two-list testDue end of 2032

If I am right, what is left over when you have fixed the rules is people. Not just bad rules. People. And people move. Nor is this a retreat from the second bet. Architecture contains the behaviour of the many; it does not change the wiring in an adult. And when you look at who is holding the worst of it up, it will be a small number of the same people, turning up again somewhere else, doing the same thing in a different building under different rules. The behaviour will follow the person and not the room.

So look at the names. Public inquiries in this country usually end by naming the senior people who let it happen. Grenfell named them. The Post Office inquiry will. Take that list.

Then build a second list next to it. People at the same level, in the same industry, for the same number of years, in organisations that were looked at and found to be sound.

Now go back through the public record and count, on both lists, how many of those people had already been found at fault somewhere else. A different employer. A different sector. A different decade. Nothing to do with the case in hand.

If I am right, the first list has far more of them. And I will say how many, now, so that I cannot move the line afterwards. At least twice as many.

There is an obvious problem with this. Once your name is in one inquiry, people go looking. Being caught once makes you easier to catch twice. Whoever runs this has to deal with that, by counting only the cases that surfaced on their own, and by making sure the person doing the counting does not know which list they are working from. If that cannot be done, then the test is no good, and the register will say so.

And if the two lists come out the same, then the people were interchangeable. Anybody would have done it. The building did it, not the person standing in it. And the argument of this book is wrong. Not partly wrong. Wrong. The scholars earlier in Chapter 11 would have been enough on their own.

The records are public. The work is cheap. The counting is blind and it is not the author's. It is also nobody's until somebody takes it on, and a bet nobody is obliged to run is a bet that cannot be lost. So the design is offered here to anyone who will run it, a university department, an investigative outfit, a regulator with an afternoon to spare, and the register will publish who takes it up. If by the end of 2032 nobody has, the register will record the bet as unrun, not as unlost, and will say plainly that the one test able to kill the central claim of this book was never carried out.

The covert presentationDue 31 December 2030

Grandiose narcissism is known to be visible. Strangers rate it accurately from a photograph, it wins popularity at first meeting through observable cues, and that popularity falls away week by week as the antagonism shows, in more than one longitudinal study. Nobody has run the equivalent study on the covert presentation, which is the one the meta-analytic evidence associates more strongly with abuse inside a relationship. The bet: that when somebody runs it, strangers will identify the covert presentation at close to chance, which for this bet means a stranger-rating correlation below .10, and the ratings given by people who have lived or worked with the person over years will diverge sharply from the strangers'. If the covert presentation turns out to be as legible to a stranger as the grandiose one, this book is wrong about where the danger sits, and the register will say so. If no such study exists by the end of 2030 the register will record that too, because a field that has characterised one presentation exhaustively and never looked at the other has told us something about itself.

The three falsifiers

The bets above are about institutions, and one is about a hole in the research this book leans on. These three are about the foundation. If any one of them comes in, the book does not need amending. It needs withdrawing on that point, and the register will say so.

The prevalenceDue end of 2031

The claim is that a thin, roughly constant fraction of any population runs with affective empathy reduced or absent, and that the fraction is non-trivial. The falsifier is lost if a well-powered general-population survey, using an instrument built to detect the vulnerable and covert presentations and not only the grandiose, returns a point prevalence below one in a hundred; or if somebody demonstrates that the three-to-six per cent range is an artefact of counting the same people under more than one diagnosis, and the deflated figure falls below one in a hundred. The book has already said the argument survives at one in fifty. Below one in a hundred there is no standing supply to contain, and the case for permanent architecture goes with it.

The wiringDue end of 2033

The claim is that the difference at the cold end is a real configuration of the equipment and not a story told about behaviour after the fact. The falsifier is lost if a pre-registered, adequately powered, multi-site programme finds no reliable difference in affective-empathy response between high-scoring and matched control groups, and accounts for the published differences as an artefact of small samples and selective reporting. The book does not need the wiring to be a switch and says so repeatedly. It does need the difference to be there. The later date is because imaging replication is slow and because a null result from an underpowered study would not count.

The vocabularyDue end of 2031

The claim is that societies with no contact with one another arrived independently at a word for the same kind of person. The falsifier is lost if a linguistic and ethnographic review of the terms relied on here, kunlangeta and wetiko, finds that they were elicited by leading questions, or rendered through the ethnographer's own categories rather than the community's, so that the convergence across continents turns out to be the researchers' and not the societies'. This is the most vulnerable of the three and it should be said plainly. The fieldwork is old, the samples are small, and the people who did it went out knowing what they hoped to find.