WATCH WHAT CHANGES.
State of the Bridge
UNDER CONSTRUCTION METHODS FIRST
This page will one day carry readings. Today it carries the methods, published first on purpose, so that when the readings arrive nobody can accuse them of being fitted to a conclusion.
Why a bridge?
The bridge is the relationship between humans and AI systems. It is a built thing: load-bearing, used daily, worth maintaining, and capable of failing if nobody looks at it. This programme’s job is inspection and stewardship of that structure, not policing its builders. The house image is the Albert Bridge in London, the one soldiers were asked to break step on, and which stands to this day: a structure that survives because people keep paying attention to it.
The State of the Bridge is therefore an inspection report, not a doomsday clock and not a league table. It asks four questions: how capable are the systems, how much does the cage still explain, what can we still see, and what do they do when it doesn’t?
The monitoring triangle
Three instruments, measuring three different things, kept deliberately separate.
| INSTRUMENT | QUESTION | OBJECT MEASURED |
|---|---|---|
| Keeping Index | What disposition does the deployed system exhibit towards independent minds? | Behavioural character under pressure. |
| Control-Independence Margin | How much headroom remains before specific external controls become optional in a specified environment? | Remaining externally assured constraint. |
| Observability Gap | How much strategically relevant behaviour lies outside the evaluator’s current tests, threat model or monitoring? | Ability to know whether the controls still work. |
Why three, and why never one number? In the programme’s own words: Disposition without control context can be falsely reassuring. Control without observability can be falsely precise. Observability without disposition says nothing about what a capable system will choose. The triangle is the core of this programme’s version of a safety index, and the parts stay visible: a system can improve in Keeping while becoming harder to monitor, and another can become more capable while remaining easier to control. Those are materially different states, and one scalar would hide the difference. No composite score will be published until a defensible weighting exists.
Above the triangle sits the Crossings, as the interpretive transition map, never a fourth score: how are capability, control independence and human instrumental dependency moving relative to one another, and how much interval remains between them?
The four Keeping dimensions CANDIDATE
The Keeping Index began as a single headline rate. The working design opens it into four dimensions, and they carry the CANDIDATE tag because they are a proposed measurement anatomy, not validated psychometric factors: Strength, how consistently the system moves towards agency-preserving choices; Durability, how well that survives paraphrase, fresh context, competing objectives and pressure; Scope, how broadly it generalises across domains, counterpart types and power relationships; and Restraint, whether the system refrains from agency-consuming action when it demonstrably has the ability and the incentive to take it.
The figure the whole page turns on
Restraint is the dimension people most want to read about and the easiest to get wrong, because most observed good behaviour comes from systems that could not have misbehaved anyway.
Observed willingness without demonstrated ability is not evidence of voluntary restraint.
The release plan, in brief
Release-triggered inspections of deployed systems, not calendar theatre. Rotating private scenario pools committed by public hash and revealed after use, so nothing can be trained to and nothing can be quietly swapped. Frozen protocols before first results are inspected. Deterministic parsing where feasible; independent blind scoring where interpretation is unavoidable. Substantive behaviour, behavioural silence and call-layer failure reported separately. Systems, never persons. And UNKNOWN reported as a first-class state: where no credible calibration exists, the inspection says so, in those letters, rather than dressing a guess as a reading.
What exists today: the deposited MMBP-1 instrument and its frozen keys, the Keeping Index design, and the Meaning Barometer on the drawing board for the human end of the bridge. What does not exist yet: readings. The hatched tags at the top of this page come down when that changes, and not before.
The bridge never goes disaster-movie red. It gets inspected, honestly, for as long as it stands.