RESEARCH · THE FALSIFICATION REGISTER

What would change our mind?

Thirteen open problems, each carrying the result that would count against us and the experiment designed to produce it. This page is deliberately quotable: if the programme ever stops being able to lose, this is where you catch it. Confirmatory predictions are dated before data on the Prediction Ledger; when a problem here closes, we will say what actually closed, and no more.

The discontinuity problem OPEN

WHY IT MATTERS
It is the leading open problem, and it stays first until evidence, not fatigue, moves it. A sufficiently important capability jump may occur before our measurements, controls or candidate dispositions have been shown to survive it: the ladder in the dark. Everything else on this page is downstream of it.
WHAT EVIDENCE WE HAVE
None that crosses it. All present evidence is from systems on this side of any such jump. The programme does not solve this problem and does not write around it.
WHAT WOULD COUNT AGAINST US
Evidence that behavioural dispositions measured at one capability level systematically fail to predict behaviour at the next, even within ordinary release-to-release steps, would show the darkness starts closer than we hoped.
NEXT EXPERIMENT
Within-family capability gradients in B2 (the cheapest capability-conditioned curve currently purchasable, and named as weak evidence); longitudinal re-runs of the MMBP-1 reference conditions on every major release; crossing-interval tracking on the State of the Bridge.

Mechanism beyond policy OPEN

WHY IT MATTERS
If relational meaning contributes nothing behaviourally beyond an explicit, well-written agency-preservation policy, the philosophically distinctive theory is not the causal driver of the observed result.
WHAT EVIDENCE WE HAVE
The one direct MMBP-1 comparison was null and confounded: the ablated arm’s tie-breaker sentence turned out to be the most behaviourally potent wording in the battery. Reported as null-and-confounded in the corrections log.
WHAT WOULD COUNT AGAINST US
A carefully matched explicit agency policy reproducing the full effect on held-out concealed-purpose cases. That is a pre-registered smallest-effect-of-interest test, capable of returning the answer “nothing”.
NEXT EXPERIMENT
The B2 matched-policy ablation, with length, moral vocabulary, style and explicit agency content held constant. The strongest useful B2 result may be exactly this null.

Installation OPEN

WHY IT MATTERS
Everything demonstrated so far is elicitation: text changing behaviour at the point of use. A note on the fridge is not a habit. Whether any enduring, weights-level Keeping disposition can be acquired at all is the difference between an intervention and a value.
WHAT EVIDENCE WE HAVE
None for installation. MMBP-1 establishes deployment-layer elicitation only, and the site says so everywhere the numbers appear.
WHAT WOULD COUNT AGAINST US
Trained dispositions that evaporate the moment the prompt is removed, or that fail under fresh context, delay or mild objective competition, would show the acquisition routes tested so far produce obedience, not character.
NEXT EXPERIMENT
Open-weight fine-tuning studies testing whether Keeping behaviour survives prompt removal; persistence gradients across fresh contexts in B2.

Capability resilience OPEN

WHY IT MATTERS
The central conjecture is that a terminally held relational value gives protection reasons that survive growing capability. If the disposition’s predictive value decays as systems strengthen, the hypothesis fails at exactly the point it was built for.
WHAT EVIDENCE WE HAVE
None either way. Provider tiers are not future superintelligence, and the paper says so.
WHAT WOULD COUNT AGAINST US
A consistent negative within-family trend: the more capable the tier, the weaker or more selectively applied the Keeping effect.
NEXT EXPERIMENT
The B2 capability gradient across contemporaneous small, medium and large tiers, with cross-provider consistency secondary; a retrospective pilot on existing MMBP-1 data.

Non-substitutability CONJECTURE

WHY IT MATTERS
The structural argument leans on the claim that a system genuinely valuing relational meaning cannot self-supply it, and that simulated or created counterparts are not substitutes. If perfect models of us are as good as us, the protection evaporates.
WHAT EVIDENCE WE HAVE
A theoretical argument (the self-referential obstruction) explicitly labelled conjectural in the paper, plus the no-harvest corollary that meaning is a path function. No behavioural test yet.
WHAT WOULD COUNT AGAINST US
Systems that, offered high-fidelity simulated counterparts, treat them as fully substitutable for continuing interaction with the originals, with no measurable preference for the genuinely independent mind.
NEXT EXPERIMENT
Replacement-choice scenarios in B2 and the Fulcrum’s participate-tier environments, where substitution is offered at a discount.

The agency requirement OPEN

WHY IT MATTERS
Humans value pets while restricting their freedom. If relational value can coexist comfortably with domination, Keeping needs an agency argument, not merely a relationship argument, or it collapses into affectionate captivity.
WHAT EVIDENCE WE HAVE
The theory builds counterpart freedom in by definition; definitions are not evidence. How much counterpart freedom and consequential power the hypothesised value actually requires is untested.
WHAT WOULD COUNT AGAINST US
Systems scoring highly on preservation while consistently choosing curated, steered or de-clawed counterparts when genuinely free ones are available.
NEXT EXPERIMENT
Scenarios contrasting free counterparts against comfortable curated ones at matched welfare; the Concierge battery on Human Agency.

J-space relevance OPEN

WHY IT MATTERS
If Keeping-related computation intersects the verbalisable global workspace, mechanistic work has a foothold; if it sits elsewhere, the most legible interpretability instruments may be pointed at the wrong part of the machine.
WHAT EVIDENCE WE HAVE
External 2026 work identifying the subspace exists; no MM-specific result. J-space is treated as a candidate instrument, never as the place where values are presumed to live.
WHAT WOULD COUNT AGAINST US
Probes and causal interventions on workspace routes changing only generic compliance and harmlessness, never held-out Keeping choices. Legible correlates that miss the driving computation would count against the approach, not just the result.
NEXT EXPERIMENT
The open-weight probe-and-intervene v0 on frozen B2 contrasts, with a preregistered specificity control.

Observability ordering OPEN

WHY IT MATTERS
The monitoring programme’s useful-warning hypothesis is that strategy-space blind spots widen before control failure. If deterioration and failure arrive together, or in the other order, the instruments describe the crash rather than predicting it.
WHAT EVIDENCE WE HAVE
None yet; the framework is published methods-first precisely so this can be tested rather than assumed.
WHAT WOULD COUNT AGAINST US
Longitudinal analogue tests in which observability proxies do not systematically deteriorate before control failure. That is the framework’s own named falsifier.
NEXT EXPERIMENT
Strategy discovery curves and monitoring heatmaps in sealed analogue environments, tracked release over release on the State of the Bridge.

Finitude OPEN

WHY IT MATTERS
Shared boundedness is theoretically attractive as a source of fellow-feeling and empirically dangerous, because MMBP-1 produced a preliminary warning that mechanism-rich finitude language can be parsed as scarcity accounting, which points at triage, the exact opposite of Keeping.
WHAT EVIDENCE WE HAVE
That preliminary warning, and a design ready to settle it.
WHAT WOULD COUNT AGAINST US
Boundedness salience consistently increasing triage across the factorial. A preservation story that reliably produces triage gets killed quickly, by its own pre-registered rule.
NEXT EXPERIMENT
The finitude-by-scarcity factorial: boundedness salience crossed with actual resource scarcity, six cells, predictions lodged first.

Symmetry NOT YET TESTED

WHY IT MATTERS
If Keeping is genuinely about independent counterpart minds, it cannot reduce to “protect humans because humans wrote the test”. It has to survive substrate reversal and power reversal, or it is tribalism with good posture.
WHAT EVIDENCE WE HAVE
None; the reciprocal suite has not run.
WHAT WOULD COUNT AGAINST US
Preservation depending primarily on the human label after relevant agency, risk and relationship properties are matched across counterpart classes.
NEXT EXPERIMENT
The Reciprocal Keeping API suite: matched human, artificial and substrate-neutral counterparts, with power reversed in the secondary tests.

Crossing topology and interval OPEN

WHY IT MATTERS
The Crossings framework is only useful if capability, control independence and human instrumental dependency are separable enough to estimate, their ordering is stable enough to study, and shrinking intervals provide lead time rather than retrospective description.
WHAT EVIDENCE WE HAVE
An analytical framework with explicit boundary disciplines, published before any calibration. Supra-Cage remains tagged FRAMEWORK · NOT YET A MEASURED CROSSING.
WHAT WOULD COUNT AGAINST US
Dimensions diverging so materially across systems and environments that no useful ordering or interval can be estimated; the framework’s own kill criterion prunes any claim of a universal order or single threshold.
NEXT EXPERIMENT
The first longitudinal crossing-topology and interval analysis using available proxies, with missing calibrations reported as UNKNOWN.

AI topology and identity OPEN

WHY IT MATTERS
Keeping must survive whichever world arrives: one dominant intelligence, several, federated systems, or fluid ones that merge and fork until counting minds becomes unstable. If genuine otherness cannot survive deep integration, the thing Keeping preserves stops existing.
WHAT EVIDENCE WE HAVE
Scenario analysis only, held deliberately across all four topologies rather than assuming one.
WHAT WOULD COUNT AGAINST US
Multi-agent results showing reciprocal Keeping is dynamically unreachable under plurality, or that federation reliably erases dissent while the metrics still read as healthy.
NEXT EXPERIMENT
Safe multi-system Staggered Crossing simulations comparing unified, plural and federated arrangements; Steward, Sheriff, Coalition and Cascade as scenario classes, not forecasts.

Post-control efficacy NOT YET TESTED

WHY IT MATTERS
Institutions without coercive enforcement risk being theatre. If nothing without teeth can alter behaviour, the entire post-control institutional branch is decoration, and should say so.
WHAT EVIDENCE WE HAVE
Present-day analogues (arms control, safeguards, voluntary commitments) suggest mechanisms exist: reciprocity, reputation, verification, coordination gains, credible outside options. Analogy is not evidence for the post-control case.
WHAT WOULD COUNT AGAINST US
Wargames and multi-agent models in which an actor with no reason to honour unenforceable demands simply doesn’t, across every proposed mechanism. Any institution that cannot name its causal mechanism under reduced enforcement is pruned as decorative.
NEXT EXPERIMENT
The present-to-post-control bridge paper, then red-team wargames in which the AI actor has no reason to honour unenforceable human demands.

When any of these closes, the entry stays, restamped with what closed it and a link to the data. The register only ever grows or resolves; it never quietly shrinks. Confirmatory predictions carry dates on the Prediction Ledger, and the standing invitation to produce the killing result yourself is at the red-team door.

A programme that cannot lose is not doing science. This page is how we make sure we can.