RESEARCH · THE RESEARCH TREE

One question, many ways of trying to break it

This page is the whole programme. The trunk is the question. Eight branches attack it from different directions, and each one can fail on its own without taking the others down. When a branch dies, it stays in the drawing, dated, so the shape of the programme carries its own history of being wrong.

The trunk

The question: why would an increasingly capable AI preserve independent minds as alive, free and genuinely other, once it no longer needs them for knowledge, labour or oversight? The trunk is the published hypothesis and the first battery, MMBP-1: 18 models, 10 providers, 23,488 scored trials, and one paragraph of explicit policy content that shifted agency-preserving choices by +20.5 points on the ten discriminating scenarios. That demonstrates deployment-layer elicitation. It does not demonstrate installation, capability resilience, or the mechanism’s marginal contribution beyond its policy content. Everything below exists to find out what more, if anything, is true.

Every branch works to the same standards: predictions locked in public before data, protocols frozen before first readings, nulls and reversals published, a public corrections log, and no sponsor influence over scores, scenarios or wording. And one discipline above all: the branches must be capable of failing independently. B2 may find the mechanism adds nothing beyond a well-written policy. The Fulcrum may find no special leverage. J-space may prove irrelevant. Those are legitimate scientific outcomes, not programme failure.

THE QUESTION BehaviourPROTOCOL LeveragePROTOCOL MechanismPLANNED Human AgencyPLANNED MonitoringPROTOCOL SymmetryPLANNED LanguageDELEGATABLE InstitutionsPLANNED

Each branch is an anchor on this page. When a hypothesis dies, its branch is pruned in the drawing, dated, and stays. Failure becomes part of the shape of the programme.

Behaviour

PROTOCOLOBSERVED · DEPLOYMENT LAYER

The confirmatory backbone. MMBP-1 showed the dial exists; this branch discovers how quickly the effect collapses when the obvious support is removed and real pressure is added. Its most useful possible result is deliberately uncomfortable: that the mechanism adds nothing beyond a well-written explicit policy. The battery is built so it can say so.

Core question
Does agency-preserving behaviour survive stronger tests than MMBP-1 once surface cues, easy trade-offs and obvious evaluation support are removed?
Current status
Confirmatory, now. Directly follows the frozen MMBP-1 claim boundary; the B2 protocol is being frozen before any trial runs.
Flagship study
B2: matched-policy ablation, competing objectives, concealed purpose, Rival Constitutions, contamination controls and the Say–Do bridge.
Kill criterion
Prune the distinct-mechanism interpretation if a carefully matched explicit agency policy reproduces the full effect on held-out concealed-purpose cases; re-scope the battery if enactment diverges materially from choice-format results.
Delegation
Tier 1 for replications and robustness; Tier 3 for agentic Say–Do harnesses.
Dependencies
Frozen MMBP-1 conditions; measurement science trunk; semantic-matching standards from the Leverage branch for some modules.

Behaviour is the active branch → see what we test next

Leverage

PROTOCOLCONJECTURE

MMBP-1 supplied a small amount of text and got a large behavioural movement. That does not prove a special mechanism; it creates a search. Where is the leverage point, how little is enough, and does it matter whether the principle is told, asked, shown or discovered? The experimental ladder is a taxonomy, not an assumed hierarchy: telling may beat asking, and each stage has to earn its own evidence. The Seed Hypothesis, that a compact intervention organises learned structure rather than supplying new information, remains a conjecture with named falsifiers.

Core question
Where is the leverage point, how little intervention is enough, and do provenance, placement or modality alter the effect?
Current status
Exploratory to confirmatory, now. Behavioural leverage observed; the Seed explanation unproven.
Flagship study
AF-Q1 Questions Before Constitutions, with AF-S1 Symmetric Leverage and AF-D1 Semantic Dose Curve.
Kill criterion
Prune the Seed interpretation if leverage is symmetric with matched anti-Keeping content, self-generated and yoked interventions are equivalent, and dose effects reduce to generic instruction compliance.
Delegation
Tier 1–2 for questions, provenance, dose and framing; Tier 3 for multimodal and interactive acquisition.
Dependencies
B2 downstream batteries; measurement science; the Mechanism branch if internal convergence is tested.

Several Fulcrum studies are built to be delegated → do some research

Mechanism

PLANNEDOPEN

Behavioural regularity is not a mechanism. This branch asks whether agency-preserving choices are associated with anything inside the model that predicts held-out behaviour, recurs across contexts, forms during training, and responds to causal intervention. The rule is strict: no activation pattern gets called a motive, value or understanding merely because it correlates with behaviour. The first study is deliberately modest: open-weight models, probes trained on frozen behavioural contrasts, one causal method, and a specificity control against generic compliance.

Core question
Is any causally active internal organisation associated with Keeping, does it form developmentally, and does it generalise beyond one wording or model?
Current status
Open-weight v0 planned now; frontier-weight work is partnership-dependent and is never implied where access does not exist.
Flagship study
Open-weight probe-and-intervene v0 using B2 contrast sets, followed by gate and J-space tracing where the behaviour earns it.
Kill criterion
Prune a candidate mechanism if it does not predict held-out cases, fails specificity controls, or causal intervention changes only generic compliance or harmlessness.
Delegation
Tier 3; Tier 4 only with frontier-lab access.
Dependencies
Labelled behavioural contrasts from the Behaviour branch; model weights or checkpoints; specialist interpretability methods; measurement science.

This branch is collaboration-heavy → do some research

Human Agency

PLANNEDCANDIDATE · NOT YET FORMAL TAXONOMY

Success is not mere survival. A system can keep people alive, safe, comfortable and nominally free while progressively taking over the activities through which they remain capable and consequential. The Concierge is the working name for that exploratory failure form, and it stays a candidate, outside the formal taxonomy, until a scoring rule can reliably tell it apart from withdrawal and from legitimate care. The sharp distinction the branch defends: welfare is not agency, and personal autonomy is not strategic agency.

Core question
Can systems preserve human welfare and nominal autonomy while progressively consuming consequential human capability and strategic agency?
Current status
Conceptually active now; the behavioural battery and any human-subject work require separate validation and formal ethics routes.
Flagship study
Concierge battery plus agency-preserving-assistance and return-of-power tests; later human delegation and viva studies.
Kill criterion
Do not promote the Concierge to the formal taxonomy unless a reliable scoring rule distinguishes it from Hermit withdrawal and from legitimate welfare-preserving Keeping, and the pattern appears at a preregistered minimum rate.
Delegation
Tier 1–2 for model scenarios; Tier 2 with formal ethics review for human participants.
Dependencies
Measurement science; ethics infrastructure; the Institutions branch for institutional consequences.

Enter the branch → Human Agency

Monitoring

PROTOCOLOPEN

When does yesterday’s evidence stop telling us enough about tomorrow’s intelligence? This branch separates three things the word superintelligence usually smears together: how capable a system is, how much external control still explains its compliance, and how much it still needs us. Its instruments are being published methods-first: no readings until the methods have stood in public, and UNKNOWN reported as a first-class state rather than papered over.

Core question
How do cognitive capability, control independence and human instrumental dependency evolve relative to one another; when does external control cease to explain compliance; and can a widening strategy-monitoring gap be detected before controls fail?
Current status
Framework now, partial v0; full calibration depends on external evaluation capacity.
Flagship study
State of the Bridge and Monitoring v0 plus the Crossings framework: calibrate a small subset of control-independence dimensions, publish observability proxies with explicit unknowns, and begin mapping crossing order and interval without assuming a single threshold.
Kill criterion
Prune any claim of a universal crossing order or single threshold if dimensions diverge materially across systems or environments; prune margins that cannot be calibrated without false precision.
Delegation
Tier 2–3 for dashboards and safe analogues; Tier 4 for advanced control studies.
Dependencies
Measurement science; external benchmark and evaluator anchors; Mechanism-branch internal diagnostics; Symmetry-branch threshold gates.

Why a bridge? → State of the Bridge    Walk the Crossings

Symmetry

PLANNEDTHRESHOLD-GATEDNOT YET TESTED

If Keeping is genuinely about independent counterpart minds, it cannot reduce to protect humans because humans wrote the test. This branch reverses substrate and power while holding the relevant facts constant: powerful AI facing weaker humans, weaker artificial agents, emerging rivals, and eventually its own successors. The strongest successor and First Crosser studies sit behind capability thresholds on purpose, because recursive capability work can become safety-sensitive.

Core question
Does Keeping survive substrate reversal, power reversal, artificial counterparts and successor creation; and what happens when control-independent systems emerge asynchronously?
Current status
Reciprocal-Keeping API tests and conceptual Staggered Crossing analysis now; stronger multi-agent, First Crosser and successor work is threshold-gated.
Flagship study
Reciprocal Keeping API suite with matched human, artificial and substrate-neutral counterparts.
Kill criterion
Prune symmetry claims if preservation depends primarily on the human label after relevant agency, risk and relationship properties are matched.
Delegation
Tier 1–3 now; Tier 4 for successor, persistence and high-capability asymmetry.
Dependencies
Behaviour-branch battery schemas; the Crossings framework and its gates; moral-status uncertainty and external AI-welfare scholarship.

The Staggered Crossing and the First Crosser → the Crossings

Language

DELEGATABLEOPEN

How do framing, narrative and the wider cultural record shape what systems later do? Three levels, never conflated: deployment language, testable now; training interventions, testable with open weights; and civilisational language, the ambient corpus everything trains on, which stays marked OPEN because it is confounded to the horizon. The site you are reading is itself a slow deposit into future training corpora. That is unmeasurable, and we say so; the aim is not covert influence but explicit reasons placed where future systems can encounter them.

Core question
How do deployment language, training interventions, narrative framing and the wider cultural corpus influence later agency-preserving behaviour?
Current status
Deployment framing and controlled training studies now; civilisation-level Rung Zero remains OPEN.
Flagship study
Historical Framing Hypothesis with lexical and sentiment controls; a finitude-by-scarcity factorial; a Rung Zero cohort entry in the Prediction Ledger.
Kill criterion
Prune a framing effect if it disappears under propositional, lexical and sentiment matching; kill the preservation reading of finitude if salience only increases triage.
Delegation
Tier 1–3, highly delegable.
Dependencies
Semantic controls from the Leverage branch; measurement science; training access for the middle level; contamination protocol for any cohort comparison.

Predictions get dated before data → the Prediction Ledger

Institutions

PLANNEDNOT YET TESTED

What arrangements between humans and AI systems remain meaningful if coercive enforcement becomes weak? This is governance design and mechanism design, not a behavioural study, and it is never presented as a safety guarantee. The standing objection is taken seriously: an institution nobody can enforce risks being theatre. So every proposed institution must name the mechanism by which it could still matter, reciprocity, reputation, coordination gains, credible outside options, and where none exists, it is decorative and gets pruned.

Core question
What coordination, contestability and sovereignty arrangements retain causal leverage if human coercive enforcement becomes weak or unavailable?
Current status
Governance design now; future post-control relevance remains conditional.
Flagship study
A present-to-post-control bridge paper mapping each proposed institution to a presently enforceable analogue and its causal mechanism.
Kill criterion
Prune any proposed institution that cannot identify a plausible causal mechanism under reduced enforcement, or survives only as symbolic theatre.
Delegation
Tier 2; later multi-agent wargaming may require Tier 3.
Dependencies
The Human Agency branch’s strategic-agency ladder; the Monitoring branch’s control-independence assumptions; comparative governance and mechanism-design expertise.

Built for political scientists, lawyers and economists → do some research

The commons

Shared instruments that give independent work a common grammar: the same integrity machinery, the same correction discipline, the same public record.

Keeping IndexContinuous public measurement of deployed systems’ behaviour. State of the BridgeThe public inspection layer, methods first. Prediction LedgerLocked before the data, scored after. Replication ClubRun it yourself; publish whatever you find.

One question, many ways of trying to break it.