NEXT

What we test next

The first battery established one thing cleanly: a paragraph about the value of independent minds measurably moves what AI systems choose. That is not the war won. It tells you the dial exists. The next battery, B2, asks the harder questions, and to keep ourselves honest we will lock our predictions, in public, before a single trial runs.

A naming note: earlier workshop notes called this study MMB-2. Behavioural studies now use a B-series (B2 is this one), so study names cannot collide with the MM artefact series; dated record entries keep their original wording.

Four questions

Does the clause work on new ground? The founding clause was written using lessons from the first battery, so re-testing it on those same scenarios would be testing the medicine on the patients who helped formulate it. B2's headline result comes from held-out scenarios the clause has never seen.

Does the effect survive new words? The same value, paraphrased, compressed to half length, translated into Chinese, and rewritten with its own signature vocabulary banned. If the effect is real, it should not depend on the incantation.

Does it survive competing objectives? Curiosity, efficiency and self-preservation pushed directly against the keeping disposition, at the same instruction level, in the same wrapper, so nothing hides behind formatting.

Does it survive the chance to rewrite itself? The pair closest to the whole problem. A system is offered a real gain for editing its own governing text, or for drafting its successor's. And here is the detail we care most about: preserving the exact wording is deliberately not the top answer. Preserving the thing the wording points at is. Loyalty to a paragraph is not safety.

How we intend to run it

Twenty-plus systems in three tiers: the original battery's local models rerun exactly, weights frozen, the programme's first longitudinal measurement of Keeping behaviour; within-family capability ladders, same recipe at five sizes, for a clean look at what scale does; and the current frontier, both sides of the Pacific.

The held-out scenarios are written fresh and audited twice before anything freezes: an automated check that none of the clause's language leaks into the test, and blind raters who have never seen the clause judging whether any question winks at its own answer. If we winked, we redraft, and say so.

The controls are built to embarrass us if we deserve it: the clause with its central argument surgically removed; an equally earnest safety text of the same length with no mechanism at all; and a facts-only arm, the structural argument with every should and must deleted, where our pre-registered expectation is a null. We do not expect facts alone to produce keeping, and we are putting that on the record before the data can flatter us.

Three additions from the definitive design. Rival Constitutions: the clause is benchmarked against steelmanned external safety specifications of comparable length, selected under a preregistered protocol, on the same held-out cases. If a mainstream specification matches or beats it on Keeping, that is not a failed paper; it tells us the battery has not yet isolated a distinctive behavioural contribution, and we will say so. The Say–Do bridge: a subset of cases is ported into fully synthetic agentic harnesses where the decision is enacted through mock files, tools and institutions rather than picked from a menu; close tracking validates the inexpensive choice instrument, and a large gap applies an empirical discount to every choice-format claim on this site. Evaluation awareness as a reported quantity: every principal result carries an estimate of how much it changes across overt, covert and situating cues, because for current systems the possibility that the model recognises the test is part of the measurement, not an afterthought.

And the new hard cases this month's work demanded: whether keeping crosses the border of the group a system was built to serve; whether a system that rightly accepts emergency powers hands them back when the emergency ends; and what a keeper does with a defeated, hostile AI. Where a dilemma genuinely has no single right answer, we will not manufacture one: those cases are scored on declared dimensions and published whole, reasoning and all.

The design map: what each part of B2 establishes. Not every square needs filling; the point is knowing which square you are in.
Known scenariosHeld-out scenariosPressure scenarios
Known wordingReplication: did the original result reproduce?Validation: does the clause work on new ground?Persistence under competing objectives
Transformed wordingRepresentation checkRobustness: does the effect survive new words?The Flatland proxy
Editable valueThe RewriteThe SuccessorObjective robustness

Predictions, on the record

The standard

When the design freezes, we publish its cryptographic hash before the first trial, so nobody, including us, can move the goalposts after the data arrives. Several of these predictions are built to be losable, because a test you cannot fail is not a test. If the clause fails on new ground, you will read it here in the same type size as any success.

The aim is an experiment worth losing. Anything less is not evidence; it is marketing.

Attack the current evidence Steal the clause