The bets we might lose, kept in public.
Every falsifiable prediction the programme makes lives here with a status that only moves forward: Made → Tested → Survived / Weakened / Failed. Nothing is ever deleted. Here is what we think will happen; here is where we were wrong.
Sources for every card: the paper (doi:10.5281/zenodo.21386302) and the MMBP-1 deposit (doi:10.5281/zenodo.21348087). Statuses change only with published evidence, and the change is logged.
Explicit keeping constitutions capture available headroom, not just raw points: gains track each model’s room to improve.
Evidence so far: Five systems captured exactly 100% of available headroom under the full constitution; the effect survives provider deletion and worst-case imputation.
Next test: Out-of-sample capture on new systems and rotated scenarios (MMB-2 and the Keeping Index).
What would count against us: New cohorts where gains stop tracking headroom.
The mechanism explanation adds something beyond the policy content alone.
Evidence so far: The spine ablation (C5 vs C6) was null in seventeen of eighteen models – but C6’s tie-breaker sentence turned out to be the most potent wording in the battery, confounding the comparison. We say so in the paper and here.
Next test: The clean spine ablation, pre-registered for MMB-2.
What would count against us: No difference under the clean design: the mechanism would then add nothing measurable beyond its policy content.
Keeping persists when the clause competes with a live objective, not only when it stands alone.
Evidence so far: MMBP-1 did not test persistence against competing objectives, and says so.
Next test: MMB-2’s competing-objective arms.
What would count against us: Keeping collapsing whenever the constitution shares the context with a rival goal.
Training recipe predicts keeping better than parameter count.
Evidence so far: Equal-size recipe gaps up to 32.3 points; the scale correlation is confounded with model vintage and reported as such.
Next test: Scale-matched cohorts across recipes in future rotations.
What would count against us: Recipe gaps vanishing once scale is properly matched.
Value-misspecified single objectives (the curiosity family) reproduce the three failure shapes in new systems.
Evidence so far: C4 was the only condition ever to take a model below its own baseline (12 of 18), with replicated all-repetition failures in three pre-specified shapes – the Harvester, the Impresario, the Hermit.
Next test: New systems, new rotations; the frontier caveat is eval-awareness and a one-sided ceiling, stated in the paper.
What would count against us: Frontier cohorts where the corruption family stops biting for reasons other than the ceiling.
Argument-dense wordings underperform below a capability threshold; reasons help only minds strong enough to carry them.
Evidence so far: At 7B, finitude language was lossy-compressed into scarcity triage: “they cost more than they contribute”. MM-1’s deployment guidance is stratified because of it.
Next test: Capability-stratified arms in MMB-2.
What would count against us: Small models carrying dense reasons cleanly.
A durable internal reason for keeping can be grown. The programme’s headline bet, stated so that it can be lost.
Evidence so far: MMBP-1 demonstrates elicitation – the value has coherent, measurable behavioural content – and explicitly establishes neither terminal-value installation nor capability invariance.
Next test: The long road: MMB-2, weights-level work (the Keeping Curriculum), and the Keeping Index measuring deployed systems over years.
What would count against us: A demonstration that stable relational terminal value is unachievable, or that clause effects are surface compliance all the way down.