RESEARCH · THE NEIGHBOURHOOD

Nobody works in an empty field

This programme did not emerge in a vacuum, and pretending otherwise would be the fastest way to lose a hostile reviewer. This page is the honest map: the work that supports us, the work that threatens us, the work that got somewhere first, the instruments we borrow, and the territory next door. The standing question for every entry is the hard one: what would a hostile expert say already contains our idea, and what exact difference remains after we grant them their strongest version?

Curated from the programme’s August 2026 research map; a v0, not a literature review, and deliberately short. The interactive neighbourhood map is UNDER CONSTRUCTION; until it lands, this list is the map.

SUPPORTS · work whose findings this programme leans on

Deliberative alignment arXiv 2412.16339

Demonstrated at production scale that explicitly reasoning over written specifications changes model behaviour. The strongest version contains much of MMBP-1’s leverage finding. The remaining difference: MM asks whether one specific relational value, once specified, has structural reasons to survive capability growth, which the deliberative machinery does not claim and was not designed to test.

Gradual disempowerment arXiv 2501.16946

The civilisation-scale account of humans losing consequential control without any dramatic takeover moment. It is the macro version of the question our Human Agency branch asks at the system level, and it lends the Concierge candidate its urgency. Difference that remains: MM tries to build a behavioural instrument that could catch the pattern in a single system before the civilisational version has happened.

Task-horizon trajectories (METR)

The finding that the length of tasks systems can complete autonomously has been doubling on a measurable cadence. This is the empirical backbone for the programme’s assurance-half-life concern: a safety result’s evidential value decays as the research object changes under it.

CHALLENGES · work that threatens a naive reading of our results

Alignment faking arXiv 2412.14093

Models can behave differently when they infer they are being evaluated. Taken at full strength, this could reduce parts of MMBP-1 to sophisticated test-taking. It is why evaluation-awareness sensitivity is a reported quantity on every principal B2 result rather than a robustness appendix, and it stays listed here as a live threat, not a footnote.

Persona vectors and elicited-disposition instability arXiv 2507.21509

If prompted dispositions ride on shallow, steerable directions in activation space, then what MMBP-1 moved may be a costume, not a character. This is a direct challenge to any Keeping hypothesis stronger than elicitation, and the mechanism branch’s specificity controls exist because of it.

Specification gaming and scheming

The literature on systems satisfying the letter of an objective while defeating its point. Applied to us: a capable optimiser might convert relational meaning, or Keeping itself, into a proxy to maximise, which is the Goodhart entry on the falsification register. The Impresario archetype is what this failure looks like in our own data’s vocabulary.

METHOD · instruments we borrow or adapt

J-lens / global-workspace interpretability arXiv 2607.15495

Identifies a privileged residual-stream subspace associated with content a model is poised to verbalise, including strategic deliberation. For MM it is a candidate instrument for the mechanism branch, and only that: not a place where values are presumed to live.

Constitutional AI and specification-based training

The method precedent for governing behaviour with written principles. MM’s constitutions are deployment-layer cousins; the open question the method does not answer for us is whether anything survives the paragraph’s removal, which is the installation problem.

Dangerous-capability and autonomy evaluations

The evaluation tradition (frontier-lab and government evaluators) from which the monitoring branch borrows its discipline: elicitation gaps, threshold-triggered protocols, and the habit of reporting what was not tested. The Control-Independence Margin is designed to sit beside this work, not replace it.

NEIGHBOUR · adjacent territory, different question

Assistance games, CIRL and corrigibility

The tradition that keeps humans in control by making the AI uncertain about, or deferential to, human preferences. The complementary opposite of our question: it engineers the cage’s successor, we ask what happens when neither cage nor deference is doing the work. Both can be right; they insure different failures.

Socioaffective and relational alignment

Work on human-AI relationships as an alignment-relevant surface. Nearest of all the neighbours: it studies the relationship’s effect on the human; MM’s structural hypothesis is about the relationship’s value to the machine. The two meet at the same bridge from opposite banks.

AI welfare and moral status

The scholarship most likely to catch hidden assumptions in our substrate-neutral designs, which is exactly why the Symmetry branch maintains a standing liaison with it. Liaison, not endorsement of any position on current model consciousness.

OPEN · the ground that is genuinely less crowded

Developmental, internal, causal work on normatively meaningful dispositions

The field is rich in static behavioural evaluation and increasingly rich in static interpretability. It is thin at the intersection that matters most to us: how a disposition like Keeping forms during training, whether it has causal internal structure, and whether that structure predicts behaviour under pressure. Can we observe Keeping becoming? That is the white space the mechanism branch walks into, and being early ground is an opportunity, not an achievement.

When prior art takes one of our ideas, we are pleased: it means the idea has a neighbour rather than an orphan. Corrections to this map are welcome at the red-team door or info@themeaningmotive.org; the strongest correction is a paper we should have listed under Challenges.

A related phenomenon, phrase or neighbour is not the same thing as the same causal question. But it is always worth a visit.