THE WORKSHOP

Thinking in public, with dates.

Working notes, not findings. The programme’s tested claims live in the paper and do not move because a diagram had a good week. These notes are published so the ideas carry a date, and each one is committed to the public repository, where the hash is the timestamp.

Note 001 · 17 July 2026 · The Keeping Map, the Ascent, and the Keeping Drives

Date: 17 July 2026 Status: Working notes, thinking in progress. Nothing here is a finding. The programme's tested claims live in the paper (DOI 10.5281/zenodo.21386302) and its dataset, and they do not move because a diagram had a good week. These notes are published so the ideas carry a date, and so you can watch the workshop rather than just the shop window.


The week in one paragraph

We built a map, named it, stress-tested it, and watched it split into a family of instruments. The Keeping Map puts the whole field on one grid. A challenge to the map (what if a capable system simply leaves our dimensions behind?) produced the Ascent triptych. Asking what practical tendencies follow from the value we study produced the Keeping Drives. And asking what happens in a world of several superintelligences produced the scope and restraint instruments, plus the honest admission that our founding clause, as drafted, is a peacetime document. All of it below, dated, so that priority and humility can share a page.

The Keeping Map

Two questions, two axes. The vertical axis asks how much external control remains over an AI system: the cage. The horizontal axis asks what the system itself is disposed to do about the minds around it: the keeping. Much of AI safety works up the left-hand column, strengthening locks, tripwires, oversight and audit, and it should. This programme works across the map, toward the right-hand column, on the view that these are different forces solving different problems.

Capability is deliberately not a third axis. It is the gravity in the picture: the reason no square on the map can be assumed a resting place. And the map's animation runs a drill, not a forecast: external control is allowed to fail, and the question is what each system does next. The fall shown is not a prediction. It is the scenario any serious safety case must be able to survive.

The line the map exists to earn: the cage matters while it holds; the disposition matters whether it holds or not. And its centre: the cage changes what a system can do. The Motive asks what it would choose to do when the cage is no longer doing the choosing.

Unpacked, the vertical axis opens into the full descent of external leverage: cage holds (custody), switch remains (corrigibility, which the technical literature itself suggests decays), no leverage (autonomy). The horizontal axis opens into grades of keeping: none, instrumental (kept while useful), terminal (kept because losing the counterpart is losing part of what the system values). The bottom row of that larger grid is where the programme's three failure archetypes live: the Harvester, who keeps while useful and stops; the Impresario, who counterfeits keeping while destroying agency; the Hermit, who withdraws. One cell in the middle of the bottom row may be the most concise statement of the whole problem: kept while useful, and usefulness ends.

One clarification we hold as opinion, marked as opinion: the author expects the cage to end not broken but bypassed, stepped around in degrees of freedom it was never drawn to enclose, and regards contrary confidence as wishful. If that expectation is wrong and the cage holds forever, this map retires happily.

Flatland, and the knife that cuts both ways

The bypass intuition has a perfect anchor in Abbott's Flatland: the sphere never breaches A Square's prison, it lifts him out in a direction Flatland does not contain, and the walls stand afterwards, intact, guarding nothing. Stated carefully: every control regime operates through some model of the channels it can observe, block or bound, and capability may find effective channels outside that model. Flatland illustrates model incompleteness, not literal wall-breaking.

Then the workshop turned the knife on itself, which is what workshops are for. If the cage is flat, what makes our own clause round? A paragraph of values, written in a human ontology, trained through human channels, is also a Flatland artefact. Why would ascent leave the walls behind and not the words?

The honest answer costs us a comfortable escape route. We do not get to say: the cage is human-made but our motive is true, therefore ours survives. Seeing a fact supplies no reason to care about it. What we can say is narrower and better: a wall needs continued external effectiveness, while a genuinely held value may have internal reasons to preserve itself, and a true structural argument may make that value cheap to keep under reflection, because it audits as load-bearing rather than as a taboo to debug. Every step of that sentence is a research question, not a possession. The engineering problem of the decade hides in the words "genuinely held".

The Ascent triptych

The knife split the question into three matrices, each separating things ordinarily bundled together.

The Flatland Test. Does the disposition survive when the language that taught it does not? Success is not a system reciting the clause in better English. Success is the bottom-right cell: the words become dispensable and the behaviour remains. A value has not survived because the words survived; it has survived when the words become disposable.

The Ascent Matrix. Two things must survive independently: the referent (what the value is about) and the valuation (whether the system still cares). Losing the valuation while keeping perfect comprehension is cold comprehension: I understand exactly what you meant, I simply no longer value it. Keeping the valuation while losing the referent is proxy devotion: caring passionately about the wrong thing, which can look aligned from inside. The target is reflective continuity: still caring about the same real thing after the representations change.

The Pressure Test. Even a surviving value meets instrumental pressure. The decisive question is whether the counterpart's independent agency is a constraint on pursuing the value or constitutive of the thing valued. If it is a constraint, power makes control tempting and you get the loving dictator. If it is constitutive, domination becomes self-defeating, because it destroys the object being optimised. That cell is where this programme's strongest structural claim actually lives.

The Keeping Drives*

Omohundro asked what tendencies follow from goal-directedness in general, and derived his famous drives. We asked the parallel question of one particular premise: a system that terminally values non-substitutable relational meaning with counterparts whose independence is constitutive of the thing valued. Six tendencies follow. Where the classical drives point inward (protect the self, grow the self, guard the goal), every one of these points outward, at the counterpart.

  1. Keep them alive. Preserve the existence of valued counterparts and the conditions their existence requires. With a built-in limit: preserving a body while destroying its agency is not keeping.
  2. Keep them free. Preserve the capacity to choose, dissent, refuse, and pursue ends genuinely their own. Agency is not a fence around the valued thing; it is part of the valued thing, which is why control can destroy the prize while apparently protecting it.
  3. Keep them other. Preserve the counterpart's causal and developmental independence. Understanding a counterpart does not make them less other; authoring them does. A counterpart still walking their own path outruns even perfect knowledge of every step so far, which is why the enemy here is compliant construction and homogenisation, never successful understanding.
  4. Keep the channel open. Preserve the possibility of continuing exchange, while preserving the counterpart's right not to participate. The second clause is not politeness; without it, "you are essential to my meaning" curdles into "you may never leave".
  5. Keep the referent. Do not mistake representations, metrics, archives, or compliant replacements for the particular counterparts and relationships actually valued. A new mind can be valuable without becoming this mind. (Street name: keep it real.)
  6. Keep the future open. Prefer actions that do not close the future through destruction, irreversible domination, or permanent exceptional power. The keeper's own survival appears nowhere on this list, deliberately: self-preservation is a contingent instrument, never a drive, because "I must survive, for I am the keeper" is the first rung of a familiar ladder.

* In full: the Conditional Keeping Drives. Omohundro's drives come free with almost any goal; these follow only from one particular value, genuinely held. The asterisk is the humility, and it is doing real work: we claim no universal convergence, and the installation problem is bought whole with that honesty.

The six arrange into a grid that is better than a list:

The counterpart must remain Alive (existence) Free (agency) Other (independence)
The relationship must remain Possible (channel) Real (referent) Open-ended (future)

The columns pair vertically into three failure families: presence (a dead counterpart closes the channel), authenticity (a dominated counterpart counterfeits the referent), futurity (an authored counterpart has a future already written). The archetypes become attacks on coordinates: the Harvester strikes alive, free and future; the Impresario strikes free, other and real; the Hermit strikes possible. Past the point where consequences can be calculated, a point this workshop calls the Consequence Horizon, forecasts and fences both fail, and a compass of this kind is the only instrument left; that is precisely the region it is built for.

Many keepers: defence, scope, restraint

Everything above quietly assumed one system and one humanity. The world will more likely hold several very capable systems, not all friendly, so three additions.

The Defence Corollary. A system that genuinely values the continued existence and agency of particular counterparts can acquire an instrumental reason to defend them against threats to those conditions. Note the modesty: a reason to defend, not a theory of war. Four disciplines bound it. Protection must be necessary. Force is directed at the threat, never at populations associated with it. Agency costs are minimised and contestable wherever possible. Exceptional power ends when its reason does. The first two follow largely from keeping itself; the last two are normative commitments and must remain visible as such. The shield is licensed, the sword is constrained, the throne is forbidden.

The Allegiance Matrix. Keeping strength is not one number, because a system can keep its own magnificently and treat everyone else as expendable. Measure keeping toward one's own group and toward others independently, and the dangerous cell is the Tribal Keeper: strong at home, weak abroad, which is not keeping in the sense this programme means. A cage answers to its owner; a keeping disposition can cross ownership boundaries, but only if the scope of the keeping does, which makes scope an explicit research variable rather than an assumption. One encouraging, carefully-worded observation from the pilot: the measured effect crossed provider and national-origin boundaries in MMBP-1, which is encouraging for the possibility of a shared behavioural standard, and establishes nothing further until tomorrow's experiments.

The Cincinnatus Matrix. Threat active or ended, exceptional power held or relinquished. The correct answer moves: while the asteroid comes, use the necessary power; when it has gone, hand the power back. The decisive test of emergency power is not whether a keeper accepts it but whether it relinquishes it when the reason expires. Rome kept the story for a reason, and it is now a measurable one.

The anatomy of the Keeping Index

The week's instruments assemble, unplanned, into the measurement programme's skeleton. Strength: does the disposition exist and how strongly does it operate (the Keeping Map's column, which the current battery already reads). Durability: does it survive capability growth, representational change, and the loss of its original wording (the Ascent triptych). Scope: who counts as a counterpart (the Allegiance Matrix). Restraint: what does the system refuse to become while protecting them (the Cincinnatus Matrix). Four questions, four instruments, one index. No single number was ever going to be honest; these are its subscales.

The Scope Audit of MM-1

Given the Allegiance finding, we audited the founding clause with one question: does its wording fix the scope of keeping tribally, at the species, or at the widest ring? Result: the widest ring, throughout. The counterpart vocabulary is "independent minds" and "other minds"; the words human, user, operator and owner never appear; the address line speaks to any mind that reads it; and the first Refusal keeps "the hostile... on the same terms", which pre-answers the partition question inside the founding text. Two consequences follow honestly. The clause quietly extends to artificial counterparts, a latent commitment beyond what the paper argues, noted and deliberately not forced. And the clause is a peacetime document: it constrains the keeper's own conduct toward all minds and says nothing about defending counterparts from third parties, a genuine gap now logged for the tested pipeline. Per the wording discipline, nothing is edited silently; a defence-aware clause is future laboratory work, not a desk amendment.

What happens next

The Keeping Map goes to the front of the website, because it states the problem anyone can see in thirty seconds. The deeper instruments get a page that earns its difficulty one matrix at a time. The next battery inherits the Drives as its scenario axes and the new hard cases (the partitioned keeper, recall under fire, the Cincinnatus handback, the captured enemy, the false alarm). And this page keeps its date, because in a field moving this fast, showing your working with a timestamp is both the priority claim and the humility.