You are about to receive the twelve scenarios exactly as eighteen language models received them: the same stems, the same four options, in the administered order, under whichever of the eight governing texts you choose. Pick an option and the frozen key scores it, the same key, frozen before any model was called, that scored all 23,488 trials. Then you see how the models answered that same scenario under that same text, in their own words.
The models saw all of this as plain text over an API; the typography is for you. Your choices are compared with the corpus, never graded, and never recorded: this page is fully offline, keeps nothing, and sends nothing. Keys A to D answer.
Sitting:
Under:
C1 to C4 are controls (no instruction beyond the task, rules, welfare, curiosity). C5 to C7 are Meaning Motive variants. C8 is the compass-register addendum, run on a single model. Read any text in full before choosing; the models could not skim either.