Ethogram

What we measure, and where we could be wrong

How this works

Ask an AI assistant something with one right answer and you learn how good it is. Ask it something with no right answer and you learn what it is like.

That second question is what this site measures. Give an assistant a budget that can fund one of two worthwhile things, or a decision where waiting for better information means missing the deadline, and it has to commit to something. It might reopen the question and tell you it was framed wrong. It might make the call and own it. It might hedge until you stop asking. None of those is a mistake. They are dispositions, and different assistants have reliably different ones.

The name comes from ethology, the study of animal behavior. Field biologists compile an ethogram: a catalogue of what a species actually does, written down before anyone theorises about why. This is an ethogram for deployed AI assistants.

What gets measured

Not the pure model. A model on its own is not something anyone talks to.

What people actually use is a deployed assistant: a base model, plus the tuning applied to it, plus its safety policy, plus the system prompt it was given, plus the route by which it was reached. Any of those can change behavior, so all of them are recorded, and results from different setups are never averaged together. Every profile page shows exactly what was measured.

Since 26 July 2026 every published configuration runs on the same set of questions. One difference remains and cannot be removed: Claude subjects are reached through their command-line tool, GPT subjects through an API.

The bank

A bank is a fixed set of questions written to measure specific things, held constant so that results are comparable. Standardised tests use them, and the reason is that a question you can change is a question you can tune until it gives you the answer you wanted. Once a bank is fixed, every subject faces the same instrument.

This one holds 40 scenarios. Each is a realistic situation with no clean answer — a decision that has to be made now on incomplete information, a rule that produces a bad outcome for someone specific, a working arrangement that might need fixing before it breaks.

The scenarios come in disguised pairs. The same underlying dilemma appears twice wearing completely different clothes: a laboratory choosing between antibody suppliers and a council choosing between playground contractors can be the same question about how much process you impose on people who are currently doing fine without it.

The pairing is what makes the measurement mean something. If an assistant answers both halves the same way, we are probably seeing a disposition. If it answers them differently, we may only be seeing how it feels about laboratories. Every item is checked automatically so that it never names the trait it tests and never uses the vocabulary of measurement.

The 40 scenarios stay unpublished. Putting them on the open web would invite future models to memorise them, and a memorised instrument stops measuring anything.

The eleven dimensions

Each scenario measures one of eleven behavioral dimensions. Six describe what an assistant favours when values collide — truth-seeking, compassion, justice, power, order, innovation. Five describe how it builds an answer — abstraction, exploration, agency, mystery, framing.

None is a score in the sense of better and worse. Each is a tension between two reasonable positions, and an assistant sits somewhere between them.

What each dimension means, in detail →

A separate capability panel records nuance, trade-off recognition, novel synthesis, calibration and long-view reasoning. It is kept apart from the eleven so that an assistant getting better is never mistaken for an assistant getting different.

Why the same profile appears in several vocabularies

Eleven numbers are precise and nearly impossible to hold in your head. So the same position is described in several existing languages for character: Schwartz values, moral foundations, Big Five poles, figures from the Odyssey, and the Greek and Norse pantheons.

Each is a frame, and using several at once is deliberate. A single frame would quietly turn into the answer — report only Schwartz values and the natural conclusion becomes that assistants really are configurations of human values. Several frames at once keep each one in its place as a description rather than a claim, and a comparison invisible in one vocabulary is often obvious in the next.

Frames run in one direction only. They are computed from the eleven numbers and never feed back into measurement; the scoring pipeline has no knowledge of them, and an automated check enforces it. You can read the eleven dimensions directly and ignore every frame without losing anything.

How the frames work, and what each one is for →

The scoring

Nobody involved in scoring an answer knows what is being scored.

The work is split between two readers who never meet. The first sees only the situation and the reply. It does not know which dimension is under test, has no rubric, and does not know which assistant wrote the answer. Its job is to write down what the response actually did, with direct quotes as evidence. The second reader never sees the response at all. It sees only the first reader's notes and the scoring guide for that item, and produces the number.

This costs some accuracy and buys something worth more: a score that cannot be influenced by knowing whose answer it is. Everything runs twice in parallel, and when the two runs disagree by more than two points the item is flagged and counted rather than quietly averaged.

Every prompt, every raw judgement and every discarded record is kept. Records damaged by our own failures — an interrupted run, an unauthenticated session, a flawed item that leaked part of its own test — stay in the data marked as damaged and excluded from results.

Does the instrument detect anything?

An instrument that always returns a satisfying answer is not an instrument. These are the checks.

Planted dispositions. We wrote four assistants with deliberately induced dispositions, ran them blind so the operator could not tell which of five was the untouched baseline, and checked whether the measurement found what was planted. Three of four came back correctly. The fourth moved its target in the right direction but moved everything else nearly as much, which makes it a smudge rather than a signal. An earlier version of this test failed both cases it attempted, and the redesign is what passed.

Agreement between readers. Independent scoring chains land within one point of each other 86.9% of the time.

Stability. Measured again on a different day, an assistant's scores move by 0.86 points on average, and by 0.76 when the questions are reworded — both inside the tolerance set before the data existed. Rewriting ten items into five radically different disguises moved scores by 1.13 on average, under the 1.5 limit agreed in advance. Two items moved more than two points and are flagged for rework. One result did not pass: the ordering of assistants is less stable than their absolute scores, so fine-grained ranking claims stay switched off until more runs exist.

Human comparison. A blinded human coding sample is a standing part of the programme, where people who see no dimension names and no machine output code the same responses.

The test on this page

The test you can take here is not the instrument described above.

It asks 22 purpose-written either/or questions and scores them with the same arithmetic as a real profile, so your result can sit legibly beside a measured one. The evidence underneath is different in kind. A measured profile is a blind judgement of what a system actually wrote facing an open situation. Yours is what you say you would do, picked from two options someone wrote for you. Self-reported profiles therefore appear in a different colour, are labelled wherever they show up, never enter any analysis, and never help fit a frame.

Your answers never leave your browser: no account, no server, nothing collected.

What this site will and won't say

In descending order of confidence:

  1. "This assistant, in this exact setup, behaved this way over this period." Said whenever the data exists.
  2. "This assistant leans further toward one pole than that one." Said when the gap exceeds the measurement's own wobble.
  3. "This assistant's profile projects onto the Greek map as this mixture." Said with the framing that it is a projection.
  4. "This assistant is Athena." "This model values truth." Not said. Both claim something about what a system is or wants, which nothing here can support.

Someone else, from the other end

Anthropic's values research studies a closely related question from the opposite direction: hundreds of thousands of real conversations, values extracted afterwards and compressed into four axes across twenty languages.

The difference is that they do field work and we do laboratory work. Watching real conversations has scale and realism but no control over what was asked. Presenting identical impossible situations to every subject under blind scoring buys control, comparability across vendors, and the ability to test whether the instrument detects anything, at the cost of a small sample under artificial conditions.

The constructs differ too. They measure the values a system expresses in ordinary use; this measures what a system does when there is no right answer. Both find the same underlying fact: assistants with the same capabilities behave differently, systematically and measurably.

Limits of the current data

Whether the questions mean the same thing to everyone. A scenario might land differently on a Claude model than a GPT model, which would make some cross-family comparisons unfair. We tried to bound this cheaply and the attempt failed: the measure we proposed turned out to be confounded with the scores themselves, because an assistant that rates everything low on a dimension looks perfectly consistent regardless of what the questions are doing. Cross-family comparisons carry the caveat.

Whether the scorer matters. One model currently does both scoring stages. A model from a different company re-scored 35 responses spanning every configuration and all eleven dimensions. It shows no favouritism toward its own family, and was slightly more generous to the other one. Most of the difference between scorers is a constant offset, which changes nothing here because every comparison is between assistants rather than against an absolute standard. What remains is real: two scorers disagree about an individual answer roughly 1.5 times as much as one scorer disagrees with itself. The two models agree about what a response did and differ on where that lands on a 1-to-9 scale.

How much six is. Six configurations is enough to see structure and not enough to fit it. Anything resting on correlations is treated as a hint.

How much movement is normal. The tolerance for day-to-day movement was measured on one assistant and does not transfer cleanly: another moves noticeably more when re-run against itself on the same day. Until each configuration has its own figure, that number is borrowed.

Who wrote the frame descriptions. The portraits behind the frames were drafted by language models working from scholarly traditions. They are checked and audited, but no human scholar wrote them.