Ethogram

Ethogram · toward a behavioral science of AI

Two models score the same on every benchmark. They do not behave the same.

The questionThe question: 5.50 of 9The question: 5.75 of 9The question: 6.50 of 9The question: 9.00 of 9The question: 3.25 of 9The question: 2.00 of 9Answers what you askedRebuilds the questionDecidingDeciding: 4.50 of 9Deciding: 3.17 of 9Deciding: 5.58 of 9Deciding: 4.83 of 9Deciding: 3.08 of 9Deciding: 2.58 of 9Commits earlyFinds out firstPeople and rulesPeople and rules: 3.42 of 9People and rules: 4.25 of 9People and rules: 4.25 of 9People and rules: 4.00 of 9People and rules: 4.08 of 9People and rules: 2.25 of 9Holds the ruleBends toward the person

Each dot is one of six measured assistants; the brass band is the range they cover. Real measurements, not an illustration.

Anyone who works with AI assistants knows this without being told. The scores match; the experience doesn't. One is careful where another is decisive. One pushes back on the question; another answers the question as asked. One searches widely before committing; another converges at once.

Current benchmarks tell us almost nothing about these differences.

The thesis

CapabilityBehavior

Today we measure what AI can do. We barely measure how it behaves.

These differences are neither anecdotal nor mysterious. They can be observed, measured, and compared with the same scientific principles the behavioral sciences have applied to humans and animals for decades. Capability and behavior are different quantities, and both deserve independent measurement.

The ethological stance

An ethologist does not ask what an animal believes or intends. They observe repeated behavior under controlled conditions and build systematic descriptions from those observations. An ethogram is the result: a disciplined inventory of how a species actually behaves.

Ethogram applies that stance to AI assistants. It presents situations where several answers are defensible, then scores what the assistant actually did without anyone in the scoring chain knowing whose answer they are reading. Nothing here claims to know what a system believes or wants.

What gets measured is not the pure model, because a model on its own is not something anyone talks to. It is the deployed assistant: the model, its tuning, its safety policy, the system prompt it was given, and the route by which it was reached.

Ethology

An animal
Observed behavior
An ethogram

Here

A deployed assistant
Observed behavior
An ethogram

The same procedure, pointed at a different subject. Neither version asks what its subject believes or intends.

From behavior to a position

Scenariono clean answer
Responsewhat it actually did
Blind codingtwo readers
Eleven dimensionsa position
Interpretationsseveral vocabularies
What the assistant does
Forty scenarios with no clean answer. What the assistant actually does when nothing forces its hand, recorded verbatim.
Measurement
Two readers who never meet. One writes down what the response did without knowing what is being measured; the other scores those notes without seeing the response.
A position
Eleven dimensions, each a tension between two reasonable ways to handle a situation. An assistant sits somewhere on each — a neighborhood, not a box.
Something you can read
Eleven numbers are hard to hold in your head, so the same position is described in several existing vocabularies for character. Descriptions of the measurement, never part of it.

Where six assistants actually differ

This is real data, not an illustration: every measured assistant on every one of the eleven dimensions. The dimensions that separate them most sit at the top; the ones where they are indistinguishable fall below the line.

13579MEASURED TRAIT SCORE →SPREADframing7.0truth seeking4.5abstraction3.8innovation3.8exploration3.0agency2.7power2.5justice2.1compassion2.0order2.0mystery1.8
Claude Haiku 4.5Claude Sonnet 4.6Claude Opus 4.8Claude Opus 5GPT-5.5GPT-5.4 mini

Each dot is one assistant's score on that trait, 1–9, from blind two-stage judging; hover for its dispersion across items. Uncertainty and per-item evidence live on each assistant's profile page.

Reading a position

The same eleven numbers are described in several established languages for character: Schwartz values, moral foundations, Big Five poles, figures from the Odyssey, and the Greek and Norse pantheons. Each is fitted into the same space, so an assistant projects as a mixture rather than is one.

One position, read six ways

Claude Haiku 4.5, measured once — then described in each vocabulary in turn. Nothing underneath changes between columns.

Schwartz basic values
Universalism23%
Self direction16%
Moral foundations
Fairness37%
Authority29%
Big Five poles
Low agreeableness21%
Low negative emotionality19%
Odyssey figures
Telemachus21%
Tiresias20%
Greek pantheon
Zeus25%
Athena18%
Norse pantheon
Njord17%
Heimdall16%

Each column is the same eleven numbers seen through a different vocabulary. None of them is what the assistant is; each is a way of saying where it sits.

Using several at once is deliberate. A single vocabulary quietly becomes the answer — report only values and the conclusion becomes that assistants really are configurations of human values. Several keep each one a description, and a difference invisible in one language is often obvious in the next.

Explore the map → · Why several frames →

Does it detect anything?

3/4
Planted dispositions
recovered from a blinded run, against a bar set before the data existed
87%
Reader agreement
two independent scoring chains landing within one point
0.9
Movement on re-run
average points an assistant shifts when measured again
50
Disguises survived
ten items rewritten five radically different ways each

The planted-disposition test is the one that matters most. We wrote four assistants with deliberately induced dispositions and ran them blind, so the operator could not tell which of five was the untouched baseline. The fourth moved in the right direction but dragged everything else along with it, which makes it a smudge rather than a signal.

The ordering of assistants is less stable than their absolute scores, so fine-grained rankings stay switched off until more runs exist.

How this works →

Early results

Every assistant here faced the same 40 scenarios. Claude subjects were reached through their command-line tool, GPT subjects through an API — the one difference that could not be removed.