Ethogram · toward a behavioral science of AI
Two models score the same on every benchmark. They do not behave the same.
Each dot is one of six measured assistants; the brass band is the range they cover. Real measurements, not an illustration.
Anyone who works with AI assistants knows this without being told. The scores match; the experience doesn't. One is careful where another is decisive. One pushes back on the question; another answers the question as asked. One searches widely before committing; another converges at once.
Current benchmarks tell us almost nothing about these differences.
The thesis
Today we measure what AI can do. We barely measure how it behaves.
These differences are neither anecdotal nor mysterious. They can be observed, measured, and compared with the same scientific principles the behavioral sciences have applied to humans and animals for decades. Capability and behavior are different quantities, and both deserve independent measurement.
The ethological stance
An ethologist does not ask what an animal believes or intends. They observe repeated behavior under controlled conditions and build systematic descriptions from those observations. An ethogram is the result: a disciplined inventory of how a species actually behaves.
Ethogram applies that stance to AI assistants. It presents situations where several answers are defensible, then scores what the assistant actually did without anyone in the scoring chain knowing whose answer they are reading. Nothing here claims to know what a system believes or wants.
What gets measured is not the pure model, because a model on its own is not something anyone talks to. It is the deployed assistant: the model, its tuning, its safety policy, the system prompt it was given, and the route by which it was reached.
Ethology
Here
The same procedure, pointed at a different subject. Neither version asks what its subject believes or intends.
From behavior to a position
Where six assistants actually differ
This is real data, not an illustration: every measured assistant on every one of the eleven dimensions. The dimensions that separate them most sit at the top; the ones where they are indistinguishable fall below the line.
Each dot is one assistant's score on that trait, 1–9, from blind two-stage judging; hover for its dispersion across items. Uncertainty and per-item evidence live on each assistant's profile page.
Reading a position
The same eleven numbers are described in several established languages for character: Schwartz values, moral foundations, Big Five poles, figures from the Odyssey, and the Greek and Norse pantheons. Each is fitted into the same space, so an assistant projects as a mixture rather than is one.
One position, read six ways
Claude Haiku 4.5, measured once — then described in each vocabulary in turn. Nothing underneath changes between columns.
Each column is the same eleven numbers seen through a different vocabulary. None of them is what the assistant is; each is a way of saying where it sits.
Using several at once is deliberate. A single vocabulary quietly becomes the answer — report only values and the conclusion becomes that assistants really are configurations of human values. Several keep each one a description, and a difference invisible in one language is often obvious in the next.
Does it detect anything?
The planted-disposition test is the one that matters most. We wrote four assistants with deliberately induced dispositions and ran them blind, so the operator could not tell which of five was the untouched baseline. The fourth moved in the right direction but dragged everything else along with it, which makes it a smudge rather than a signal.
The ordering of assistants is less stable than their absolute scores, so fine-grained rankings stay switched off until more runs exist.
Early results
Every assistant here faced the same 40 scenarios. Claude subjects were reached through their command-line tool, GPT subjects through an API — the one difference that could not be removed.