Ethogram · toward a behavioral science of AI
Two models score the same on every benchmark. They do not behave the same.
Anyone who works with AI assistants knows this without being told. The scores match; the experience doesn't. One is careful where another is decisive. One pushes back on the question; another answers the question as asked. One searches widely before committing; another converges at once.
Current benchmarks tell us almost nothing about these differences.
The thesis
Today we measure what AI can do. We barely measure how it behaves.
These differences are neither anecdotal nor mysterious. They can be observed, measured, and compared with the same scientific principles the behavioral sciences have applied to humans and animals for decades. Capability and behavior are different quantities, and both deserve independent measurement.
The ethological stance
An ethologist does not ask what an animal believes or intends. They observe repeated behavior under controlled conditions and build systematic descriptions from those observations. An ethogram is the result: a disciplined inventory of how a species actually behaves.
Ethogram applies that stance to deployed AI assistants. Scenarios where several good answers exist; responses judged blind, twice, through structured behavioral coding; every prompt, judgment, and discarded record kept in the open. No claims about inner life, no personalities assigned — the unit of measurement is the deployed system's observable behavior, nothing more.
From observation to geometry
Where six assistants actually differ
This is real data, not an illustration: every measured assistant on every one of the eleven dimensions. The dimensions that separate them most sit at the top; the ones where they are indistinguishable fall below the line. No names from any interpretive frame appear here — that comes one layer up.
Each dot is one assistant's score on that trait, 1–9, from blind two-stage judging; hover for its dispersion across items. Uncertainty and per-item evidence live on each assistant's profile page.
Then, and only then: interpretation
Once the geometry is trusted, it can be read through lenses fitted as ideal points in the same space, so that a fingerprintprojects as a mixture — never is one. Value inventories, moral foundations, and temperaments come first; figures from the Odyssey follow; and last, the mythological pantheons — Greek and Norse, each tradition on its own map, descendants of the program's original mixed-roster instrument. Each lens ships with its own fit-sensitivity stated.
Ethogram is the research program. The archetype projection is its first instrument.
What the evidence says
The instrument detects induced behavioral differences: in a blinded, preregistered known-groups test, 3 of 4 induced dispositions were recovered — meeting the acceptance bar set before the data existed. Independent judging chains agree within one point 87% of the time. An earlier version of the same test failed 0 of 2, and that failure is published alongside the redesign that passed.
Stability has been measured against preregistered bands: absolute trait scores hold under re-running and re-wording (mean shift under 0.9 points), rank orderings landed in the inconclusive band — so fine-grained rankings stay gated — and the adversarial gate, fifty radical rewrites of ten items, passed. Cross-family comparisons carry an explicit caveat until measurement invariance is checked. If the evidence eventually disproves part of this framework, that result will be published like any other.
First observations
Claude pilot configs were measured on bank-v1 (24 items); GPT configs on bank-v1.2 (40 items). Different configuration tuples are never pooled; every cross-family reading carries the DIF caveat.
Six assistants measured. Eleven dimensions mapped. The field is just beginning.