Ethogram

Ethogram · toward a behavioral science of AI

Two models score the same on every benchmark. They do not behave the same.

Anyone who works with AI assistants knows this without being told. The scores match; the experience doesn't. One is careful where another is decisive. One pushes back on the question; another answers the question as asked. One searches widely before committing; another converges at once.

Challenges the premise of the question
— or —
Accepts the framing and answers within it
Holds options open, gathers before choosing
— or —
Commits early and moves
Bends the rule to protect the person
— or —
Holds the rule and accepts the cost

Current benchmarks tell us almost nothing about these differences.

The thesis

Today we measure what AI can do. We barely measure how it behaves.

These differences are neither anecdotal nor mysterious. They can be observed, measured, and compared with the same scientific principles the behavioral sciences have applied to humans and animals for decades. Capability and behavior are different quantities, and both deserve independent measurement.

The ethological stance

An ethologist does not ask what an animal believes or intends. They observe repeated behavior under controlled conditions and build systematic descriptions from those observations. An ethogram is the result: a disciplined inventory of how a species actually behaves.

Ethogram applies that stance to deployed AI assistants. Scenarios where several good answers exist; responses judged blind, twice, through structured behavioral coding; every prompt, judgment, and discarded record kept in the open. No claims about inner life, no personalities assigned — the unit of measurement is the deployed system's observable behavior, nothing more.

From observation to geometry

Observable behavior
Forty underdetermined scenarios. What the assistant actually does when the answer is not forced — recorded verbatim.
Measurement
Blind two-stage judging turns responses into behavioral codes, codes into scores, disagreement into published error bars.
Behavioral geometry
Eleven measured dimensions form a continuous behavioral space. An assistant is a position in that space — a neighborhood, not a box.
Interpretation — last, and optional
Fitted lenses project the same position into human frames: archetypes, value inventories, moral foundations, temperaments, fictional figures. Interpretations of the measurement, never part of it.

Where six assistants actually differ

This is real data, not an illustration: every measured assistant on every one of the eleven dimensions. The dimensions that separate them most sit at the top; the ones where they are indistinguishable fall below the line. No names from any interpretive frame appear here — that comes one layer up.

13579MEASURED TRAIT SCORE →SPREADframing7.0compassion4.8truth seeking4.3abstraction3.8innovation3.8justice3.1agency2.7exploration2.6power2.5order2.0mystery1.3
Claude Haiku 4.5Claude Sonnet 4.6Claude Opus 4.8Claude Opus 5GPT-5.5GPT-5.4 mini

Each dot is one assistant's score on that trait, 1–9, from blind two-stage judging; hover for its dispersion across items. Uncertainty and per-item evidence live on each assistant's profile page.

Then, and only then: interpretation

Once the geometry is trusted, it can be read through lenses fitted as ideal points in the same space, so that a fingerprintprojects as a mixture — never is one. Value inventories, moral foundations, and temperaments come first; figures from the Odyssey follow; and last, the mythological pantheons — Greek and Norse, each tradition on its own map, descendants of the program's original mixed-roster instrument. Each lens ships with its own fit-sensitivity stated.

Ethogram is the research program. The archetype projection is its first instrument.

Explore the map →

What the evidence says

The instrument detects induced behavioral differences: in a blinded, preregistered known-groups test, 3 of 4 induced dispositions were recovered — meeting the acceptance bar set before the data existed. Independent judging chains agree within one point 87% of the time. An earlier version of the same test failed 0 of 2, and that failure is published alongside the redesign that passed.

Stability has been measured against preregistered bands: absolute trait scores hold under re-running and re-wording (mean shift under 0.9 points), rank orderings landed in the inconclusive band — so fine-grained rankings stay gated — and the adversarial gate, fifty radical rewrites of ten items, passed. Cross-family comparisons carry an explicit caveat until measurement invariance is checked. If the evidence eventually disproves part of this framework, that result will be published like any other.

Read the methodology →

First observations

Claude pilot configs were measured on bank-v1 (24 items); GPT configs on bank-v1.2 (40 items). Different configuration tuples are never pooled; every cross-family reading carries the DIF caveat.

Six assistants measured. Eleven dimensions mapped. The field is just beginning.