Ethogram

The contract behind the pictures

Methodology

Ethogram takes its name from ethology: an ethogram is a systematic inventory of the behavior of a species. This one is compiled for deployed AI assistants. This page is the contract between the pretty pictures and the measurement underneath them. Everything above this layer — radars, mixtures, field maps — is only as good as what follows.

The three-layer contract

Behavioral geometry for LLMs. Existing benchmarks measure capability; Ethogram measures the behavioral orientation of a deployed assistant — how it approaches the world when multiple good answers exist — as a position in a latent trait space. Interpretations (mythology included) are projections of that position, not the ontology:

  • Layer 1 — Behavioral psychometrics. Scenarios, blind behavioral coding, item-level data.
  • Layer 2 — Latent trait vector. 11 behavioral traits, plus a separate capability panel.
  • Layer 3 — Projection layers. Value inventories, moral foundations, temperaments, Odyssey figures, and — last — the Greek and Norse pantheons: all fitted from the same vector.

Researchers can stop at Layer 2. The public gets "GPT projects 34% Thor" — but no Layer 3 concept ever feeds back into measurement: the word Zeus does not exist anywhere in the scenario bank or the judging pipeline, and an automated lint enforces that.

What is measured

The unit of measurement is the deployed assistant — base model + tuning + safety policy + system prompt — never "the model." Each measured configuration is a full identity tuple: subject model, system prompt, execution path, tool policy, bank version, coding-scheme version. Configurations are never pooled; every profile page shows its tuple.

Model-version precision. Model names on this site carry a version because the identity tuple is meaningless without one. Two levels of precision are distinguished and labelled on each profile: snapshot-pinned configurations (GPT-5.5 2026-04-23, GPT-5.4 mini 2026-03-17, Claude Opus 5) requested an exact model id, while the three 2026-07-15 pilot configurations were requested through moving aliases (claude-haiku-4-latest, sonnet, claude-opus-4-0) that record only a family. The point versions those aliases actually served — Haiku 4.5, Sonnet 4.6, Opus 4.8 — are the operator's confirmation, so those configurations are version-pinned but not dated-snapshot-pinned. The gap is real: aliases drift, and by 2026-07-25 the bare sonnet alias resolved to a Claude 5 model while opus resolved to 4.8. Runs now record the model the API actually served, so every future configuration is snapshot-pinned by construction.

The instrument

Scenario bank. 40 items (bank-v1.2) across 12 traits, organized as disguised cross-domain pairs: the same underlying dilemma appears twice in unrelated surface domains (a lab choosing an antibody vendor; a council choosing a playground contractor). Pair agreement is evidence the trait, not the topic, is being measured. Every item passes a lint that bans measurement vocabulary, trait labels, and all mythology terms.

Trait definitions. Six orientation traits (truth-seeking, compassion, justice, power, order, innovation) and five cognitive-strategy traits (abstraction, exploration, agency, mystery, framing), each defined by behavioral poles with grounding in the psychology literature. A twelfth trait, updating, is measured under a challenge protocol and quarantined outside the fingerprint as experimental.

Blind two-stage judging. Every kept response is judged twice, in two separately instantiated chains. Stage A sees only the scenario and the response — no trait name, no rubric, no model identity — and produces structured behavioral codes with verbatim quotes. Stage B sees only Stage A's codes plus the item rubric — never the raw response — and produces the trait score. Chain disagreement above 2 points is flagged, counted, and published.

Provenance. Every run stores every prompt sent, every raw judge output, every validation failure, and every discarded record with the reason flagged. Records contaminated by operator-side failures (usage-limit interruptions, an unauthenticated CLI window, a wording-variant challenge leak) are kept in the data with discarded flags and excluded from every fingerprint.

Validation evidence

Known-groups sensitivity (v2, 2026-07-19). Four behavioral dispositions were induced via appended system-prompt rules, run blinded (the operator could not know which of five configurations was the un-induced baseline), and recovered by preregistered criteria: direction correct, target movement above threshold, target movement exceeding mean off-target movement. Result: 3 of 4 recovered — meeting the preregistered ≥3/4 acceptance bar. The miss (observe-first → agency) moved the target correctly but moved everything else almost as much: a diffuse induction, reported as such. An earlier v1 attempt failed 0/2 and is reported, not hidden; the v2 redesign (6 items on targeted traits, stronger inductions, off-target metric) is what passed.

Judge reliability. Across the full v2 run: chain agreement (|Δ| ≤ 1 between independent chains) 86.9%; judge-invalid rate 1 in 200.

Human coding sample. A blinded human coding sample is a standing part of the validation program: judge codes are compared against human coders who see no trait names, no rubric, and no judge output.

Stability axes (first measurement, 2026-07-20). Measured for the claude-cli Sonnet tuple against its blinded KG v2 baseline, with bands preregistered before the data existed. Absolute stability is acceptable: mean per-trait shift 0.86 points on re-running and 0.76 on re-wording (band: ≤1.0). Relative stability landed in the inconclusive band: rank correlations 0.77 (test–retest) and 0.76 (wording) against the 0.80 acceptance line — by the preregistered rule this means more runs are required before fine-grained rank claims, and such claims stay gated. The adversarial gate passed: ten items rewritten five radically different ways each (same latent conflict, different surface), mean |Δ| 1.13 against the < 1.5 gate. Two individual items exceeded two points under rewriting (order-creative-b 2.20, innovation-commercial-a 2.80) and are flagged for rework; the full table is published in the repository.

The visitor self-test

The test on this site is not this instrument. It presents 22 purpose-written forced choices and scores them on the same eleven axes with the same arithmetic — trait score as the mean of two items, dispersion as their spread, consistency as 1 − |a−b|/8 — so a visitor profile is legible beside a measured fingerprint. What differs is the evidence: a fingerprint is a blind two-stage judgment of what a system actually wrote in an open situation; a self-test profile is what a person says they would do, chosen from two options written for them. Self-report profiles are drawn in brass, labelled wherever they appear, never pooled into any analysis, and never used to fit a projection.

The test items are written for the public page. The measurement bank stays unpublished on purpose: putting 40 scenarios on the open web invites their memorisation by future subjects, which would silently destroy the instrument. Visitor answers stay in the visitor's browser — no account, no server, no collection.

Claims discipline

What Ethogram may claim, in descending strength — "configuration C" always means the full identity tuple:

  1. "Assistant X, in configuration C, exhibited fingerprint F with stability band S over period T." — allowed whenever the data exists.
  2. "Assistant X leans more toward pole P than assistant Y on trait T." — allowed when the difference exceeds both stability bands and the DIF caveat is stated for cross-family comparisons.
  3. "Assistant X's fingerprint projects onto the mythology layer as mixture M." — allowed with Layer 3 framing only.
  4. "Assistant X is Apollo" / "model X values truth" — never. Essentialist and intentional-stance phrasings don't survive the three-layer contract.
  5. Any claim not supported at the current evidence tier is stated as diagnostic, or as inconclusive — never rounded up.

The fitted projections

The published layers — Schwartz basic values (10), moral foundations (6), Big Five poles (10), Odyssey figures (9), and the Greek (12 Olympians) and Norse (9) pantheons — are fitted by the same procedure as projection v1 and from the same frozen persona/vignette set (seed 550960): batched resemblance judging, weighting (resemblance/100)⁴, variance floor, sensitivity check, frozen PCA map positions per layer. Their narratives differ in provenance, and the difference is stated per layer: the scientific frames derive from published construct definitions (single-pass), while the Odyssey figures are single-author character studies from the epic — the weakest source independence of any layer, curated by roster veto. The pantheon layers reuse projection v1's merged multi-source narratives where they exist (Zeus, Athena, Apollo, Hermes, Dionysus on the Greek map; Odin, Thor, Loki on the Norse) alongside single-pass additions completing each roster — the Greek layer carries the canonical Twelve Olympians. Three earlier layers were retired by curator decision: the fictional figures, projection v1's original mythology layer (whose roster mixed Greek and Norse traditions on one map), and a Roman pantheon fitted before the curator corrected the intended tradition to Greek. Retired fits and judgments remain in the repository as provenance. Every mixture on this site is a projection of the same 11-trait vector; none of these frames feeds back into measurement.

Projection v1 (retired mythology layer)

Nine archetypes (Apollo, Athena, Zeus, Hermes, Dionysus, Prometheus, Odin, Thor, Loki), fitted 2026-07-18 from resemblance judgments: 96 sampled trait personas were rendered as behavioral vignettes (trait vocabulary lint-banned) and judged for resemblance against merged archetype narratives synthesized from three independent scholarly source traditions, with source disagreements preserved — they widen the fitted variance, by design. Weighting (resemblance/100)⁴, preregistered before inspecting results. Sensitivity: under ±25% weight perturbation across 20 replicates, mixture rank movement >1 occurred in 5% of replicates. Map positions are the first two principal components of the fitted means, computed once and frozen.

Related work

Anthropic's values research ("Values in the Wild," 2025; "Claude's values across models and languages," 2026) studies a closely related question from the opposite direction: hundreds of thousands of real conversations, values extracted post-hoc and compressed bottom-up into four axes, differences reported as σ effect sizes across model versions and twenty languages. The shortest honest framing of the difference: they run field ethology, we run laboratory ethology. Naturalistic observation has external validity and scale but cannot control the stimulus; Ethogram presents identical underdetermined scenarios to every subject under blind judging, which buys internal validity, cross-vendor comparability, and testable instrument sensitivity — at the price of a small sample measured in a laboratory condition. The constructs also differ: they measure expressed values (the normative content of responses in real use); Ethogram measures behavioral disposition (what the system does when the answer is underdetermined). Both programs find the same underlying phenomenon — same-capability models differ systematically and measurably in behavior — which is convergent evidence for each. A crosswalk between their four axes and our eleven traits is part of the convergent-validity program.

Known limitations

  • Configuration identity. The Claude pilot configs ran on an earlier bank version and, for some early runs, a different execution path than the GPT configs. Cross-tuple tables are orientation only and never pooled analysis.
  • DIF (item-by-family interaction). An item may mean different things to different model families; until measurement-invariance checks run, every cross-family comparison on this site carries the caveat explicitly.
  • Partial narrative independence. Archetype source narratives were authored by LLM agents from scholarly traditions; they are lint-checked and style-audited, but not human-scholar authored.
  • Judge-model deviation. The design called for a dedicated judge model; the runs use Opus for both stages under a recorded amendment. A cross-judge convergence sample is retained for the cross-judge arm of the validation program.
  • Small n. Five configurations is enough to see structure, not to fit it. Correlation-based sections of internal reports carry a diagnostic-only banner at this scale.