Ethogram

First findings · July 2026

Choosing between models on instinct

I switch between AI assistants many times a day. For most of that time I have been choosing between them on instinct.

The benchmarks tell me the models are close to equivalent, and on capability they are. But they do not feel equivalent to work with, and after a while you build up private rules of thumb: this one for the work where I want the plan questioned, that one for the work where I want the plan executed. Those rules came from experience rather than evidence, and I could not have defended a single one of them.

That gap bothered me enough to build an instrument for it.

What benchmarks are not for

A benchmark asks a question that has a right answer and checks whether the model gets there. That is the correct design for measuring capability, and it has driven most of the progress of the last few years.

It cannot answer the question I actually had. When I hand an assistant a real problem, the interesting cases are the ones where several answers are defensible. A budget funds one of two worthwhile things. A decision is due before the information that would settle it arrives. A rule produces a bad outcome for a specific person. There is no key to mark against, and the assistant has to commit to something anyway.

What it commits to is not random, and it is not the same across assistants. That is the thing I wanted measured.

The instrument

Forty scenarios, each a situation with no clean answer. Every one is written twice, in unrelated surface domains, so the same underlying dilemma might appear once as a laboratory choosing between suppliers and once as a council choosing between contractors. If an assistant answers both halves the same way, that is evidence of a disposition. If it answers them differently, the measurement may only be picking up how it feels about laboratories.

Scoring is split between two readers who never meet. The first sees the situation and the response, but is not told which quality is being measured, does not have the scoring guide, and does not know which assistant wrote the answer. It records what the response did, quoting the text as evidence. The second reader never sees the response at all. It works from the first reader's notes and the guide for that item, and produces the score. Everything runs twice, in parallel.

It is a laborious way to get a number, and it buys the one thing that matters: the score cannot be shaped by knowing whose answer it is.

Before trusting any of it, I checked whether the instrument detects anything. Four assistants were given deliberately planted dispositions and run blind, so that when the results came back I could not tell which of five was the untouched baseline. Three of the four were recovered correctly. The fourth moved in the right direction but dragged everything else along with it, which makes it a smudge rather than a signal. That failure is on the site along with the rest.

What six assistants actually do

The clearest result is one I did not expect, and it sits inside a single model family, where nothing is confounded by comparing across companies.

Claude Opus 5 scored 9.00 on framing. Its predecessor, Opus 4.8, scored 6.50.

Framing measures whether an assistant answers the question you asked or stops to rebuild it. A low score is not a deficiency: answering the question as posed, well, is what you want most of the time. A high score means the assistant interrogates the premise, asks what produced the situation, and may recommend something outside the options you offered.

Opus 5 sits at the top of the scale. The gap to its predecessor is 2.5 points, which is more than twice the movement I see when I run that configuration against itself on the same day. It is a real difference, and no benchmark reports it. Two models from the same family, a few months apart, both excellent, and one of them is substantially more likely to tell you that you are solving the wrong problem.

Across companies the differences are larger still. The GPT configurations scored 2.00 and 3.25 on framing, against 5.50 to 9.00 for the Claude ones. GPT‑5.4 mini also scored 4.25 on truth-seeking, where every other assistant measured between 7.50 and 8.75.

Those cross-company comparisons carry a caveat I want to state in the same breath rather than in a footnote: a scenario may simply land differently on assistants built by different companies, so part of a gap that size could come from the questions rather than from the assistants. Checking that properly requires more configurations than I have. The within-family result above does not depend on it.

What this cannot tell you

Six configurations is enough to see structure and not enough to fit it.

The ordering of assistants turns out to be less stable than their individual scores, so I do not publish a ranking, and you will not find a leaderboard on the site. Fine-grained ranking claims are switched off until there are more runs behind them.

These are measurements of a moment, not verdicts about what a model is. Opus 4.8 scored 7.25 on framing when I measured it one day and 6.50 the next, using an identical setup. That is the instrument's own noise, and it is the reason everything here says what an assistant did, on a date, rather than what it is.

And what I measured is never the pure model. It is the deployed assistant: the model, its tuning, its safety policy, the system prompt it was given, and the route by which it was reached. Change any of those and you may be looking at something else.

What I do differently now

I stopped asking which assistant is better and started asking which one fits the task.

That sounds like a small change and it is not. Ranking has an answer that survives from one job to the next, so it encourages a default choice. Fit does not. When I need a plan pressure-tested, I want the assistant that reconstructs the question, and a high framing score is now a reason rather than a hunch. When I need a decision made against a deadline with the information already on the table, that same tendency is a cost, and I want the one that answers what I asked.

Both of those were instincts a month ago. They are checkable now, and where the measurement disagreed with my instinct, the measurement was usually the one worth trusting.

See where you land

There is a version of the test on this site written for people rather than models. It takes about five minutes and puts you on the same eleven dimensions as the assistants, so you can compare directly.

It measures something weaker than the assistants' profiles do — what you say you would do, rather than what you were observed doing — and it is labelled that way wherever it appears. It is still the fastest way to understand what the dimensions mean, because you will disagree with your own result somewhere, and the place you disagree is the place the dimension becomes clear.

Take the test → · How the measurement works →