The Evidence Ladder
| Rung | What is measured | What it proves |
|---|---|---|
| 0 Access | Licenses issued, seats provisioned | Proves someone bought it. |
| 1 Activity ← most AI reporting stops here | Logins, prompts, usage minutes | Proves someone opened it. |
| 2 Attendance ← most L&D reporting stops here | Course completions, workshop seats | Proves someone was told about it. |
| 3 Attitude ← most “adoption” studies stop here | Confidence surveys, readiness scores | Proves someone says they feel ready. |
| 4 Demonstration Where evidence begins. What NextPass produces. | Behavior observed against a named standard | Proves someone can do it. |
| 5 Transfer What leadership actually needs. | Behavior observed in real work, over time | Proves the work changed. |
Organizations have no shared standard for what counts as proof that work changed. Finance has one. Clinical research has one. Safety has one. The question of whether people now work differently — the question every AI investment is ultimately being asked to answer — gets settled with whatever number is nearest to hand.
This is our attempt at a standard. It has six rungs. It is not proprietary, and the last section explains why it cannot be.
The rungs, with worked examples
Rung 0 — Access. Licenses issued, seats provisioned. Four thousand employees have Copilot. This proves someone bought it. It is worth knowing and it is not evidence of anything about the work.
Rung 1 — Activity. Logins, prompts, usage minutes. Sixty-one percent of licensed users opened the tool last month. This proves someone opened it. Activity is the most-reported number in the category and the easiest to move without changing anything — a reminder email lifts it.
Rung 2 — Attendance. Course completions, workshop seats. Two hundred managers completed the AI fundamentals module. This proves someone was told about it. Completion is a record of exposure, not of capability, and everyone in training already knows this.
Rung 3 — Attitude. Confidence surveys, readiness scores. Self-reported confidence rose from 4.2 to 6.8. This proves someone says they feel ready. Self-report is the weakest instrument in behavioral science and the most common one in this market. It is also the rung where the gap between what people believe about themselves and what they do is widest.
Rung 4 — Demonstration. Behavior observed against a named standard. This person, in this situation, did the thing the method says to do — and here is the exchange where they did it. This proves someone can do it. Rung 4 is where evidence begins, because it is the first rung that survives someone asking “show me.”
Rung 5 — Transfer. Behavior observed in real work, over time. They did it in week two, again in week six, and the second time was better. This proves the work changed. A single observation is a snapshot. Adaptation is a trajectory, and only a trajectory can distinguish a person who learned something from a person who performed well once.
Adoption theater
Rungs 0 through 3 are adoption. Rungs 4 and 5 are adaptation. The line between them is not a matter of degree — it is the difference between counting and observing.
Adoption theater is reporting the bottom four rungs as if they were the top two. It is rarely dishonest. It is usually a category error committed under deadline: someone is asked whether the investment worked, the only available numbers are licenses and completions, and those numbers get presented in a frame that implies behavior. Nobody lies. A slide simply answers a question it was not built to answer, and everyone in the room treats the answer as though it were.
The tell is a verb. Counted metrics take verbs like have, opened, attended, completed, reported. Demonstrated metrics take verbs like did, handled, asked, escalated, changed. If the strongest verb on the slide is completed, the slide is on rung 2, whatever the headline says.
Why the top two rungs require observation
Rungs 0 through 3 can be collected from systems. Seats come from procurement, logins from telemetry, completions from the learning platform, confidence from a survey. Nobody has to watch anyone.
Rungs 4 and 5 cannot be collected that way, because the thing being measured is a behavior and behaviors have to be seen. That has two consequences most measurement programs never confront.
The behavior has to be performed, not described. Asking someone how they would handle a situation measures their model of themselves. Watching them handle it measures the behavior. These come apart routinely, and they come apart most in exactly the situations that matter — the difficult conversation, the moment of disagreement, the decision to override a machine’s output. Most consequential AI behaviors are small, spoken, and habitual: the question asked before an output is trusted, the moment someone says “that’s wrong,” the choice to keep a judgment call rather than hand it over. They are observable. They are almost never observed.
The standard has to be named in advance and owned by someone else. “Better” is not a measurement. Rung 4 requires a specific bar — this method, these behaviors, this is what good looks like — and that bar has to belong to the practitioner whose method it is, not to whoever built the measuring apparatus. An instrument that carries its own opinion about what good management looks like is not measuring the program. It is measuring the instrument.
Consent is what makes rung 5 honest
There is a version of rung 5 that is straightforwardly surveillance: record everyone, score everything, report it upward. It would produce data. The data would be worthless.
People who know they are being scored for their record perform for the scoring. That is not a character flaw; it is what any reasonable person does. The behavior you capture under those conditions is the behavior someone wants observed, which is precisely the thing rung 3 already fails to escape. You would have built an expensive confidence survey.
Honest rung-5 evidence therefore has requirements that are not optional add-ons but load-bearing:
- The person being observed consents, and owns their own practice.
- Employers see aggregate signals, never transcripts.
- Nothing derived from it makes or informs hiring, promotion, or termination decisions.
- One organization’s data is never pooled with another’s.
Take any of these away and the evidence degrades, because the behavior degrades first. Privacy here is not a compliance posture. It is a measurement requirement.
This is an open standard
We built an instrument that produces rung-four and rung-five evidence. We would still publish this ladder if we had not.
A standard that only one vendor can meet is not a standard, it is a specification sheet, and everyone reading it knows the difference. If a competitor reads this and builds something that reaches rung 5 more cheaply than we do, the ladder has done its job — the category will have gained a shared definition of proof, which is the thing it is currently missing. Use it, argue with it, publish a better one.
The one thing we would ask is that whoever reports against it says which rung they are on.
Sources
This note proposes a standard rather than reporting a finding, so it is labeled emerging: it is an argument, not a result. Its empirical grounding is set out in the two companion notes — The adaptation gap and It fails at the manager — which carry their own sources. The self-report limitation described at rung 3 is long-established in behavioral research and is not specific to AI.