Skip to content
HDX Simulations

The four evaluation levels, and which two are worth your time

Everyone can name the four levels. Far fewer can say what each one costs to collect, which is the only question that decides what you actually measure.

15 September 2026 5 min read

Everyone in corporate learning can name Kirkpatrick's four levels. Reaction, learning, behaviour, results. It is on the second slide of every evaluation deck written since 1959.

Far fewer people can say what each level costs to collect, and that is the only question that decides what gets measured. Levels 1 and 2 are not popular because they are informative. They are popular because they are nearly free.

Level 1 is a room-temperature reading

Reaction is a form at the end of the day. It costs nothing, it comes back immediately, and the scores cluster between 4.2 and 4.7 regardless of what happened. That range is so consistent it functions as a constant rather than a measurement.

This does not make it useless. It makes it useful for exactly one thing: catching a session that went badly wrong. A 2.9 means something. A 4.5 means the room was comfortable, which is not the same as the room being changed.

The trap is that a 4.6 feels like evidence. It gets put in a slide deck next to a spend figure and nobody asks what the relationship between the two is, because the number is high and high numbers end conversations.

Collect it. Spend ninety seconds on it. Do not report it as impact.

Level 2 measures whether they can repeat it back

Learning is a test. Did the content land, can it be recalled, can the model be applied to a written scenario.

For a knowledge intervention — a compliance module, a new system, a product update — this is the right measure and it is close to free. For an experiential session it is close to meaningless, because the thing being learned is not recall. Someone can describe a framework for handling a disagreement perfectly and still not handle one differently.

There is a version of level 2 that is worth collecting from a simulation, and it is not a test. It is the facilitator's record of what a named team actually did during the session: the decision they made without consulting the person holding the relevant information, the round where they stopped talking to each other, the moment they abandoned the plan. That is observed behaviour under pressure, recorded at the time, and it belongs in the file even though it does not look like a level 2 measure.

Level 3 is where the money was

Behaviour is what the sponsor was actually buying. It is also the first level that costs something real to collect, because it requires access to the workplace ninety days after the session and someone willing to look.

The cost is not the measurement. The cost is the setup, and it is front-loaded. You need a named, observable behaviour agreed before the room, a place where that behaviour leaves a trace, and a person who has agreed to look at that trace in ninety days. Arrange those three things and the collection itself is an hour of work.

Arrange none of them and level 3 becomes impossible, which is why most organisations conclude it is impractical. It is not impractical. It is impractical retrospectively, and that is when almost everybody first attempts it.

Level 4 is usually already being collected

Results is the business number. Cycle time, rework rate, escalations, handover defects, attrition on a specific team, days to close.

The common mistake is to invent a new metric for the programme. You almost never need to. The organisation is already tracking something the target behaviour plausibly moves, and using an existing number has two advantages that a new one cannot have: it has history, so you can see the trend before the intervention, and nobody can accuse you of having designed the measure to flatter the result.

The honest limit at level 4 is attribution. A number moved; the session was one of many things that happened. Say so, and say it before anyone asks — which is the subject of the ROI piece in this cluster.

What this means in practice

For a knowledge intervention: levels 1 and 2, cheaply, and stop.

For an experiential session: level 1 as a safety check, the facilitator's behavioural record in place of a test, and levels 3 and 4 designed before the room in a single conversation with the sponsor. That conversation is forty minutes and it is the whole method.

The reason this framework survives sixty-five years of criticism is not that the four levels are the right four. It is that the numbering exposes the gap. Everyone collects 1 and 2 and reports them as though they were 3 and 4, and the framework makes that substitution visible enough to be embarrassing.

More on measuring impact

All of measuring impact
Fernando Mendes, HDX Simulations

Book a call with Fernando

Thirty minutes on whether the catalogue fits your market. No pitch deck — bring the clients you are trying to win.

With
Fernando Mendes
Timezone
GST (Dubai, UTC+4)
Length
30 min
Book a call