Measuring behaviour change, not satisfaction
Level 1 tells you whether they enjoyed it. Levels 3 and 4 are the ones your sponsor is paying for, and they are measurable if you set them up before the session rather than after.
9 September 2026 — 5 min read
Almost every corporate training session is evaluated the same way: a form at the end asking how useful it was, out of five. The scores come back between 4.2 and 4.7, they always do, and they tell the sponsor nothing they could act on.
That is Kirkpatrick level 1 — reaction. It measures whether people enjoyed the day. Level 2 measures whether they can recall the content. Neither is what the budget was approved for. The sponsor approved it because they wanted something different to happen at work, which is level 3, and because that difference was meant to show up in a business number, which is level 4.
Levels 3 and 4 have a reputation for being impractical. They are not. They are impractical after the session, which is when most people first think about them.
Everything is decided before the room
If you want to say anything defensible about behaviour change, you need a description of the current behaviour written down before the intervention. Afterwards, you are asking people to remember how they used to be, and they cannot — the session itself has changed what they recall.
The setup conversation is with the sponsor and it takes about forty minutes. Three questions.
What will people do differently? Not "improve collaboration". Something observable, that a third party could confirm happened. "Cross-team requests get a yes or no within two working days instead of going quiet." "Project leads state the decision owner at the start of a kickoff." If the sponsor cannot name one, the session does not yet have a purpose and you should say so.
How would we know? Where does that behaviour leave a trace? A ticket system, a calendar, a meeting note, a handover doc, a monthly report. If the behaviour leaves no trace anywhere, you will be relying on observation, which is fine, but you need to arrange who is observing before the session rather than after.
What business number does it touch? This is level 4, and it is nearly always a number the organisation already tracks. Cycle time, rework rate, escalations, handover defects, attrition in a specific team. You do not need a new metric. You need an existing one that the behaviour plausibly moves.
Write those three answers down and send them to the sponsor. That document is the evaluation design, and it takes forty minutes because you asked the questions in the right order, not because the method is complicated.
What the simulation gives you that a course does not
An experiential session has one enormous evaluation advantage: you observe the behaviour happening during the session, under pressure, in front of you.
That is a genuine baseline. Not a self-report, not a survey — a record of what a named team actually did when a decision was time-boxed and the information was incomplete. If your facilitator notes say that team three made four decisions in round two without consulting the person holding the relevant role, you have a concrete, specific starting point that no classroom day can produce.
Run the same observation ninety days later against real work and you have a before and after where both halves describe behaviour rather than opinion.
The ninety-day check
Level 3 needs elapsed time. Anything measured in the week after the session is measuring enthusiasm, which decays.
Ninety days is the working standard: long enough that novelty has worn off, short enough that people still connect the change to the session. The check itself should be small — three questions to the participant's manager, not a survey to the participant.
- Have you seen [the specific behaviour] since [month]?
- Can you describe one occasion?
- Is it happening more, less, or about the same as before?
Twelve managers, five minutes each. That is an hour of someone's time and it produces a level 3 finding that will survive a challenge in a budget meeting, because it rests on described occasions rather than on ratings.
The "describe one occasion" question is the one that does the work. A manager who says yes but cannot produce an example has told you something important, and you should record it as a no.
Being honest about attribution
Level 4 is where evaluation gets oversold. A team's escalation rate falls 19% in the quarter after a simulation. Did the session cause that?
Probably not on its own. There was also a reorganisation, a new team lead, and the summer. Claiming the 19% is dishonest and, worse, it is fragile — the first person who points at the reorganisation has destroyed the credibility of everything else you reported.
The defensible version states the number, states the confounders, and lets the reader weigh it:
Escalations from the three participating teams fell from 41 to 33 per month across the quarter following the session, against a flat trend in four comparable non-participating teams. A team lead changed in one of the three during the same period.
That is a weaker claim and a far stronger document. Sponsors are not naive about attribution; they deal with it constantly. What damages your credibility is not a modest result, it is an overclaimed one.
Where you can, find a comparison group. Organisations almost always run these sessions in waves, which means for a few months there is a population that has not been through it yet. That is a natural control and it costs nothing to use. Ask for the same number for both groups.
What to hand back
The report the sponsor actually reads is about two pages.
- What we said would change, in the sponsor's own words from the setup conversation
- What was observed in the room, with specifics from the facilitator notes
- What the ninety-day check found, including the occasions managers described
- The business number, with its confounders stated plainly
- What did not change, and the most likely reason
Point 5 is not a confession. It is the part that makes the other four believable, and it is usually the part that determines whether there is a second session.
Every HDX simulation ships with client-ready reporting built around this structure, so a partner has something to hand the sponsor that is neither a satisfaction score nor a marketing document. See the catalogue, or talk to us about licensing.

