AIforBC source-backed guide

How to measure whether AI training actually worked

A practical measurement model for AI workshops, workplace cohorts, and community programs that goes beyond attendance and satisfaction.

Last reviewed August 13, 202611 minute readPrimary sources listed below

Direct answer

Measure four layers: capability, approved adoption, work or learner outcome, and risk behaviour. Establish a baseline before training, observe a realistic task, follow up after people return to their work or daily lives, and include review time, corrections, exceptions, and incidents. Attendance and satisfaction are useful—but they do not prove capability or value.

What matters most

  • Define the behaviour and outcome before the course.
  • Use a realistic task, not a recall quiz alone.
  • Separate capability, adoption, outcome, and risk.
  • Include the time and effort needed to verify and correct AI output.
  • Make expansion decisions from evidence, not enthusiasm.

Start with one observable promise

A training promise should describe what a learner can do in context: complete a bounded task, protect information, verify an important claim, review an output, or design a controlled pilot. ‘Understand AI’ is too broad to evaluate consistently.

Write the task, available information, constraints, output standard, required review, and success measure before delivery. The same task frame can support a baseline, supervised practice, and follow-up.

Measure four different layers

Capability asks whether the learner can perform the method. Adoption asks whether an approved method is repeated after training. Outcome asks whether the work or learner result improved. Risk asks whether use stayed inside the agreed boundaries.

Combining these into one score hides important trade-offs. A team can improve skill but receive no permission to use the tool. Usage can rise while quality falls. Drafting can become faster while review and correction consume the savings.

  • Capability: successful realistic task, correct source behaviour, correct escalation.
  • Adoption: approved users repeating the workflow correctly during the follow-up period.
  • Outcome: full-cycle time, quality, throughput, response, confidence, or task independence.
  • Risk: unapproved information, missed review, material correction, exception, complaint, or incident.

Use baseline, completion, and follow-up evidence

A before-and-after confidence question is useful when paired with behaviour. At baseline, ask the learner to describe or attempt the task. At completion, observe the same capability with a comparable scenario. At follow-up, determine whether the skill transferred to approved real use.

For workplace training, the follow-up period might examine a small number of controlled workflow repetitions and an expand-revise-stop decision. For community learning, follow-up may focus on confidence, safe task completion, verification, and where additional human support is needed.

Protect the integrity of the measurement

Do not let the facilitator mark success only because the course was completed. Define criteria in advance, preserve failures and corrections, and distinguish self-report from observation. Avoid collecting personal or sensitive information that is unnecessary for evaluation.

Small pilots can identify friction and support product decisions, but they should not be presented as representative population research. Report the setting, learner group, sample size, method, missing data, limitations, and curriculum version.

  • Do not compare unlike tasks or ignore changes in review standard.
  • Record the authoritative source and who completed the review.
  • Include accessibility support when interpreting task independence.
  • Separate commercial interest from learning evidence.

Turn evidence into a product decision

At the end of the measurement period, decide whether to expand, revise, stop, or choose another workflow or learning design. Expansion requires both useful outcome evidence and acceptable risk behaviour.

For commercial validation, add partner and buyer signals: paid or sponsored cohort, learner recruitment, attendance, delivery cost, facilitator burden, renewal, referral, and the reason a qualified buyer declined. Those signals should not replace learning quality, but they determine whether the program can operate sustainably.

Continue with a practical resource

Frequently asked questions

Answers before you take the next step

Is a satisfaction survey enough?

No. It measures reaction. Pair it with an observed task, safe-use behaviour, and a follow-up outcome.

What if the workflow becomes faster but needs more correction?

Measure the full cycle. If review and correction remove the time or quality gain, the workflow has not yet improved.

Should every course use the same metric?

Use a common four-layer model, but choose task and outcome measures that match the audience, context, consequence, and program promise.

Can we publish the results?

Only with appropriate privacy, methodology, sample-size, and representativeness disclosures. A small operating pilot should be described as a local pilot.

Source trail

Primary sources used in this guide

  1. Guide on the use of generative artificial intelligenceGovernment of Canada

    Official guidance on accountable use, oversight, training, change management, monitoring, and realistic expectations.

  2. AI Risk Management FrameworkU.S. National Institute of Standards and Technology

    Voluntary framework for governing, mapping, measuring, and managing AI risk.

  3. Workplace artificial intelligence use: A profile of sociodemographic and job characteristicsStatistics Canada

    Official worker-level research that helps distinguish employee use from formal business adoption.

AIforBC uses official and primary sources where practical. This guide provides general operational education, not legal, privacy, cybersecurity, or financial advice.

Need help choosing?

Describe the workflow before choosing the vendor.

Request an AI solution match