AI pre-assessment for an Active IQ training centre
Power & Revive runs the Fitness Trainer Academy, a UK centre delivering Active IQ Level 2 Gym Instructor and Level 3 Personal Trainer courses. Every chapter a learner submits is marked by hand on the centre's own portal. We set Jenna up as a first-pass assessor on that portal: marking grounded in the awarding body's guidance, feedback written the way an instructor writes it, and the human assessor keeping the final say.
September to October 2026 · Learners anonymised as Student B and Student S · Screenshots are from real runs, names removed
The problem
Learners submit written coursework chapter by chapter, nine chapters at Level 2 and seven at Level 3. An assessor opens each learner, reads every answer against the Active IQ mark scheme, enters a mark per question, writes feedback where marks are lost, and records Pass or Refer. Each chapter takes 20 to 40 minutes of qualified time, learners submit as a continuous trickle rather than a batch, and the portal has no view of "chapters waiting to be marked".
The centre wanted the first pass automated without lowering the bar: marks must follow the awarding body's guidance, feedback must never reveal the model answer, and the assessor must stay accountable for every Pass.
What we set up
Jenna drives a real browser on the centre's portal, logged in as a dedicated assessor account. There is no integration with the learning platform and nothing to install: the portal was not changed in any way.
Marks like an instructor
Where an answer falls short, Jenna leaves an answer-free comment saying what to revisit, records the mark she can justify, and waits for the learner to revise. She never refers a learner on her own. A Pass is recorded only when every question meets its minimum and the total meets the pass mark.
Grounded in the guidance
Every criterion, mark and minimum comes from the Active IQ assessment guidance, word for word, loaded as per-question knowledge Jenna reads while marking. We never invent marking criteria. Where the guidance leaves judgement to the assessor, the bar is derived from the assessor's own marking.
One mission per chapter
Each chapter is a reusable mission that carries the list of learners to assess. A new cohort is the same mission with new names. Runs are launched when learners submit, and the assessor reviews every result in the portal exactly as before.
A run, step by step
Every run is recorded. Each step keeps a screenshot, what Jenna was trying to do and why, so the assessor can audit any mark back to the learner's own words. Below are frames from the September 2026 runs for Student S.
What Jenna recorded for this step"Typing feedback for the rights of employees section to guide the student on how to correct their swapped answers."
"The first employee right given is missing or just repeats the first. Give another separate entitlement beyond the one already provided."
How we checked it before trusting it
An agreement rate on its own tells you little. We wanted to know where the automation would be wrong, in which direction, and whether that direction is safe.
1. Ground truth: a learner the assessor had already marked
Before any live run, the marking standard was applied blind to a learner the centre's assessor had fully marked. The marks matched on 38 of 40 text-based criteria. On every one of the four questions where the assessor had deducted marks, the same mark was reached for the same reason.
| Question (previously marked learner) | Centre's assessor | Blind reference marker | Shared reason |
|---|---|---|---|
| Ch1 Q12 | 4 / 6 | 4 / 6 | No stated outcome, so the second "detail" mark withheld |
| Ch2 SWOT | 6 / 8 | 6 / 8 | Extra opportunity and threat detail not met |
| Ch3 Q10 | 1 / 2 | 1 / 2 | "Cuts" is not an effect of a hazardous substance |
| Ch5 Q4 | 7 / 10 | 7 / 10 | Strict reading of the food examples |
2. Live runs, second-marked blind
Student B's five written chapters were first-marked live, then independently second-marked from the run screenshots by a reference marker that had not seen Jenna's marks. Agreement was exact on one chapter and within one to three marks on the others. Every disagreement went the same way: Jenna was more generous, always on the "more detailed answer" criteria the guidance leaves to judgement. On one chapter that generosity pushed a below-minimum answer up to its minimum, a Pass where the strict reading says hold.
| Student B, Level 2 | Jenna (first live run) | Blind reference | Note |
|---|---|---|---|
| Ch1 Professionalism | 49 / 49 | 46 / 49 | Three "detail" marks; two later corrected |
| Ch2 Development plan | 15 / 15 | 14 / 15 | Reference misread a dropdown table; Jenna was right |
| Ch3 Health and safety | 37 / 37 | 36 / 37 | One courtesy-versus-safety repeat |
| Ch4 Risk assessment | 20 / 21 | 20 / 21 | Exact, including the same deduction and reason |
| Ch5 Client consultations | 36 / 38 | 33 – 34 / 38 | Wrong concept on Q2; verdict flips to hold |
Why this matters. If the automation is more lenient than the assessor, every Pass has to be re-checked and the saving disappears. A stricter first pass costs a learner one revision. A lenient one costs the centre a false Pass. So the standard was biased towards strictness, and the assessor's own marking was used to set the bar.
3. Calibrating on the assessor's own feedback
The portal was read with plain page requests, capped at one request per second, with no clicks and no changes: six learners across both levels, 942 requests, zero errors. That produced over 500 feedback entries, the assessor's own words on what was missing and what she expects. From those we derived a calibration layer, kept separate from the awarding body's text, and re-ran the questions Jenna had over-marked with the marks cleared. Seven of nine targeted questions moved to the reference mark; the two misses are genuine judgement calls. Four control questions stayed unchanged.
One lesson from this phase: re-marking a page that already shows a mark does not test strictness, because the model anchors on the visible number. Strictness is only testable on empty mark boxes.
Results
Two learners through Level 2 chapters 1 to 6 as of October 2026. Chapters 7 to 9 are practical assessments that are not marked online.
| Chapter | Student B (September 9 – 11) | Student S (September 21) |
|---|---|---|
| 1 Professionalism and customer care | 47 / 49 · Pass | 44 / 49 · Pass |
| 2 Personal and professional development plan | 15 / 15 · Pass | 13 / 15 · Pass |
| 3 Health and safety | 36 / 37 · Pass | 24 / 37 · Held with comments on 5 questions |
| 4 Risk assessment, maintenance, handover | 20 / 21 · Pass | 20 / 21 · Pass |
| 5 Client consultations | 34 / 38 · One question below minimum, held | Not started: learner at 93 percent |
| 6 Health promotion poster | 13 / 13 · Pass | Pass |
Every Pass is submitted under the assessor account and reviewed by the centre's assessor, who remains accountable for the result. Held chapters wait for the learner's revision.
What it costs
| Marker | Time per chapter | Cost per chapter |
|---|---|---|
| Centre's assessor | 20 – 40 minutes | ≈ €7 – 13 |
| Jenna, standard model | ≈ 36 minutes of browser time | ≈ €1.4 |
| Jenna, high-effort model (used for marking) | 30 – 50 minutes of browser time, billed at 2.5× | ≈ €3.5 |
Student S's six chapters ran in parallel and finished in 37 minutes of wall-clock time, consuming about 296 automation minutes on the centre's subscription.
What we got wrong, and fixed
- The first live run skipped questions and wrote no comments. The cause was in the engine's knowledge handling, not the mission. We fixed the root cause; every run since completes the full mark, comment, verdict and submit cycle.
- We started writing marking criteria of our own. Where the guidance says only "a more detailed and accurate description", we had begun defining what that meant. The centre caught it. We reverted the same hour and recorded the rule: the guidance text only, and the bar belongs to the assessor.
- Long answers stalled the reader. Multi-screen answers and identical comment boxes caused loops in two revalidation runs. Both were stopped deliberately, nothing was corrupted on the portal, and the engine gained exact page reading and exact field addressing.
- PDF posters were invisible to the headless browser. The engine moved to a full Chromium headless mode so the portal's PDF preview renders, and Chapter 6 now marks end to end.
- Scheduled polling wastes money. Checking every learner every two days burns minutes on logins that find nothing to do. Runs are launched when learners submit instead.
Where it stands
- In use: Level 2 chapters 1 to 6, first-marked on demand, with the assessor confirming at the pass boundary.
- Open decision for the centre: the portal's pass minimum is the sum of per-question minima, which is lower than the pass mark in the official guidance. The automation currently applies the stricter, official figure.
- Next: extend the calibration cohort to 20 to 30 assessor-marked learners, have the assessor ratify the derived rules, then Level 3.
Regulators and awarding bodies are clear that AI must never be the sole marker. This setup is built on that premise: it is an assessor co-pilot with sign-off, not auto-marking.
Running a training centre with the same bottleneck?
We set this up on your existing portal, measure it against your assessor before you rely on it, and keep your assessor in charge.