Jenna FREN Book a call
Case study · Assessment automation

AI pre-assessment for an Active IQ training centre

Power & Revive runs the Fitness Trainer Academy, a UK centre delivering Active IQ Level 2 Gym Instructor and Level 3 Personal Trainer courses. Every chapter a learner submits is marked by hand on the centre's own portal. We set Jenna up as a first-pass assessor on that portal: marking grounded in the awarding body's guidance, feedback written the way an instructor writes it, and the human assessor keeping the final say.

September to October 2026 · Learners anonymised as Student B and Student S · Screenshots are from real runs, names removed

12chapter assessments first-marked for two learners
38 / 40criteria where the marking standard matched the centre's assessor
€1.4 – 3.5per chapter, against €7 – 13 of assessor time
37 minwall-clock to first-mark six chapters in parallel

The problem

Learners submit written coursework chapter by chapter, nine chapters at Level 2 and seven at Level 3. An assessor opens each learner, reads every answer against the Active IQ mark scheme, enters a mark per question, writes feedback where marks are lost, and records Pass or Refer. Each chapter takes 20 to 40 minutes of qualified time, learners submit as a continuous trickle rather than a batch, and the portal has no view of "chapters waiting to be marked".

The centre wanted the first pass automated without lowering the bar: marks must follow the awarding body's guidance, feedback must never reveal the model answer, and the assessor must stay accountable for every Pass.

What we set up

Jenna drives a real browser on the centre's portal, logged in as a dedicated assessor account. There is no integration with the learning platform and nothing to install: the portal was not changed in any way.

Behaviour

Marks like an instructor

Where an answer falls short, Jenna leaves an answer-free comment saying what to revisit, records the mark she can justify, and waits for the learner to revise. She never refers a learner on her own. A Pass is recorded only when every question meets its minimum and the total meets the pass mark.

Standard

Grounded in the guidance

Every criterion, mark and minimum comes from the Active IQ assessment guidance, word for word, loaded as per-question knowledge Jenna reads while marking. We never invent marking criteria. Where the guidance leaves judgement to the assessor, the bar is derived from the assessor's own marking.

Operation

One mission per chapter

Each chapter is a reusable mission that carries the list of learners to assess. A new cohort is the same mission with new names. Runs are launched when learners submit, and the assessor reviews every result in the portal exactly as before.

The portal's chapter list for a learner, showing student progress, assessor progress and status per chapter
Jenna reads the learner's chapter list to decide what is ready: a chapter is assessed only when the learner has completed it and no assessor has marked it yet. Chapter 5 here is at 93 percent, so it is left alone.

A run, step by step

Every run is recorded. Each step keeps a screenshot, what Jenna was trying to do and why, so the assessor can audit any mark back to the learner's own words. Below are frames from the September 2026 runs for Student S.

A Chapter 3 question with the learner's three answers and an assessor comment left under one of them
Chapter 3, question 1. The learner swapped employee rights and responsibilities. Jenna leaves a comment under the weak answer, saves it with the portal's own Update button and records a provisional mark. The comment says what is missing, never what the answer is.
What Jenna recorded for this step

"Typing feedback for the rights of employees section to guide the student on how to correct their swapped answers."

"The first employee right given is missing or just repeats the first. Give another separate entitlement beyond the one already provided."

Chapter 3 summary page with per-question totals, the Pass/Refer selector left on Choose and the feedback box empty
End of Chapter 3. Five questions sit below their minimum, so the chapter is held: the Pass/Refer choice is left untouched, nothing is submitted, and the learner is notified of the comments. On the next run Jenna re-marks only the answers that were revised.
Chapter 1 summary page showing per-question totals, a total of 44 marks against a minimum of 32, and the Submit button
End of Chapter 1. Every question meets its minimum and the total clears the pass mark, so Jenna selects Pass, writes the overall feedback and submits. Four questions were docked, each with a comment.
The portal's image preview showing a learner's health promotion poster being read inside the assessor view
Chapter 6 is a poster. Jenna opens the preview, scrolls through the whole image or PDF and marks the five sections of the mark-scheme grid from what she sees.

How we checked it before trusting it

An agreement rate on its own tells you little. We wanted to know where the automation would be wrong, in which direction, and whether that direction is safe.

1. Ground truth: a learner the assessor had already marked

Before any live run, the marking standard was applied blind to a learner the centre's assessor had fully marked. The marks matched on 38 of 40 text-based criteria. On every one of the four questions where the assessor had deducted marks, the same mark was reached for the same reason.

Question (previously marked learner)Centre's assessorBlind reference markerShared reason
Ch1 Q124 / 64 / 6No stated outcome, so the second "detail" mark withheld
Ch2 SWOT6 / 86 / 8Extra opportunity and threat detail not met
Ch3 Q101 / 21 / 2"Cuts" is not an effect of a hazardous substance
Ch5 Q47 / 107 / 10Strict reading of the food examples

2. Live runs, second-marked blind

Student B's five written chapters were first-marked live, then independently second-marked from the run screenshots by a reference marker that had not seen Jenna's marks. Agreement was exact on one chapter and within one to three marks on the others. Every disagreement went the same way: Jenna was more generous, always on the "more detailed answer" criteria the guidance leaves to judgement. On one chapter that generosity pushed a below-minimum answer up to its minimum, a Pass where the strict reading says hold.

Student B, Level 2Jenna (first live run)Blind referenceNote
Ch1 Professionalism49 / 4946 / 49Three "detail" marks; two later corrected
Ch2 Development plan15 / 1514 / 15Reference misread a dropdown table; Jenna was right
Ch3 Health and safety37 / 3736 / 37One courtesy-versus-safety repeat
Ch4 Risk assessment20 / 2120 / 21Exact, including the same deduction and reason
Ch5 Client consultations36 / 3833 – 34 / 38Wrong concept on Q2; verdict flips to hold

Why this matters. If the automation is more lenient than the assessor, every Pass has to be re-checked and the saving disappears. A stricter first pass costs a learner one revision. A lenient one costs the centre a false Pass. So the standard was biased towards strictness, and the assessor's own marking was used to set the bar.

3. Calibrating on the assessor's own feedback

The portal was read with plain page requests, capped at one request per second, with no clicks and no changes: six learners across both levels, 942 requests, zero errors. That produced over 500 feedback entries, the assessor's own words on what was missing and what she expects. From those we derived a calibration layer, kept separate from the awarding body's text, and re-ran the questions Jenna had over-marked with the marks cleared. Seven of nine targeted questions moved to the reference mark; the two misses are genuine judgement calls. Four control questions stayed unchanged.

One lesson from this phase: re-marking a page that already shows a mark does not test strictness, because the model anchors on the visible number. Strictness is only testable on empty mark boxes.

Results

Two learners through Level 2 chapters 1 to 6 as of October 2026. Chapters 7 to 9 are practical assessments that are not marked online.

ChapterStudent B (September 9 – 11)Student S (September 21)
1 Professionalism and customer care47 / 49 · Pass44 / 49 · Pass
2 Personal and professional development plan15 / 15 · Pass13 / 15 · Pass
3 Health and safety36 / 37 · Pass24 / 37 · Held with comments on 5 questions
4 Risk assessment, maintenance, handover20 / 21 · Pass20 / 21 · Pass
5 Client consultations34 / 38 · One question below minimum, heldNot started: learner at 93 percent
6 Health promotion poster13 / 13 · PassPass

Every Pass is submitted under the assessor account and reviewed by the centre's assessor, who remains accountable for the result. Held chapters wait for the learner's revision.

What it costs

MarkerTime per chapterCost per chapter
Centre's assessor20 – 40 minutes≈ €7 – 13
Jenna, standard model≈ 36 minutes of browser time≈ €1.4
Jenna, high-effort model (used for marking)30 – 50 minutes of browser time, billed at 2.5×≈ €3.5

Student S's six chapters ran in parallel and finished in 37 minutes of wall-clock time, consuming about 296 automation minutes on the centre's subscription.

What we got wrong, and fixed

  • The first live run skipped questions and wrote no comments. The cause was in the engine's knowledge handling, not the mission. We fixed the root cause; every run since completes the full mark, comment, verdict and submit cycle.
  • We started writing marking criteria of our own. Where the guidance says only "a more detailed and accurate description", we had begun defining what that meant. The centre caught it. We reverted the same hour and recorded the rule: the guidance text only, and the bar belongs to the assessor.
  • Long answers stalled the reader. Multi-screen answers and identical comment boxes caused loops in two revalidation runs. Both were stopped deliberately, nothing was corrupted on the portal, and the engine gained exact page reading and exact field addressing.
  • PDF posters were invisible to the headless browser. The engine moved to a full Chromium headless mode so the portal's PDF preview renders, and Chapter 6 now marks end to end.
  • Scheduled polling wastes money. Checking every learner every two days burns minutes on logins that find nothing to do. Runs are launched when learners submit instead.

Where it stands

  • In use: Level 2 chapters 1 to 6, first-marked on demand, with the assessor confirming at the pass boundary.
  • Open decision for the centre: the portal's pass minimum is the sum of per-question minima, which is lower than the pass mark in the official guidance. The automation currently applies the stricter, official figure.
  • Next: extend the calibration cohort to 20 to 30 assessor-marked learners, have the assessor ratify the derived rules, then Level 3.

Regulators and awarding bodies are clear that AI must never be the sole marker. This setup is built on that premise: it is an assessor co-pilot with sign-off, not auto-marking.

Running a training centre with the same bottleneck?

We set this up on your existing portal, measure it against your assessor before you rely on it, and keep your assessor in charge.