← Back to School Blog

AI Marking for A Level and IGCSE: How Accurate Is It and How Should Schools Validate It?

How accurate AI marking is for IGCSE and A Level answers, what Ofqual's 2026 working paper and Cambridge research say, how human marker agreement compares, and a step-by-step protocol schools can use to validate any AI marking tool.

Last reviewed:

ai marking for a levelai exam marking softwareai marking for teachersai marking igcseai grading software for teachers

There is no single answer: accuracy depends on the question type and the tool, and no independent study has measured AI marking on IGCSE or A Level at scale. Ofqual concluded in January 2026 that evidence “does not support an overall or general case for the validity of AI marking in high stakes qualifications”. Test any tool against your teachers’ marks.

How accurate is AI marking for IGCSE and A Level?

Unproven in general, and only as good as the evidence for a specific tool and question type. Ofqual’s working paper, Principles of AI use in marking, found evidence that “points to the need for context-specific evidence”. It also warns that “empirical studies have identified variable AI marking performance across demographic groups”. The published accuracy figures available to schools come from vendors’ own tests on small samples. We compare them in the best AI marking tools for IGCSE and A Level.

Cambridge’s own researchers are cautious too. A 2025 paper in Research Matters notes that large language model marking shows “potential in terms of increasing accuracy in comparison to previous automarking methods”. It still argues that “human examiners have a higher potential for trustworthiness over LLM-based automarkers” (Morley & Walland, 2025). It cites research where adding irrelevant words to answers reduced the accuracy of some earlier automarkers by 10 to 22%.

How accurate are human markers?

It helps to know the benchmark. Human marking is very consistent on objective items and less so on extended answers. A Cambridge review summarised mean agreement between markers of 0.97 for objective items, 0.82 for points-based items and 0.77 for levels-based items (Curcin, 2010). Ofqual found that the probability of a candidate receiving the “definitive” grade ranged from 0.96 in a mathematics qualification to 0.52 in an English language and literature qualification. The probability of the definitive or an adjacent grade was above 0.95 for all qualifications (Ofqual, 2018).

Question typeHuman marker agreementWhat it means for AI marking
Objective (e.g. multiple choice, single answers)Very high (0.97)AI should match humans closely; errors are easy to spot
Points-based (short answers against a mark scheme)High (0.82)AI can be useful for practice; check against teacher marks
Levels-based (essays, extended writing)Lower (0.77)Hardest for AI and humans; teacher judgement essential

What does Ofqual say a school should check?

More than agreement. Ofqual’s paper says “agreement with human marks alone is insufficient for assuring validity”. It lists the measures used to report agreement: “percentage agreement”, “kappa”, “quadratic weighted kappa (QWK)”, “mean squared error” and “correlation coefficients”. It also stresses “looking at score distributions, and the size and nature of marking discrepancies (including outliers)”. There is no single pass mark: “an appropriate threshold needs to be defined and justified”. For short answers, it cites the aim that “at least human-level performance should be aimed for in general”.

It also flags two traps when humans check AI marks. Reviewing every AI mark gives “no efficiency gain”. It also creates an “anchoring effect”, where the reviewer accepts the AI’s mark too readily. Marking in parallel, with humans and AI independently, “avoids anchoring”.

How should a school validate an AI marking tool?

Run a blind, side-by-side test before using it for anything that matters. A protocol based on the measures above:

  1. Choose the question types you want AI to mark, such as short-answer science or maths working. Test each type separately.
  2. Take a real sample: at least one full class set of your own students’ answers.
  3. Mark in parallel. Teachers mark to the official mark scheme without seeing the AI marks, and the AI marks without teacher input.
  4. Measure: exact agreement, agreement within one mark, and the average difference, positive or negative, which shows systematic over- or under-marking.
  5. Check fairness. Compare differences for different groups, such as students writing in an additional language.
  6. Read the feedback, not just the marks. Is it correct and specific?
  7. Set a threshold and a role. Decide the level of agreement you need for each use, such as practice only or first-pass marking with teacher review. Record the decision.

Repeat the test when the tool updates or the syllabus changes.

What are the rules for using AI marking?

AI may support marking but may not be the sole marker of work that counts. Cambridge’s 2026 Handbook says “teachers must not use artificial intelligence (AI) as the sole or primary means of marking candidates’ work”. JCQ says “an AI tool cannot be the sole marker”, and Pearson’s guidance agrees. Ofqual expects AI marks to become “more acceptable in contexts such as formative and low-stakes assessments”. Practice, quizzes and homework are where schools can use AI marking now.

How schools do this with AI Buddy

AI Buddy marks students’ practice answers to past-paper-style questions instantly, with feedback, for formative use, not for coursework or final assessments. We encourage schools to run the validation protocol above on AI Buddy’s marking for their own subjects before relying on it, and teachers see every mark on their dashboards. We have not published an independent accuracy study, and we would rather schools test AI Buddy on their own students than take our word for it.

Frequently asked questions

How accurate is AI marking for A Level?

It varies by tool and question type, and there is no independent large-scale study for A Level. Ofqual says current evidence does not support a general case for AI marking in high-stakes qualifications.

Can AI exam marking software mark IGCSE answers?

It can mark practice answers against mark schemes. Cambridge rules say AI must not be the sole or primary marker of candidates’ assessed work.

How should teachers check AI marking?

Mark a class set in parallel with the AI, blind to its marks, then compare exact agreement, agreement within a mark, systematic differences and feedback quality, by question type.

Is AI marking as reliable as a human examiner?

Not proven. Human agreement is itself lower on essays than on short answers, and Cambridge researchers argue human examiners currently have “a higher potential for trustworthiness”.

Discover how AI Buddy helps schools strengthen teaching, learning and evidence-informed school improvement. Or start a short consultation with our schools team using the form below — we will get back to you directly.

Sources

Explore how AI Buddy supports international school implementation.

View case studies
See AI Buddy in action Request a Demo
Book a Tutor