How accurate is our marking?
Marking you can’t trust is worse than no marking. So we test Redpen against Cambridge’s own mark schemes, and publish what we find, including where it gets things wrong.
How we tested it
- 1
Real questions and real mark schemes
We took written-answer questions from 30 recent Cambridge papers: 14 for A Level 9618 and 16 for IGCSE 0478, from 2023 to 2026. Each question part came with its published mark scheme: the mark points, and the notes on what to accept and what not to accept.
- 2
Three answers to every question
For each question part we wrote three answers: one that earns full marks, one that earns some of them, and one that should earn nothing because it uses wording the scheme says not to accept. The expected mark for each answer was worked out from the scheme’s own rules.
- 3
Marked the way students are marked
Each answer went through exactly the marking code students get, given the scheme’s real mark points. Then we compared the mark it gave with the mark the scheme gives. A question is only right if the total matches exactly.
- 4
Published, mistakes included
The numbers below are the whole run, not the best questions. The 9618 run was done after the marker was built, with no changes made for it. The 0478 run was used to tune how the marker applies 0478 conventions, which took exact agreement from 97.4% to 98.1%.
The full results
| A Level 9618 | IGCSE 0478 | |
|---|---|---|
| Answers marked | 351 (117 question parts) | 378 (126 question parts) |
| Exact mark | 98.6% | 98.1% |
| Within one mark | 99.7% | 99.7% |
| Gave more marks than the scheme | 0.3% | 1.1% |
| Gave fewer marks than the scheme | 1.1% | 0.8% |
| Exact: full-mark answers | 99.1% | 100% |
| Exact: part-mark answers | 96.6% | 96.0% |
| Exact: answers worth nothing | 100% | 98.4% |
| Total marks | 532 expected, 528 awarded | 557 expected, 560 awarded |
Where it goes wrong
Part-mark answers are the hardest, because an answer that half-makes a point has to be judged against the scheme’s notes. Fewer than one part-mark answer in 25 got a different mark from the scheme, and nearly all of those were one mark out.
On 9618, when the marker was wrong it was more often too strict (1.1% of answers) than too generous (0.3%). On 0478 the two were closer: 1.1% too generous and 0.8% too strict. The single worst case was a 0478 answer the scheme gives nothing that the marker gave 3 out of 6.
Code, SQL and logic are run, not read
Pseudocode answers are run against test cases, SQL queries are run against the question’s table, and logic expressions are checked against every row of the truth table. Full marks for code are decided by running it, not by an AI’s judgement.
Examiners also give part marks: a mark for each thing a program does right, even when one bug or a missing keyword stops it working. So when pseudocode fails a test, an examiner-style marker checks it against each requirement in the question, credits the ones it meets, and says what to fix. It can’t award full marks while a test fails, and never fewer than the share of tests the code passed.
Diagrams marked from a photo
For diagram questions, students upload a photo. We tested this on 30 diagrams with known correct marks: 5 logic circuits, 4 flowcharts and a structure diagram, each drawn in several versions from fully correct to wrong. After two changes it marked all 30 correctly: we reworded three mark points that depended on each other, and when two marking passes disagree on a point, a third now decides it. The diagrams were generated, not photographed by students, and the set is small.
What we haven’t measured yet
- Answers from real students. The test answers were written for the test, so they are cleaner than a real script written under time pressure.
- Photos of real handwriting, taken on real phones in poor light.
- The 15-mark 0478 programming scenario, which is marked in bands. It has been checked on examples, not benchmarked like the questions above.
- Part marks for pseudocode that fails a test. These are new, and checked on examples so far.
When we have run these, the results will go on this page.
Why we don’t claim 100%
No marker is perfect, human or not. Exam boards check samples of their examiners’ marking for exactly this reason. A number you can check, with the method next to it, is more useful to you than a promise. The exam wording itself stays private, because it belongs to Cambridge, but everything else about the test is described here.
See it mark your own answers.
The free audit is six questions from your papers. No card needed.
Redpen is independent and is not affiliated with, or endorsed by, Cambridge Assessment International Education. These tests were run by Redpen, not by Cambridge or an independent auditor.