Skip to content

AI marking for teachers: how Rubrical works

Rubrical is an AI marking assistant OpenKit built for the DfE. Teachers report 52% less marking time, and every grade stays teacher-signed.

Ibrahim Mizi Ibrahim Mizi  · 5 min read Updated
Rubrical AI Marking Dashboard for Teachers

Rubrical is an AI marking assistant OpenKit built for the Department for Education, starting with GCSE Geography. It drafts rubric-aligned feedback that the teacher edits and signs, and teachers report a 52% reduction in marking time in the developer’s tests, while the moderation trail stays defensible. This guide is the honest version of how it works, how accurate it is, and where a teacher stays in control.

Last updated 10 August 2026.

Why marking is the problem worth solving

Workload is the reason teachers most often give for leaving, and marking is one of its heaviest components, a pattern the DfE’s own Working Lives of Teachers and Leaders survey keeps reporting. Long-form subjects like geography make it worse, because grading an extended answer well is slow and grading it fast usually means grading it thinly.

The Department for Education ran a programme to test whether AI could take the administrative weight off marking without taking the judgement away from teachers. OpenKit was selected to build a tool for it, and the brief was specific: assess handwritten work and hand-drawn geography maps, give useful feedback, and never let the machine be the one that decides a grade.

What Rubrical does

Rubrical takes a student’s work, typed or handwritten, and returns draft feedback aligned to the marking rubric. The teacher reviews it, changes what they disagree with, and signs the grade. That is the whole loop, and the order of it is the point: the model drafts, the teacher decides.

Rubrical dashboard showing a teacher’s list of classrooms.

It reads handwriting through text recognition, so a scanned script goes in the same way a typed one does. It syncs classes from Google Classroom and Microsoft Teams, or runs on a manual setup for schools that use neither. And it reports back across a class, so a teacher can see where a whole cohort missed the same thing rather than rediscovering it thirty times.

The accuracy had to be earned, not claimed

Before building anything, we benchmarked. We took GCSE Geography long-form responses already graded by senior teachers, held them out, and scored generic AI tools against the same rubric. They matched the human grader 67% of the time. That is fine for a demo and indefensible at moderation.

Rubrical, a private retrieval-augmented system answering from the DfE content store and evaluated continuously against real marks, reached 92% on the same held-out set. That gap is the argument for building against your own domain instead of prompting a general model. The programme’s public target was a 50% reduction in marking time, and teachers report a 52% reduction in the developer’s tests, which the Department for Education records on page 98 of its February 2026 schools white paper, Every child achieving and thriving. Feedback came back around 2.5 times more detailed than teachers had time to write by hand, with satisfaction at 95%.

Rubrical’s marking interface showing draft feedback on a student’s geography answer.

The teacher signs every grade

The design principle that survived every iteration is that nothing reaches a student unsigned. Rubrical drafts, the teacher overrides and signs, and the moderation trail records all of it. In a setting where a trust has to answer to governors, parents, and exam boards, an audit trail of who marked what is part of what the school is actually buying.

That is also why Rubrical runs on a private stack in a UK Azure tenant rather than a public model. Student work stays inside controls a school can evidence: OpenKit holds ISO 27001 and ISO 9001, the deployment operates to UK GDPR, and it passed a DfE data protection impact assessment before it scaled.

As Reuben McQueen, OpenKit’s CTO, put it at the programme announcement: “Our goal is reducing marking time by 50% while maintaining high-quality feedback, with level-appropriate explanations and follow-up exercises tailored to each student.”

Where this fits

If you want the engineering detail, inside the Rubrical build covers how the private stack and the evaluation framework were put together, and the full Rubrical case study carries the numbers and the rollout. The same pattern, a private model answering from approved content with a human signing every decision, is how we approach AI in education generally, not just marking.

Start with an audit

Most engagements start with an AI Audit and Transformation, a fixed-scope, fixed-fee piece of work that finds where AI earns its place in your business and where it does not, then puts the first of it into service. You leave with a written report and a prioritised 12 month roadmap.

Ibrahim Mizi

Ibrahim Mizi

Co-founder & CEO · Full-Stack AI Engineer · OpenKit

Co-founded OpenKit in 2020 and runs the consultancy side end to end. Eight years of full-stack development, then production AI for SMEs and the public sector.

Is AI marking accurate enough for teachers to trust?

On held-out GCSE Geography responses already graded by senior teachers, Rubrical matched the human grader 92% of the time, against 67% for generic AI tools with no domain tuning. The number matters because a teacher reviews, edits, and signs every grade before it reaches a student, so the model drafts and the teacher decides.

Does the teacher stay in control of the grade?

Yes. Rubrical drafts rubric-aligned feedback, and the teacher edits it, overrides it where they disagree, and signs each grade. Nothing reaches a student unsigned, and the moderation trail records who changed what. It is slower than full automation on purpose, because that is the version a school can defend at moderation.

Is student data safe with Rubrical?

Rubrical runs on a private retrieval-augmented stack in a UK Azure tenant, answering from the DfE content store rather than sending work to a public model. OpenKit holds ISO 27001 and ISO 9001, the deployment operates to UK GDPR, and it passed a DfE data protection impact assessment. Student work stays inside controls a trust can evidence.

Which subjects and levels does Rubrical cover?

Rubrical was built for the Department for Education programme starting with KS4 Geography long-form responses, including handwritten answers read through text recognition. The approach, a private stack answering from approved content with continuous evaluation, is not tied to one subject, but the current tool is built around the geography rubric it was evaluated against.

Does Rubrical work with Google Classroom and Microsoft Teams?

Yes. Rubrical syncs classes and rosters from Google Classroom and Microsoft Teams, with a manual option for schools that use neither. Typed documents, images, and handwritten scripts all go in, and rubric-aligned draft feedback comes back for the teacher to review.

Find your first workflow.

We start with a conversation, audit where AI actually pays back, and build the first automation into how your team already works. We reply within one working day.