Skip to content

Inside the Rubrical build: measuring the 52%

How OpenKit built Rubrical for the DfE: a private AI stack, continuous evaluation, and teachers signing every grade. Teachers report 52% less marking time.

Ibrahim Mizi Ibrahim Mizi  · 5 min read Updated
Rubrical AI marking platform built by OpenKit for the Department for Education

The Department for Education contracted OpenKit to change how formative assessment marking works, starting with GCSE Geography. The public framing was time, because marking is the load that eats teachers’ evenings. The engineering framing was trust, because a system that touches assessed work has to keep the student declaration trail intact and stand up to anyone who later asks how a mark was reached, and that constraint ruled out every off-the-shelf tool before we started. This is how Rubrical got built, and why it looks the way it does.

Published 8 July 2026.

The number that had to be earned

Before building anything, we benchmarked. We took GCSE Geography long-form responses already graded by senior teachers, held them out, and scored generic AI tools against the same rubric. They matched the human grader 67% of the time, which is impressive at a demo and indefensible at moderation.

Rubrical, a private stack that retrieves from the DfE content store and is evaluated continuously with the KILN framework, reached 92% on the same held-out marks. That 25-point gap is the whole argument for building against your own domain instead of prompting someone else’s model. The programme’s public target was a 50% reduction in marking time, and teachers report a 52% reduction in marking time in the developer’s tests, which the Department for Education records on page 98 of its February 2026 schools white paper, Every child achieving and thriving. Feedback came back around 2.5 times more detailed than teachers had time to write by hand, with satisfaction at 95%.

The teacher signs every grade

The design principle that survived every iteration is that the model drafts and the teacher signs. Typed and handwritten responses go in, rubric-aligned draft feedback comes out, and the teacher edits, overrides, and signs each grade while the moderation trail records all of it. Nothing reaches a student unsigned.

It is slower than full automation, and it is the version a trust can defend to governors, to parents, and to an exam board, because every mark has a named human behind it and a record of how it was reached. In regulated work the audit trail is part of what a school is actually buying.

Private by necessity, not preference

Student work could not leave controls the trust could evidence, so Rubrical runs on a private stack in a UK Azure tenant rather than a public API. OpenKit holds ISO 27001 and ISO 9001, the deployment operates to UK GDPR, and it passed a DfE data protection impact assessment before it scaled.

The rollout followed the same discipline. A closed pilot in three schools, the ISO 27001 audit on the deployment closed before scale-up, then fourteen schools in production behind a trust-wide moderation framework and service-level agreements. The work continues, with model refreshes against new marking data and a jointly authored impact report.

As Reuben McQueen, OpenKit’s CTO, put it at the programme announcement: “Our goal is reducing marking time by 50% while maintaining high-quality feedback, with level-appropriate explanations and follow-up exercises tailored to each student.”

What the build teaches

Rubrical is the pattern we use when the stakes are real: a private model answering from approved content it retrieves at the point of use, evaluated against ground truth instead of vibes, with a human signing every decision that matters. The full Rubrical case study carries the rollout detail, and AI marking for teachers is the teacher-facing version. It is also the approach behind our wider work on AI in education, where the same rule holds: the tool earns trust or it does not ship.

Start with an audit

Most engagements start with an AI Audit and Transformation, a fixed-scope, fixed-fee piece of work that finds where AI earns its place in your business and where it does not, then puts the first of it into service. You leave with a written report and a prioritised 12 month roadmap.

Ibrahim Mizi

Ibrahim Mizi

Co-founder & CEO · Full-Stack AI Engineer · OpenKit

Co-founded OpenKit in 2020 and runs the consultancy side end to end. Eight years of full-stack development, then production AI for SMEs and the public sector.

What is Rubrical?

Rubrical is an AI marking assistant OpenKit built for the Department for Education, starting with GCSE Geography. It drafts rubric-aligned feedback that a teacher edits and signs, so the model does the volume and the teacher keeps the judgement. Teachers report a 52% reduction in marking time in the developer's tests, a figure the Department records in its February 2026 schools white paper, Every child achieving and thriving.

How accurate is Rubrical compared with generic AI?

On GCSE Geography responses already graded by senior teachers and held out as a test set, generic AI tools matched the human grader 67% of the time. Rubrical, which retrieves from the DfE content store and is evaluated continuously, reached 92% on the same set. That 25-point gap is the case for building against your own domain instead of prompting a general model.

Why does the teacher sign every grade?

Because a school has to defend its marks to governors, parents, and exam boards. Rubrical drafts, the teacher overrides and signs, and the moderation trail records who changed what. Full automation would be faster and indefensible at moderation, so the slower design is the deliberate one.

Was Rubrical safe to deploy on student data?

It runs as a private retrieval-augmented stack in a UK Azure tenant rather than sending work to a public model. OpenKit holds ISO 27001 and ISO 9001, the deployment operates to UK GDPR, and it passed a DfE data protection impact assessment before it scaled to production.

Find your first workflow.

We start with a conversation, audit where AI actually pays back, and build the first automation into how your team already works. We reply within one working day.