An AI-assisted marking platform. EMQN · discovery sprint, multilingual benchmark and technical specification
A benchmark against the client’s own reports, in every language their schemes run in, so the board could see where AI marking holds and where it does not before commissioning anything.
6 models tested 10 runs per model UK and EU hosting only
Where AI could take the first pass at marking.
EMQN is a community interest company that works as an exam board for laboratories, running external quality assessment schemes for human genetic testing. Expert volunteers mark every report. A large part of that marking is a set of objective, rule-based clerical criteria that assessors hand-check every time, and which they describe as the dullest and least scientifically interesting part of the job.
They engaged OpenKit to find out whether AI could take the first pass at those criteria without weakening a process that laboratory accreditation depends on. The deliverables were a strategy report for the board and a technical specification for a build team, both resting on a benchmark run against EMQN’s own reports rather than a demonstration.
- 01
Five interviews before any model touched any data: the chief executive, the assessment lead, the scientific team, the head of IT, and the project manager who owns the platform integration.
- 02
Six models, each run ten times against fifteen representative reports in all six scheme languages, with the full spread of results reported.
- 03
Seven hosting options assessed against the jurisdiction rules the schemes operate under, so procurement had its answer in the same document.
The constraints the design had to meet.
EMQN’s schemes sit under accreditation and its members are diagnostic laboratories, so there was little room for a pilot that only looked promising in a demonstration. The requirements were fixed before any technology conversation started, and the architecture was designed to help meet them from the first sketch.
- 01
Laboratory data stays inside UK or EU jurisdiction, at every tier of the hosting stack.
- 02
The platform takes the first pass at the clerical criteria, and a volunteer assessor signs every mark.
- 03
The same criteria hold in all six languages the schemes run in.
Per-criterion accuracy across the six supported languages.
Every model was run ten times so the sprint would show how far its answers moved between runs. Per-criterion accuracy landed high enough to be useful, while strict accuracy, meaning every criterion correct on the same report, dropped sharply. That gap is what put numbers behind the human-in-the-loop requirement, and it is why the recommendation buys assessors time while the judgement stays with them.
The most accurate model in the set was ruled out anyway, because its licence does not permit deployment by an EU-based organisation, so the recommendation went to the next best one EMQN can actually run. The IT lead’s condition shaped the rest: a healthcare system cannot be benchmarked once and trusted thereafter, so re-verification against pre-assessed baselines is designed in as a normal operation rather than an audit exercise.
The headline 93-96% is the per-criterion range across the benchmark as a whole. The chart cuts the same results by language, where German, French, Italian and Portuguese sit at 93-95% and English and Spanish reach 99-100%.
Tested in all six scheme languages
- English
- Spanish
- German
- French
- Italian
- Portuguese
Each language was benchmarked on EMQN’s own reports.
What EMQN took to its board.
The board came out of the sprint with a strategy report and a technical specification a build team can work from, alongside interface mockups and an API questionnaire for the third-party platform the marking sits inside. Together they describe a platform that clears the clerical criteria first and leaves the scientific judgement to the assessor.
The hosting analysis went in the same pack, weighing seven providers on US CLOUD Act exposure, UK GDPR compliance, running cost and the self-hosted alternative. The options it recommends are hosted in the UK or EU by design, so they sit outside US CLOUD Act scope.
Working with OpenKit has been a genuinely positive experience. Their team quickly understood the unique challenges of our business and the problem we were trying to solve, and delivered a thorough, evidence-based strategy for our AI-assisted marking platform. We were particularly impressed by their transparent approach, technical expertise, and commitment to long-term partnership and support.
What the AI audit produced.
- 01
A model benchmarking framework
Six frontier models evaluated against fifteen representative reports, ten independent runs each, so EMQN can rerun the comparison as models change rather than trusting a one-off result.
- 02
A multilingual evaluation suite
Accuracy read separately in all six languages the schemes run in, with the variance left in, because a model that only performs in English would fail the organisation’s real workload.
- 03
The data residency analysis
Seven hosting options weighed on US CLOUD Act scope, UK GDPR, running cost and the self-hosted alternative, with UK and EU hosted options recommended by design.
- 04
A traffic-light confidence design
The assessor-facing UI concept: every draft arrives with its confidence made visible, so a person decides what to trust before anything is signed.
- 05
A third-party API risk register
The risks a community interest company would carry against any external AI provider, written down so procurement and governance start from the same page.
The audit pack
- AI model benchmarking framework
- Multilingual evaluation suite
- EU data residency analysis
- Traffic-light confidence UI design
- Third-party API risk register
OpenKit certifications
- ISO 27001 · ISO 9001, UKAS-accredited
- Cyber Essentials
Controls on this project
- UK GDPR
- UK and EU data residency
More of the work.
Find your first workflow.
We start with a conversation, audit where AI actually pays back, and build the first automation into how your team already works. We reply within one working day.
Start the conversation Governance & compliance