Skip to content
Discuss your AI plans 

AI marking for lab quality assessment.

EMQN · discovery sprint, multilingual benchmark and technical specification

EMQN’s board could see how far AI marking goes before commissioning anything. The benchmark ran against their own reports, in every language their schemes run in.

EMQN's website homepage, shown in a browser window against a sculpted teal backdrop

Overview

EMQN is a community interest company that works as an exam board for laboratories, running external quality assessment schemes for human genetic testing. Expert volunteers mark every report. A large part of that marking is a set of objective, rule-based clerical criteria that assessors hand-check every time, and which they describe as the dullest and least scientifically interesting part of the job.

They engaged OpenKit to find out whether AI could take the first pass at those criteria without weakening a process that laboratory accreditation depends on. The team wrote a strategy report for the board and a technical specification for a build team, both resting on a benchmark run against EMQN’s own reports.

  • Five interviews before any model touched any data: the chief executive, the assessment lead, the scientific team, the head of IT, and the project manager who owns the platform integration.
  • Six models, each run ten times against fifteen representative reports in all six scheme languages, with the full spread of results reported.
  • Seven hosting options assessed against the jurisdiction rules the schemes operate under, so procurement had its answer in the same document.

93-96%

Per-criterion accuracy across the benchmark as a whole

45%

Strict accuracy, meaning all 17 criteria correct on the same report

6 models

Each run ten times against fifteen representative reports

6 languages

Every language the schemes run in, read separately

7 providers

Hosting options assessed, with UK and EU hosting only

Challenge

The constraints the design had to meet

EMQN’s schemes sit under accreditation and its members are diagnostic laboratories, so there was little room for a system that only looked promising in a demonstration. The requirements were fixed before any technology conversation started, and the architecture was designed to help meet them from the first sketch.

  • Laboratory data stays inside UK or EU jurisdiction, at every tier of the hosting stack.
  • The platform takes the first pass at the clerical criteria, and a volunteer assessor signs every mark.
  • The same criteria hold in all six languages the schemes run in.

Approach

What the AI audit produced

A model benchmarking framework

Six frontier models evaluated against fifteen representative reports, ten independent runs each, so EMQN can rerun the comparison as models change.

A multilingual evaluation suite

Accuracy read separately in all six languages the schemes run in, with the variance left in, because a model that only performs in English would fail the organisation’s real workload.

The data residency analysis

Seven hosting options weighed on US CLOUD Act scope, UK GDPR, running cost and the self-hosted alternative, with UK and EU hosted options recommended by design.

A traffic-light confidence design

In the assessor-facing UI concept, every draft arrives with its confidence made visible, so a person decides what to trust before anything is signed.

A third-party API risk register

The risks a community interest company would carry against any external AI provider, written down so procurement and governance start from the same page.

Findings

Per-criterion accuracy across the six supported languages.

Every model was run ten times so the sprint would show how far its answers moved between runs. Per-criterion accuracy landed high enough to be useful. Strict accuracy, meaning every criterion correct on the same report, dropped sharply. That gap is what put numbers behind the human-in-the-loop requirement, and it is why the recommendation buys assessors time and leaves the judgement with them.

The most accurate model in the set was ruled out anyway, because its licence does not permit deployment by an EU-based organisation, so the recommendation went to the next best one EMQN can actually run. The IT lead’s condition shaped the rest. A healthcare system cannot be benchmarked once and trusted thereafter, so re-verification against pre-assessed baselines is designed in as a normal operation.

The headline 93-96% is the per-criterion range across the benchmark as a whole.

All six scheme languages were tested: English, Spanish, German, French, Italian and Portuguese. Each language was benchmarked on EMQN’s own reports. The same results, cut by language:

  • Strict accuracy (all 17 criteria): 45%. Why human oversight stays in the loop.
  • German, French, Italian, Portuguese: 93-95%. Compound technical terms and diacritics.
  • English and Spanish: 99-100%. Highest-confidence languages.

Result

What EMQN took to its board

The board came out of the sprint with a strategy report and a technical specification a build team can work from, alongside interface mockups and an API questionnaire for the third-party platform the marking sits inside. Together they describe a platform that clears the clerical criteria first and leaves the scientific judgement to the assessor.

The hosting analysis went in the same pack, weighing seven providers on US CLOUD Act exposure, UK GDPR compliance, running cost and the self-hosted alternative. The options it recommends are hosted in the UK or EU by design, so they sit outside US CLOUD Act scope.

Two controls were set on this project: UK GDPR and data residency inside the UK or EU. OpenKit holds ISO 27001 and ISO 9001, both UKAS-accredited, plus Cyber Essentials.

Voice

Working with OpenKit has been a genuinely positive experience. Their team quickly understood the unique challenges of our business and the problem we were trying to solve, and delivered a thorough, evidence-based strategy for our AI-assisted marking platform. We were particularly impressed by their transparent approach, technical expertise, and commitment to long-term partnership and support.
Simon Patton · CEO, EMQN CIC

Other engagements

What could your business do with AI?

Find out with experts who become part of your team. We uncover opportunities, get ideas working, and help your people build on the results. We reply within one working day.

Get in touch  Governance & compliance