Kaption. Agentic AI for controls assurance
A per-control AI agent that reads the evidence, reruns the numbers in an isolated sandbox, and returns a verdict in which every claim is a clickable citation, down to the exact spreadsheet cell or the boxed region of a document page.
Minutes per control tested Whole population, every period Concept to production platform
The problem.
Internal audit and controls teams spend most of their time on mechanical work: opening evidence files, re-adding spreadsheet totals, hunting for sign-offs, and writing up what they found. Testing a single control can take hours, and testing every control every period is impossible by hand, so teams sample, and exceptions slip through the gaps between samples.
Our client asked a direct question: could an AI agent perform this testing end to end, at a standard an experienced reviewer would sign off, with every conclusion traceable back to its source?
- 01
One agent per control, re-performing the work rather than summarising what the files say.
- 02
Every factual claim is a live citation into the cell or the page region that supports it.
- 03
Anything uncertain goes to a human reviewer, and the system is calibrated to flag a control rather than pass it.
A pipeline that works like a careful reviewer.
Rather than asking one large model for an opinion, we built a multi-stage pipeline that moves through a control the way a careful human does, with each stage answering one question before the next one starts.
- 01
Validate the rule
What exactly does this control require, and is the requirement grounded in the policy and standards that apply?
- 02
Verify the sources
Are the evidence files authentic, complete, and the right ones for the period under test?
- 03
Re-perform the work
Recalculate the numbers, trace the approvals, check the dates. Nothing is taken on trust.
- 04
Conclude
A verdict with plain-language reasoning: effective, exception raised, or refer to a human reviewer.
Each stage runs in isolation and produces its own cited finding, so a failure at any point is visible, explained and attributable. Uncertain results are routed to a human rather than resolved by guesswork, and the calibration deliberately over-flags: the system will raise a control for review long before it will wave one through.
Every claim is one click from its evidence.
The defining feature of the system is the citation trail, and every factual claim the agent makes is a live reference. Click a reference in a finding and the source opens beside the conversation: a spreadsheet opens on the right sheet with the cited cells highlighted, rendered with the workbook’s original formatting, and a document opens on the right page with the cited region drawn as a bounding box.
Checking a claim costs the reviewer one click.
Numbers are recomputed in code.
Arithmetic does not go through the language model. Wherever a control involves numbers, the agent writes and runs code in an isolated sandbox: it re-adds the schedule, reconciles the totals, and compares its own result with the workbook’s. The verdict follows from that computation, and when the agent’s total matches the cited cell, the agreement becomes part of the evidence trail.
Questioning and overriding the agent.
Testing runs as a conversation. The reviewer hands the agent a control and its evidence, watches the working notes stream in plain language, and can interrupt at any point to ask why a stage failed, request a re-test with corrected evidence, or challenge a conclusion. A reviewer can override a verdict with a written justification, and both the agent’s original finding and the override stay on the record.
Measured against experienced reviewers.
We validated the system the way an audit methodology would demand. Controls were tested repeatedly and the agent’s verdicts were compared, stage by stage, against determinations made by experienced human reviewers. Where the agent disagreed with the humans, every mismatch was root-caused and written up, and the causes fed directly back into the design. The validation work settled how the system behaves:
- 01
Fail safe
The dangerous error is a false pass. The system is calibrated so that when it errs, it errs towards flagging a control for review.
- 02
Misses are investigated
Disagreements with human reviewers are treated as findings about the system, investigated with the same discipline the system applies to controls.
Built for a regulated environment.
The platform was built for a regulated financial-services context from the first commit, so the things a compliance team asks about last were settled first.
- Residency
- EU data residency throughout, on serverless cloud infrastructure.
- Retention
- Zero-retention document processing: evidence is parsed, never kept by third-party services.
- Traceability
- Versioned prompts and full execution traces, so any verdict can be reproduced and inspected after the fact.
- Isolation
- Isolated, ephemeral sandboxes for all agent-executed code.
- Access
- Human-in-the-loop controls, role-based access and a complete override audit trail.
The build
- Multi-stage agent pipeline, one agent per stage
- Isolated, ephemeral code sandboxes
- Document and spreadsheet extraction with zero retention
- Cell-level and page-region citation layer
- Versioned prompts with full execution traces
- Serverless cloud, EU region
OpenKit certifications
- ISO 27001
- ISO 9001
- Cyber Essentials
Built into this platform
- EU data residency
- Zero-retention document processing
- Reproducible execution traces
- Human-in-the-loop review with role-based access
- Override audit trail
ISO 27001, ISO 9001 and Cyber Essentials are OpenKit certifications, and we operate to UK GDPR. The controls beside them are the ones built into this platform for this engagement.
From concept to commercial product.
per control tested, where the same work by hand ran to hours.
the whole control population tested, where teams used to settle for a sample.
the platform is live and generating revenue for the client.
The client now has working papers in which every claim is one click from its evidence, and reviewers get their time back for the part of the job that actually needs human judgement.
OpenKit carried this engagement from first concept to a full production platform, combining strategic consulting with hands-on AI software development. That platform now sits at the heart of the client’s commercial offering, and it has helped them win opportunities and contracts with large enterprise customers.
Alongside delivery, we supported the client in building their own internal development capability, so the product keeps growing in their hands.
More of the work.
Find your first workflow.
We start with a conversation, audit where AI actually pays back, and build the first automation into how your team already works. We reply within one working day.
Start the conversation AI Engineering