Pipeline integrity. International Oil and Gas Service Provider · AI audit and on-premises pilot specification
An audit of how the gap analysis actually gets done, from the first request for information through to the final report, and an honest answer on whether a language model could carry any of it, written down well enough to defend either way.
Whole workflow audited, RFI to report 5 interviews across the lifecycle Concept validated, roadmap costed
What the audit looked at.
The client provides pipeline integrity consultancy to operators worldwide, and their audits sit underneath capital decisions, insurance positions and regulatory submissions. The methodology is proprietary and was refined over decades of field work. The step that matters most, a senior consultant cross-referencing an operator’s safety documentation against licensed industry standards clause by clause, was slow, repetitive and entirely dependent on the availability of engineers who were approaching retirement.
The brief had a succession edge to it. The sponsor wanted to know whether AI could take part of that load, and wanted the answer written down well enough to defend either way. Hosted assistants were out before the design conversation started, because audit content cannot leave the client environment.
- 01
Five stakeholder interviews across the consultancy and engineering teams, including ex-operators who had built management systems from the customer side.
- 02
The delivery process mapped end to end, marking where a model could realistically intervene and where human judgement has to stay.
- 03
A specification and a costed roadmap, with no pilot run during the engagement.
The constraints the design had to meet.
The list came out of the interviews and the client’s contractual position, and every design call was tested against it. The first line alone settled the hosting comparison.
- 01
Audit content stays inside the client environment, end to end.
- 02
Every finding stays the consultant’s call, with the system preparing the evidence behind it.
- 03
Every drafted finding cites the source pages it came from.
- 04
The proprietary methodology never reaches a hosted model.
- 05
The recommendation has to be defensible to procurement, engineering and a regulator.
Inside the two-stage pilot.
Stage one
Evidence retrieval, proved on real audits.
Stage two
Findings drafted, and marked by a consultant.
- 01
Stage one: finding the evidence
Staff can search the audit estate in plain language and get back the passages that answer the question, so assembling evidence stops taking days. It also puts the client’s own hardware through real workloads before anything riskier sits on top of it.
- 02
Stage two drafts, the consultant decides
A gap-analysis engine breaks the framework into testable criteria, drafts a finding against the evidence index for each, and routes every draft into a review interface. The review step is built into the architecture, so a draft cannot reach a customer-facing report without a consultant marking it.
- 03
The gate between the stages
If stage one cannot hit the agreed retrieval accuracy and stability targets against historical audits, stage two does not start. The cheaper, safer half proves itself before the judgement half is funded.
- 04
Client-hosted, with an audit trail
The pilot runs primarily on the client’s own GPU and orchestration stack, with any overflow inheriting the same data-handling posture. Retrieval splits documents on meaning rather than token counts, and an immutable trail records every accept, edit and reject.
How the hosting was decided.
Three deployment postures were compared with a total cost of ownership behind each, across open-source and commercial models, and assessed against the specific demands of reading regulatory text rather than general benchmarks. The comparison did not take long, because the first constraint had already decided most of it.
- 01
Pure cloud was eliminated by the first constraint: audit content cannot leave the client environment.
- 02
Pure on-premises, chosen for the pilot. The existing GPU capacity made it the cheapest route to a working prototype as well as the safest.
- 03
Hybrid, held in reserve for production, on the condition that overflow inherits the same data-handling posture.
What the study settled.
The result was a report, shared across the business, on connecting the estate’s data layers and sources, so that every department we interviewed can work from the same connected evidence rather than their own copy of it.
sessions across the audit lifecycle, each paired with an evidence request, so every claim in the report pointed at something we could look at.
pilot stages with a gate between them, sequenced so the retrieval half funds its own case before the judgement half is commissioned.
the recommended path end to end, because audit content leaving the client environment is a contractual line rather than a preference.
hosted models in the recommended pilot architecture. The methodology that took decades to build never reaches one.
We worked with OpenKit, as we wanted a company with domain expertise in LLMs to look at our strategy and test the concept. The study was focussed and really gave us a validation of our concept, a technical roadmap and prioritisation of the developments. The OpenKit team engaged well during the project, and through capture of workflows and interviews gained a good understanding of what we do. OpenKit provided a good technical study and exceptional value for money.
Specified
- Private retrieval over audit content
- Open-weight LLM on client GPU
- Kubernetes platform, client-owned
- Consultant review interface
- Immutable audit trail
OpenKit certifications
- ISO 27001
- ISO 9001, UKAS-accredited
- Cyber Essentials
Controls on this project
- UK data residency
- Client data sovereignty
More of the work.
Find your first workflow.
We start with a conversation, audit where AI actually pays back, and build the first automation into how your team already works. We reply within one working day.
Start the conversation AI Audit