Skip to content
Discuss your AI plans 

AI EngineeringLLM development & fine-tuning

LLM development & fine-tuning

When an off-the-shelf model isn’t accurate enough for your work, we tune one that is, and prove it against your own numbers. We become part of your team to build it. The AI Audit finds and costs the first one; building it is a separate decision.

Discuss your AI plans 

Most tasks don’t need a tuned model, and we say so.

Tuning earns its place when the language is specialist and the accuracy bar is high enough to measure. It pays at volume: marking, extraction, review and classification.

What we build

OpenKit builds and tunes language models on your own examples, and hands over the weights, the evaluation suite and the numbers they were judged against.

Fine-tuned models
Trained on your examples, for your task.
Evaluation suites
Your baseline turned into a benchmark the model must beat.
Prompt & pipeline engineering
Often the honest first answer before any tuning.
Open-weight deployments
Tuned models you own, run where you choose.
Continuous evaluation
Accuracy tracked in production.

Language models are one of four builds under AI engineering, alongside agents, retrieval systems and voice.

When general is not good enough

A custom LLM is worth the effort when the domain is narrow, the data is sensitive, or the same language task repeats often enough that accuracy and privacy start to matter more than breadth.

Domain language
Legal, clinical, engineering, or financial text where a general model misreads the terminology and a tuned one reads it the way your experts do.
Document extraction
Pulling structured fields out of long, messy documents, each one carrying a citation back to the source line so a person can check it. On the BAiSICS lease pipeline that runs at 96% extraction accuracy.
Private deployment
Data that cannot leave your walls: the model runs inside your own controlled environment, on hardware you own, so nothing is sent to an outside service.
Drafting in house style
First versions of routine documents written in your language and format, so a person edits rather than starts from a blank page.
Grounded answers
Questions answered from your own knowledge base with sources attached, usually by searching your documents first and letting a tuned model answer from what it finds.

Tuned and proven once the audit has costed it

  1. In the audit

    Ranked

    Task defined, examples gathered, ranked against the rest of the work

    • Human baseline measured
  2. On the roadmap

    Costed

    A cost, a place in the order of work and the hours it gives back, before anything is built

    • Build or buy settled
    • Sequenced by payback
  3. In implementation

    Built

    First tuned model and evaluation suite, inside an agreed delivery allocation or a separately scoped build

    • Error review with your experts
  4. In use

    Measured

    Benchmarked against the human baseline, monitored after launch

    • Reviewed with the experts who check its output

The smaller answer, when it exists

Four cases where tuning is the wrong spend, and we say so before anyone commits.

If retrieval is what you actually need, start with retrieval-augmented generation instead: grounding a model in your data is often the cheaper, more maintainable route.

Where a tuning engagement sits

Scoped in an AI Audit

The audit finds the task, measures the baseline, and costs the tuning on the roadmap before anything is built. Building it is a separate decision.

How the audit runs 
Built with an Embedded AI Lead

Iterative tuning cycles for marking, review or extraction work, run against a backlog you agree with us and inside the delivery allocation your engagement reserves.

How the embedded engagement runs 

The honest range for UK LLM work

Published UK market ranges for what a competent partner charges to take a model from scope to a working deployment. Building a frontier model from scratch costs millions and is almost never the right answer: fine-tuning or grounding an open model is where the money should go. OpenKit scopes each build to an outcome rather than a day rate, and our own fees are on the pricing page.

Fine-tuning & adaptation
£20k-£80k
Adapting an open model to your data and terminology.
Knowledge / retrieval system
£15k-£50k
Grounding answers in your documents; depends on data cleanliness.
Regulated environment
+10-20%
Security documentation, access controls, audit trails, evidence.

Source: OpenKit AI development cost guide, published UK market ranges.

…delivered a thorough, evidence-based strategy for our AI-assisted marking platform. We were particularly impressed by their transparent approach [and] technical expertise…
Simon Patton, CEO, EMQN CIC

Rubrical is a DfE-backed private marking assistant for GCSE geography, and teachers report a 52% reduction in marking time in the developer’s tests. EMQN’s marking platform for genetic testing laboratories achieved 93-96% per-criterion accuracy.

BAiSICS
A custom OCR-plus-LLM pipeline for commercial-lease review that reads poor-quality documents at 96% extraction accuracy, beating general models on the firm’s own leases. Every output is verifiable against the source.
International Oil and Gas Service Provider
An LLM strategy and an on-premises pilot specification for pipeline-integrity audits, with private retrieval over the audit content and an open-weight model on the client’s own GPUs. The specification also sets out an immutable audit trail, and we wrote it precisely enough for their own team to build from.

What teams ask before tuning anything

Do we need our own model, or is GPT enough?

Usually GPT-class models with good engineering are enough. The audit says which side of the line your task sits on.

How much data do we need?

Often only hundreds of good examples, which beat thousands of poor ones for fine-tuning. The first pass over your examples tells us how many yours are.

Whose model is it afterwards?

Yours. Tuned weights on open-weight bases are handed over like any other build, with the documentation to run them. The engagement covers access, handover and who runs it. The base model keeps its own licence.

How do you measure accuracy?

Against humans doing the same task. We measure the baseline before any tuning starts, and the model has to beat it before rollout.

Does it keep learning?

Not silently. Retraining is a deliberate, evaluated step: accuracy is tracked in production and tuning cycles are scheduled, not automatic.

What is a custom large language model, and why build one?

A custom large language model is one adapted to your data, terminology, and workflows rather than the general internet. You build one when a general model keeps getting your domain wrong, or when your data cannot leave your environment. OpenKit scopes custom LLM development around a specific business problem: reading your documents, drafting in your house style, or answering with your policies.

What is the difference between a custom LLM and using ChatGPT or Claude?

A general assistant is trained on public data and, on the public tiers, your prompts may leave your control. A custom LLM is adapted to your terminology and runs where you choose: on your own infrastructure the data never leaves it, and where a build uses a hosted provider we contract on terms that keep your prompts out of anyone’s training data. For narrow, repeated tasks a tuned model is more accurate and more private; for open-ended work, a general model is often the right tool, and we will say so.

Do you train a model from scratch or fine-tune an existing one?

Almost always we adapt an existing open-weight model rather than train from scratch. Training a frontier model from zero costs millions and rarely beats fine-tuning or grounding an open model on your data. OpenKit fine-tunes, prompts and grounds open-weight models so your team can run and update the result under the arrangements we agree.

What can a custom LLM do for my business?

The strongest cases are document-heavy and language-heavy: extracting fields from contracts, classifying and routing correspondence, drafting first versions of routine documents, and answering questions from your own knowledge base. It understands your context, which is where general tools fall down, and where the BAiSICS legal platform, for example, surpassed general models on the firm’s own leases.

How do you keep our data private?

Client data is used only to deliver the engagement. Prompts and outputs are not retained afterwards and are never used to train external models. The model itself runs where your data already lives, in a UK-region cloud or on your own servers.

Is a custom LLM secure enough for regulated data?

It can be, when built for it. OpenKit is ISO 27001 and ISO 9001 certified and holds Cyber Essentials. We add role-based access, data-residency controls and an audit trail of every prompt and response. We design to operate under UK GDPR.

What does large language model development cost in the UK?

Published UK market ranges put fine-tuning and adaptation work at roughly £20,000 to £80,000, with a knowledge or retrieval system from around £15,000 depending on how clean the data is (OpenKit AI development cost guide). Regulated environments add ten to twenty percent. OpenKit scopes each build against a defined outcome rather than a day rate, and our own fees are on the pricing page.

Can the model run on our own infrastructure?

Yes. An open-weight model can run inside your own controlled environment, on hardware you own or in a private cloud, so prompts and documents stay on your network. That is the pattern OpenKit specified for an International Oil and Gas Service Provider. It takes more setting up than a hosted deployment, so we recommend it where the sensitivity or the volume justifies it and a UK region where they do not.

Not ready to talk? The free AI readiness check scores where you stand in about five minutes.

What could your business do with AI?

Find out with experts who become part of your team. We uncover opportunities, get ideas working, and help your people build on the results.

We reply within one working day.