Skip to content
Get in touch 

Where AI fits in a manufacturer's back office

Invoice matching, order entry, supplier chasing and quality records: where AI holds up in a UK manufacturer's back office, and why the old RPA kept breaking.

Ibrahim Mizi Ibrahim Mizi  · 6 min read
A conveyor belt with pale roller rings carrying three documents, the middle one azure

The back office of a mid-sized manufacturer runs on documents that never stop arriving. Orders come in as PDF attachments and someone re-keys them into the ERP, invoices wait in a queue to be matched against purchase orders and delivery notes, and a decent slice of every week goes to chasing suppliers who haven’t confirmed. The quality file for each batch gets assembled by hand from records that live in three systems and a shared drive.

Manufacturers can see what that costs, which is why the industry surveys read two ways at once. In this year’s Make UK and PwC executive survey, 65% of manufacturers believe the opportunities will outweigh the risks in 2026 and the sector is putting money into digital and AI, yet the same survey puts a number on the catch: “When it comes to AI and automation, 60% say that skills are the major barrier to adoption.”1 This article is for the person who runs IT and operations inside that gap: where AI holds up in the back office, and what it takes to run it when IT is three people with no spare capacity.

Why the RPA you tried kept breaking.

Most manufacturers that automated the back office once already have scar tissue. The RPA generation of tools worked by position: the bot knew the PO number lived in the top right of the supplier’s invoice, so it read the top right. When a supplier redesigned their template the bot either stopped or, worse, kept going and filed the wrong number, and every layout change across a few hundred suppliers became a ticket for whoever set the rules up. The economics died on upkeep, long before anyone questioned the automation itself.

What changed is that language models read a document the way a person does, by what the fields mean rather than where they sit. An invoice with the PO number moved, renamed or buried in a footer still reads correctly, because the model is looking for a purchase order reference instead of a pixel position. That is the honest extent of the improvement. These models are probabilistic, so a small share of reads will be wrong, and a serious build treats that as a design fact: every extracted value carries a confidence score, low-confidence documents route to a person before they touch the ERP, and the straight-through rate is measured live rather than promised in a sales deck.

Four back-office jobs worth pricing.

When the ONS asked UK businesses what prevents or delays AI adoption, the most commonly reported factor was difficulty identifying use cases, ahead of cost and lack of expertise.2 The general method for finding and ranking candidates is in how to identify AI use cases; what follows is the manufacturer’s shortlist drawn in advance, because these four jobs are where the hours sit in almost every back office we look at.

Invoice matching

Three-way matching is the strongest first candidate in most manufacturing back offices because the rules already exist: the invoice has to agree with the purchase order and the delivery record, and today a person does the agreeing. A model reads the invoice whatever its layout, proposes the match, and posts only what clears the confidence bar, while everything else lands in a queue with the reason attached. The saving is counted in clerk hours per week, which is exactly the kind of number a payback case can carry.

Order entry

Orders that arrive as PDFs or emails and get re-keyed into Dynamics are the same read-and-transcribe problem from the other direction. The added work is validation: part numbers checked against your catalogue and prices against the agreed list before a person approves the order the system has drafted. The person stays in that loop because a mis-keyed order costs more than a mis-keyed invoice, and the cost of a wrong answer is what decides how much checking each workflow keeps.

Supplier chasing

The least discussed candidate is often the quickest to pay for itself. Chasing is composition plus bookkeeping, writing the same email again and keeping the list of who hasn’t replied, and an assistant can draft every chase from the open order book, track the responses and escalate the ones that go quiet. Your buyer still owns the relationship and reads what goes out; what they stop doing is the remembering.

Quality documentation

Quality records carry the heaviest document load, because a batch file is assembled from certificates of conformity, inspection records, test results and delivery paperwork, each arriving in whatever format the supplier or the line produces. The pattern that works here is retrieval over your own documents: the system finds and cites the source record instead of retyping it, and an engineer signs the assembled pack. We ran the audit version of this for an International Oil and Gas Service Provider whose consultants read dense technical documents against standards all day; the recommendation was private retrieval over their own audit content on their own hardware, specified precisely enough that their in-house team took the pilot forward without us in the room.

The engineering side of the paperwork, technical documentation and tender answers, is its own subject and lives on our manufacturing and mobility page.

What low maintenance means with three people in IT.

With three people in IT, anything that depends on a standing in-house engineering effort is disqualified before it starts, and any vendor you talk to should be able to answer who runs this after handover in one breath. Our answer is the one we give Embedded AI Lead clients: we plan the handover from the first month and train your people to run and extend each system as it ships, so when the retainer ends nothing goes dark. Day to day, the exception queue belongs to the operations people who already own the work, and IT’s involvement is the same as for any other system they administer rather than a new discipline.

Low maintenance also has a technical meaning you can test in a reference, which is what the system does when conditions get worse. Kioti, an IoT operations platform we built for commercial kitchens, ingests over 1,000 data points an hour in live service at 99.9% uptime, and its kitchen tablets keep working through dropped connections and reconcile when the network returns. The question to put to anyone proposing back-office AI is the same one that design answered: when the read fails or a supplier does something new, does the work land in a queue a person sees, or does it stop silently. That difference is most of what separates automation that lasts from the RPA that didn’t, and the patterns that kept recurring across our own engagements are written up in what seven UK AI audits kept finding.

The number your finance director will ask for.

A payback case for this work stands on one discipline: the baseline is measured before anything goes live. Somebody counts the hours the matching and the chasing take in a normal month, timed rather than recalled, and the saving is claimed against that count instead of against a vendor’s estimate. A system that goes live without a baseline can never prove anything, however well it runs.

We hold ourselves to the same standard on data. At FW Thorpe’s Thorlux, a UK-listed lighting group, our discovery sprint validated around 6.5 million occupancy records against the live dashboard before any recommendation was written, so the pilot specification rested on numbers both sides had checked.

That counting is the first part of an AI Audit and Transformation: a fixed fee from £10,000, agreed in writing before anything starts, delivered within four weeks, and closing with a written report and a prioritised 12 month roadmap with a cost against each line. The fee structure is on the pricing page, and if you want to see how the wider market prices this work before you compare quotes, what an AI consultancy costs in the UK sets out the ranges with our own figures on the table.

References

  1. PwC UK and Make UK. (2026). Make UK / PwC Executive Survey 2026, accessed on 31 August 2026, https://www.pwc.co.uk/industries/manufacturing/insights/make-uk-executive-survey.html
  2. Office for National Statistics. (2026). Artificial intelligence in UK businesses: 2023 to 2026, released 20 July 2026, accessed on 31 August 2026, https://www.ons.gov.uk/businessindustryandtrade/business/businessservices/articles/artificialintelligenceinukbusinesses/2023to2026
Ibrahim Mizi

Ibrahim Mizi

Co-founder & CEO · Full-Stack AI Engineer · OpenKit

Co-founded OpenKit in 2020 and runs the consultancy side end to end. Eight years of full-stack development, then production AI for SMEs and the public sector.

Does this work with Dynamics 365 and Microsoft 365?

Yes, and that combination is the normal case rather than a complication. Microsoft 365 is evidenced in delivered work; Dynamics 365 is an integration target scoped at the audit stage, where we confirm which modules and customisations are in play before any build is priced. Nothing gets replaced: the assistant reads and writes through the interfaces your systems already expose, and its entries go through the same validation as a person's.

What does this cost for a manufacturer our size?

The audit that starts every engagement is a fixed fee from £10,000. If you keep us on afterwards, the Embedded AI Lead retainer runs from £5,000 a month and renews monthly, and any build is priced in the roadmap the audit produces, so each figure arrives with the reasoning behind it. Company size moves the audit fee less than scope does: what sets it is how many workflows and systems you want in range.

Do we need a data team before any of this is worth doing?

No. The data these workflows need already exists in the ERP and the inboxes where the documents arrive, and the audit's job is to check it is reachable rather than to commission a data programme. Where a candidate's data turns out to be missing or unusable, that moves it down the ranking while the others proceed; it does not add a data science team to your payroll.

Take the question to an audit.

If this raised a question about your own operation, the AI Audit and Transformation is where we answer it. It completes within four weeks, and your first automation is live before it ends. You leave with a written report your board can read in one sitting alongside a prioritised 12 month roadmap. Your fee is fixed and agreed before anything starts.

Find your first workflow.

We start with a conversation, audit where AI actually pays back, and build the first automation into how your team already works. We reply within one working day.