AI EngineeringLLM development & fine-tuning
LLM development & fine-tuning
Models tuned to your domain, benchmarked on your baseline. When an off-the-shelf model isn’t accurate enough for your work, we tune one that is, and prove it against your own numbers.
Most tasks don’t need a tuned model, and we say so.
Tuning earns its place when the domain language is specialist, the accuracy bar is high and measurable, and the volume justifies the work: marking, extraction, review and classification at scale.
What we build.
OpenKit builds and tunes language models on your own examples, and hands over the weights, the evaluation suite and the numbers they were judged against.
- Fine-tuned models
- Trained on your examples, for your task, in your language.
- Evaluation suites
- Your baseline turned into a benchmark the model must beat.
- Prompt & pipeline engineering
- Often the honest first answer before any tuning.
- Open-weight deployments
- Tuned models you own, run where you choose.
- Continuous evaluation
- Accuracy tracked in production, not assumed.
Language models are one of four builds under AI engineering, alongside agents, retrieval systems and voice.
When general is not good enough.
A custom LLM is worth the effort when the domain is narrow, the data is sensitive, or the same language task repeats often enough that accuracy and privacy start to matter more than breadth.
- Domain language
- Legal, clinical, engineering, or financial text where a general model misreads the terminology and a tuned one reads it the way your experts do.
- Document extraction
- Pulling structured fields out of long, messy documents at accuracy a firm can rely on, with a citation back to the source line rather than a confident guess.
- Private deployment
- Data that cannot leave your walls: the model runs inside your own controlled environment, on hardware you own, so nothing is sent to an outside service.
- Drafting in house style
- First versions of routine documents written in your language and format, so a person edits rather than starts from a blank page.
- Grounded answers
- Questions answered from your own knowledge base with sources attached, usually by searching your documents first and letting a tuned model answer from what it finds.
Tuned and proven in four weeks.
- Week 1
Identify
Task defined, examples gathered
- Human baseline measured
- Week 2
Build
First tuned model
- Evaluation suite running
- Week 3
Prove
Benchmarked against the human baseline
- Error review with your experts
- Week 4
Scale
Production and monitoring
- Team trained
The smaller answer, when it exists.
Four cases where tuning is the wrong spend, and we say so before anyone commits.
- A general model already does the job. If ChatGPT or Claude handles the task well and privately enough, a custom model is money spent for no gain.
- The answers already sit in your documents. Pulling them out with sources attached is retrieval-augmented generation, which is usually cheaper to build and easier to maintain than fine-tuning.
- There is too little clean data to learn from. A handful of messy examples produces a weak custom model; a well-prompted general one beats it.
- Nobody will keep it fed. A model that is never re-evaluated drifts as your business changes: if there is no owner for it, do not build it.
If retrieval is what you actually need, start with retrieval-augmented generation instead: grounding a model in your data is often the cheaper, more maintainable route.
Where a tuning engagement sits.
- Inside an AI Audit and Transformation
-
The audit finds the task and measures the baseline; the first tuned model ships in the same engagement.
How the audit runs - With an Embedded AI Lead
-
For marking, review or extraction products: iterative tuning cycles, run inside your team.
How the embedded engagement runs
The honest range for UK LLM work.
Published UK market ranges for what a competent partner charges to take a model from scope to a working deployment. Building a frontier model from scratch costs millions and is almost never the right answer: fine-tuning or grounding an open model is where the money should go. OpenKit scopes each build to an outcome and never publishes its own prices.
- Fine-tuning & adaptation
- £20k-£80k
- Adapting an open model to your data and terminology.
- Knowledge / retrieval system
- £15k-£50k
- Grounding answers in your documents; depends on data cleanliness.
- Regulated environment
- +10-20%
- Security documentation, access controls, audit trails, evidence.
Source: OpenKit AI development cost guide, published UK market ranges.
…delivered a thorough, evidence-based strategy for our AI-assisted marking platform. We were particularly impressed by their transparent approach [and] technical expertise…
Rubrical is a DfE-backed private marking assistant that returns half of GCSE geography marking time; EMQN’s marking platform for genetic testing laboratories achieved 93-96% per-criterion accuracy.
- BAiSICS
- A custom OCR-plus-LLM pipeline for commercial-lease review that reads poor-quality documents at 96% extraction accuracy, surpassing general models on the firm’s own leases, with every output verifiable against the source.
- International Oil and Gas Service Provider
- An LLM strategy and on-premises pilot specification for pipeline-integrity audits: private retrieval over audit content, an open-weight model on the client’s own GPUs, and an immutable audit trail, written precisely enough for their team to build.
What teams ask before tuning anything.
Do we need our own model, or is GPT enough?
Usually GPT-class models with good engineering are enough: the audit says which side of the line your task sits.
How much data do we need?
Less than you’d think for fine-tuning: hundreds of good examples often beat thousands of poor ones. Week one tells us.
Whose model is it afterwards?
Yours: tuned weights on open-weight bases are handed over like any other build.
How do you measure accuracy?
Against humans doing the same task: the baseline is measured in week one and the model has to beat it before rollout.
Does it keep learning?
Not silently. Retraining is a deliberate, evaluated step: accuracy is tracked in production and tuning cycles are scheduled, not automatic.
What is a custom large language model, and why build one?
A custom large language model is one adapted to your data, terminology, and workflows rather than the general internet. You build one when a general model keeps getting your domain wrong, or when your data cannot leave your environment. OpenKit scopes custom LLM development around a specific business problem: reading your documents, drafting in your house style, or answering with your policies.
What is the difference between a custom LLM and using ChatGPT or Claude?
A general assistant is trained on public data and, on the public tiers, your prompts may leave your control. A custom LLM is adapted to your terminology and runs where you choose: on your own infrastructure the data never leaves it, and where a build uses a hosted provider we contract on terms that keep your prompts out of anyone’s training data. For narrow, repeated tasks a tuned model is more accurate and more private; for open-ended work, a general model is often the right tool, and we will say so.
Do you train a model from scratch or fine-tune an existing one?
Almost always we adapt an existing open-weight model rather than train from scratch. Training a frontier model from zero costs millions and rarely beats fine-tuning or grounding an open model on your data. OpenKit fine-tunes, prompts, and grounds open-weight models so the result stays private and your team can run and update it themselves.
What can a custom LLM do for my business?
The strongest cases are document-heavy and language-heavy: extracting fields from contracts, classifying and routing correspondence, drafting first versions of routine documents, and answering questions from your own knowledge base. It understands your context, which is where general tools fall down, and where the BAiSICS legal platform, for example, surpassed general models on the firm’s own leases.
How do you keep our data private?
The model runs where your data already lives, in a UK-region cloud or on your own servers, and client data is used only to deliver the engagement: prompts and outputs are not retained afterwards and never used to train external models. For the most sensitive work, an open-weight model on your own hardware means your documents never leave the building, the pattern OpenKit specified for an International Oil and Gas Service Provider.
Is a custom LLM secure enough for regulated data?
It can be, when built for it. OpenKit is ISO 27001 and ISO 9001 certified and holds Cyber Essentials. We add role-based access, data-residency controls, and an audit trail of every prompt and response, and we design to operate under UK GDPR. The deployment posture (on-premises or UK region) is chosen to fit your regulatory obligations, not ours.
What does large language model development cost in the UK?
Published UK market ranges put fine-tuning and adaptation work at roughly £20,000 to £80,000, with a knowledge or retrieval system from around £15,000 depending on how clean the data is (OpenKit AI development cost guide). Regulated environments add ten to twenty percent. OpenKit scopes each build against a defined outcome and never publishes its own rate card.
Can the model run on our own infrastructure?
Yes. An open-weight model can run inside your own controlled environment, on hardware you own or in a private cloud, so prompts and documents stay on your network. You keep control of the data and the running costs are predictable rather than metered per call. It takes more setting up than a hosted service, so we recommend it where the sensitivity or the volume justifies it, and a UK-region hosted deployment where it does not.
Not ready to talk? The free AI readiness check scores where you stand in about five minutes.
Find your first workflow.
We start with a conversation, audit where AI actually pays back, and build the first automation into how your team already works.
We reply within one working day.