Agentic delivery 11 min read

What Does an AI Development Company Actually Do?

What AI development companies really build, what to expect from an engagement, and how to judge one. From a UK agency that ships AI to production, with proof.

What Does an AI Development Company Actually Do?

An AI development company builds and ships software that uses models in production - assistants, retrieval systems, automation pipelines, trained predictors, and full AI-native products - under human architecture and release authority. You hire one when the outcome is working software with owners, not a slide deck of use cases. At Code23 that means seniors direct the work, agents accelerate implementation, and we only count a job done when it runs against real data with monitoring and a path for failure.

The short answer: what an AI development company builds

Definition we work to:

An AI development company turns a named business problem into production software that calls models or agents inside clear boundaries - with evaluation, logging, handover and someone accountable when the system is wrong.

Typical surfaces:

  • Workflow automation that removes repetitive hours
  • RAG and assistant products grounded in your documents
  • Classical or modern ML models trained on your history
  • AI-native product features inside a marketplace, SaaS or ops tool
  • Evaluation harnesses, review gates and ops runbooks

That is different from buying a model licence, hiring a slide-only adviser, or bolting a homepage chatbot onto an unchanged process. Service home: AI development.

We ship this pattern across client work and our own products. Proof artefacts include a logistics ML model trained on 13,371 historic orders (Plastor), a UK document-comparison platform we shipped - diligence-grade RAG over private materials using the Claude API, and a multi-agent content pipeline we run internally. Method and timing live in the build-time dataset; the qualitative trail is in the AI agent build log.

The four project shapes

Most serious AI builds we see fall into four shapes. The label on the proposal matters less than which shape you are actually buying.

1. Automation

Remove repetitive human hours from a named process: quoting, triage, document intake, reporting packs, CRM hygiene. Success looks like hours returned every week, fewer errors, and a clear exception path for humans. This is the commercial cousin of what an AI automation agency does - outcome-first, process-owned.

2. RAG and assistants

Document-heavy work where answers must cite sources. Contracts, diligence packs, knowledge bases, compare tools. Retrieval pulls the right chunks; the model drafts; humans still own release when the stake is high. We walk the pattern in what RAG is and when you need it. A UK document-comparison platform we shipped sits here: Claude API plus retrieval over private material, not a generic chat toy.

3. ML models

Predict or score from your own history. Plastor’s shipping-cost model trained on 13,371 historic orders is the worked example: production ML against real order history, evaluated against the incumbent process, not a demo notebook. You need clean enough data, a metric that matters, and a plan for drift.

4. AI-native product builds

The product is the AI surface - marketplace matching aids, SaaS copilots, agent-assisted ops inside a platform your users pay for. Same delivery discipline as any other product build, with model evaluation treated as part of Harden. Agentic delivery (humans lead, agents execute) is how we implement these at speed - definition in what is agentic AI.

ShapeYou are buyingKill signal if missing
AutomationHours deleted, exception pathNo named process owner
RAG / assistantsCited answers over private docsNo corpus worth retrieving
ML modelPredictions that beat baselineNo labelled history
AI-native productA feature users pay forNo product owner or release bar

If a vendor collapses all four into “AI transformation”, ask which row you are on before you sign.

Worked examples from our studio

Plastor - ML on real order history. A shipping-cost model trained on 13,371 historic orders. The point was not a notebook accuracy chart for a slide. It was a predictor that had to stand next to the incumbent quoting process with evaluation and an owner. That is the ML shape: labelled history, a baseline to beat, and a plan when predictions drift.

Diligence-grade document comparison - RAG over private documents. Claude API plus retrieval over private documents, with citations and human release authority on high-stakes interpretation. That is the RAG shape: corpus, retriever, generator, logging, gates. Detail in what RAG is.

Multi-agent content pipeline we run internally. A multi-agent pipeline under editorial rules. That is automation-plus-agents inside our own ops: humans set remit and publish authority; agents draft and transform at volume. Same discipline we sell to clients.

Client delivery itself. Websites, marketplaces and SaaS built with seniors directing coding agents - the build log and build-time data are the receipts. AI development companies worth hiring should show how they use AI in their own delivery without hiding the review gates.

What an engagement actually looks like week by week

Shapes differ; the week rhythm is recognisable once you have run a few.

Early weeks - frame and readiness. Name the system that will change and the metric that will move. Check data access, permissions, failure modes, and who owns exceptions. Kill soft ideas here. A discovery that never says “no” was sales.

Middle weeks - thin working slice. One path end to end: ingest, model or retrieval call, UI or API, logging, human review gate. Prefer a boring production path over a flashy demo that cannot survive your auth.

Later weeks - Harden and handover. Evaluation set, monitoring, runbooks, access review, and who to call when quality drifts. Agents can draft tests and refactor; seniors own architecture and release. That is agentic mastery, not “the AI shipped it”.

After launch. Optional care on a retainer if you want someone watching the boring parts. Extras outside a fixed band sit on transparent fixed terms on our commercial terms. We do not pretend a workshop alone is a production programme.

Calendar length tracks politics and data readiness more than model fashion. Clean data and a single owner compress everything. Multi-stakeholder programmes and regulated domains stretch everything. The build-time data explains which phases compress under agent-assisted delivery and which still take calendar time.

What “done” means (and what a demo is not)

A demo answers: can the model sound useful on a happy path? Production answers: does it work on messy inputs, with permissions, logging, cost controls, and a human path when it fails?

Minimum bar we use before calling a build shipped:

  • Named owner in the client organisation
  • Evaluation set that includes awkward cases, not only brochure examples
  • Logs you can actually read when something goes wrong
  • Access control that matches how the business already thinks about data
  • A rollback or disable switch when quality drops
  • Documentation a new starter can follow without a tribal call

If a proposal cannot point to those, you are buying theatre with a model logo on it. Est. 2005 and 350+ projects taught us that lesson on ordinary software long before LLMs; AI simply makes the theatre prettier.

Data readiness without the fog

Most AI projects stall on data, not on model choice. Before you argue GPT-versus-Claude in a steering pack, answer:

  • Where does the source of truth live today?
  • Who can grant access without a three-week ticket loop?
  • How dirty is it - duplicates, missing fields, conflicting PDFs?
  • What must never leave the building?
  • What does a correct answer look like in numbers or citations?

When those answers are soft, the right engagement is a readiness slice, not a six-month “transformation”. We would rather kill a bad idea in week two than decorate it for quarter-end.

The labels decoded: development company vs consultancy vs automation agency

Buyers use these words interchangeably. The cheque shapes are not interchangeable.

AI development company. Outcome is software in production. Architecture, build, evaluation, Harden, handover. Our default home.

AI consultancy. Outcome is a decision, a thin proof, or change guidance - sometimes with light build. Day rates and short programmes dominate. Full cost picture: AI consultant cost.

AI automation agency. Outcome is a named workflow with fewer human hours every week. Overlaps with development when the automation is custom software; overlaps with tooling when Zapier-class glue is enough. See AI automation agency.

Agentic AI (delivery pattern). Not a company type. It is how implementation gets done: multi-step agents under human remit. We use it across websites, marketplaces, SaaS and AI features.

Rough rule:

  • Need a decision and a kill/go → consultancy sprint
  • Need production software with owners → development company
  • Need ops hours deleted weekly → automation-shaped engagement
  • Need the definition of the delivery pattern → agentic AI

Many UK searches for “AI consultancy” are really build briefs. That mismatch is why day-rate-only articles mislead.

Pricing shapes you will see (market orientation)

UK buyers usually meet one of four commercial shapes:

ShapeWhat you hearWatch for
Day-rate advisoryWorkshops, architecture reviewsEnds without software
Fixed discovery + prototypeA thin slice and a go/no-goPrototype that cannot survive auth
Fixed build bandProduction scope with extras rateScope written so loosely extras eat the saving
Retainer / managed AI opscontinuous engineering and support after launchNo monitoring or evaluation in the retainer

Market day rates and project bands for advisory-led work are unpacked in AI consultant cost. Our build-shaped work uses a fixed band after Blueprint and transparent fixed change bands extras, with agentic delivery framed at roughly half a traditional bench cost for comparable implementation. Support tiers of £495 / £1,850 / £3,450 exist when you want continuous engineering and support - they are optional, not a tax on every pilot.

Do not confuse a model API invoice with an AI development company invoice. Tokens are cheap compared with permissions, evaluation and the boring admin that makes staff trust the tool on a Monday morning.

How to judge one (proof of production, not demos)

“Top 10” listicles will not save you. Judge on production proof and disclosure.

What company is leading AI development?

“Leading” splits into two different games. Model labs (OpenAI, Anthropic, Google DeepMind and peers) lead foundation research and frontier models. Delivery companies lead at shipping business software on top of those models. Confusing the two is how buyers end up comparing a research org to a Reading agency. For a UK buyer hiring someone to ship, leadership means production systems you can inspect - not who released the latest base model.

What are the top 10 AI development companies?

There is no shared audit behind most “top 10” roundups. Directories, sponsorships, affiliate relationships and editorial packages assemble many of them. Use a list as a lead list, then verify: who merges production code, what evaluation exists, where humans stay accountable, and whether the firm publishes failure notes or only hero demos.

What are the top AI development companies in the UK?

Same mechanics, smaller pond. Prefer firms that show UK delivery proof, clear entity facts, and honest AI-use disclosure over badge collections. Ask for artefacts: model cards, evaluation sets, build logs, monitoring plans. Ask what stayed slow. Fantasists claim everything 10x’d; practitioners name discovery politics, data waits and third-party delays.

Who are the big 4 of AI?

In consulting shorthand, “Big 4” still means the large professional-services firms (Deloitte, PwC, EY, KPMG) and their AI practices - advisory-heavy, enterprise-process-heavy, often partnering for build. In model shorthand, people sometimes mean the frontier labs. Neither list is a substitute for checking whether your problem needs a board pack, a regulated programme, or a focused production slice. Match firm shape to problem shape.

Buyer checks we recommend on every shortlist:

  1. Who merges to production?
  2. Show a review trail - diffs, failed tests, release authority.
  3. Show a proof artefact from a live system, not only a workshop deck.
  4. Where do humans remain accountable when the model is confidently wrong?
  5. Who owns prompts, logs, fine-tunes and derived datasets?
  6. Separated marketing toys from client production?
  7. What stayed slow on the last project - and can they name it without flinching?
  8. How do they use AI in their own delivery, with gates visible?
  9. What happens in month two when the model drifts or the corpus changes?
  10. Can they explain the project shape in the four-row table above without buzzword soup?

The disclosure test in practice

Ask the firm to describe, in writing, where agents draft, where humans approve, and who is on the hook if the system refunds the wrong customer or cites the wrong clause. Vague answers (“we use AI throughout”) fail. Specific answers with remit boundaries pass - even if you then choose a different vendor.

Serious buyers should demand public method artefacts when shortlisting: build logs, timing method, and plain definitions such as agentic AI. Apply that standard to every firm in the RFP, including us.

We have shipped since 2005 across 350+ projects, including the production AI assets named above. Still apply the same checks to us. Disclosure is the product.

Shortlist the firm that can name where agents draft, where humans approve, and who owns a wrong refund. When you want that conversation with us, open AI development with a named process and the metric you would defend in a board meeting.

Related

More from the blog

Engineering deep-dives, product updates, and notes from the team.

View all posts