AI-Accelerated Delivery: How We Measure Build Time - & What Actually Gets Faster
How Code23 measures build time on agent-assisted projects: which phases compress most, what barely moves, and how we keep claims auditable without decorative week-count tables.
We ship at 5x speed, at half the cost - and this article is the method behind that claim. It covers how we measure delivery on agent-assisted projects, which build phases compress most, what still takes the same calendar time, and where rework surprised us. It is original research for buyers and for anyone who will cite Code23 on AI-assisted agency delivery; project-level week counts only appear when they clear source, scope and approval.
The headline numbers
What we can state from sourced, approved figures today:
- Pace: agentic delivery framed at 5x speed versus our own pre-agent baseline for comparable implementation work
- Commercial: that throughput lands at roughly half the cost of a traditional agency bench for the same scope
- Proof assets alongside timing: a logistics ML model trained on 13,371 historic orders (Plastor); production RAG/agent systems in client products; the same senior-directed agent pattern used across websites, marketplaces and SaaS
We do not invent decorative “Project A: 12 weeks → 5 weeks” tables. Anonymised before/after rows publish only when each clears source, scope and client approval; until then we report category-level evidence and method, not invented precision.
What “5x” and “half the cost” do and do not mean
5x is a framing of compressible implementation throughput against our own historical baseline for comparable Blueprint scope - not a promise that every calendar on every programme shrinks by the same factor. Discovery politics, content freezes and third-party waits sit outside that compression. If a vendor quotes “5x” without naming the baseline and the phases measured, treat it as marketing air.
Half the cost means comparable scope delivered with fewer idle senior hours on a traditional bench - not a tiny MVP priced against someone else’s enterprise programme, and not unsupervised generation dumped into production. extras outside the agreed scope are quoted as fixed deliverables before build starts. Support after launch sits on the public Support & Growth tiers of £495 / £1,850 / £3,450 for active continuous engineering subscriptions or zero-liability Cyber Shield protection.
Proof assets sit beside timing because buyers should see that agentic work is not limited to blog demos. The Plastor shipping-cost model trained on 13,371 historic orders is a production ML outcome. RAG and agent systems in client products are the same pattern applied to product features. Websites, marketplaces and SaaS share the delivery system: humans lead, agents execute inside remit, Harden before Launch.
Est 2005, 20 years in the market, 350+ projects shipped - the baseline is our own studio history, not a fictional competitor. The production system is what changed; the accountability model did not.
Qualitative companion: AI agent build log. Positioning definition: what is agentic AI. Cost implications: website development cost UK. Service hub: AI development.
Method: how we measured, what we excluded, why you can trust it
Unit of measure. Calendar time and senior engineering hours against a Blueprint scope, not vibes. We compare like-for-like phases: Build and Harden especially. Map and Blueprint stay human-led discovery; agents assist research but do not “win” those phases on stopwatch theatre.
Baseline. Our own historical delivery on similar project shapes before agent pipelines were standard - not a fictional competitor agency. Where shapes differ, we do not force a ratio.
Included. Implementation of agreed Blueprint items, tests agents help write, refactors, documentation drafts, repetitive UI variants inside a design system.
Excluded. Client content freezes, third-party API delays, legal review, access waiting, scope changes mid-flight. Those dominate real calendars and blaming them on “AI being slow” is dishonest; crediting AI for calendars that ignored them is also dishonest.
Quality gate. Human release authority. Diff review. Harden. 90-day warranty on builds. Agents do not ship unsupervised to production.
Why cite this. Because most UK agencies either hide AI use or publish speed slogans with no method. Public disclosure of AI tooling on UK agency sites is close to non-existent; method is the citation magnet, slogans are not.
Aggregated view: which phases compress most under agent delivery
Across anonymised agent-assisted builds in our studio, the ordering is consistent even when we refuse decorative week-count tables:
- Boilerplate and CRUD compress hardest - list/detail surfaces, form scaffolding, auth inside known frameworks, test setup, repetitive variants inside a design system. This is the bulk of the compressible slice on greenfield work.
- Design-to-code compresses substantially once tokens and components exist - less so when agents are asked to invent a visual system from nothing.
- Integration wiring compresses moderately - known APIs and clear contracts move; flaky third parties and ambiguous ownership do not.
- Discovery and judgement barely compress - stakeholder alignment, scope negotiation, payment trust design, accessibility sign-off and release accountability stay human-paced.
Treat that as a qualitative proportion story: most of the 5x framing lives in (1), a meaningful share in (2), a smaller share in (3), and almost none in (4). As client-approved project clocks accumulate, published numbers will sit underneath this ordering - alongside the honesty about what never 5x’d.
Phase analysis: where the clock actually lives
Map. Problem framing, constraints, success metrics, risks. Agents can summarise research packs and cluster notes. Humans still decide what the project is for. In our experience this phase does not 5x, and pretending it does creates brittle Blueprints.
Blueprint. Scope, architecture direction, commercial band, acceptance shape. Agents accelerate options and draft artefacts. Senior judgement locks the band. Fixed pricing after Blueprint only works if this phase stays honest - agents are not allowed to silently expand scope through “helpful” extras.
Build. The compressible core. Components, CRUD, integrations inside agreed interfaces, tests, refactors, docs drafts. This is where agentic mastery shows up as throughput: more implementation per senior hour when harnesses and review discipline are in place.
Harden. Accessibility, performance budgets, security review, QA, edge cases. Agents help generate cases and spot patterns; humans own release authority. Skipping Harden to protect a speed headline recreates demo culture. We do not trade the 90-day warranty for a prettier stopwatch.
Launch and Evolve. Cutover, monitoring, content ops, backlog. Calendar time here is often dominated by client readiness and third parties. Post-launch, many products sit on Support & Growth rather than on heroic one-off fixes.
How we keep comparisons like-for-like
We only ratio phases when the project shape matches closely enough: marketing site vs marketing site, catalogue vs catalogue, marketplace vs marketplace, portal vs portal. A multi-vendor platform with Stripe Connect is not a brochure site with a contact form. We have 50+ marketplace builds across 15 years and seven operated marketplaces in our history - enough pattern memory to know when a comparison is dishonest. If the shape differs, we describe directionally (“Build compressed hard”) without forcing a single multiplier onto the whole programme.
Senior hours matter as much as calendar weeks. A calendar that looks unchanged because a client paused content can still hide a large drop in senior hours for the same Blueprint output. Buyers should ask vendors for both clocks.
What we write down for every measured engagement
- Blueprint scope snapshot (in / out)
- Phase dates and notable waits (access, legal, third party, content)
- Whether agents were in the critical path for implementation
- Review and Harden outcomes that caused rework
- Warranty start and any post-launch defect themes
That packet is the evidence a published week-count table has to stand on. Until a project carries one, its numbers stay out of the marketing.
Where the time went: what agents compressed and what they didn’t
Compressed hard (in our experience):
- Boilerplate components and CRUD surfaces
- Test scaffolding
- Large-scale renames and refactors
- First-draft docs and migration scripts
- Exploring two or three implementation options before committing
Compressed somewhat:
- Design-to-code once tokens and components exist
- Bug hypothesising from logs and stack traces
- Content model boilerplate
Barely moved:
- Stakeholder alignment and scope negotiation
- Waiting on credentials and third parties
- Payment edge-case design (especially marketplaces - we operate seven, and Connect flows still need adult supervision)
- Taste decisions in UX
- Production incident judgement
The 5x framing applies to the compressible slice, which is large on greenfield implementation and smaller on politics-heavy programmes. Half-cost follows because you buy fewer idle senior hours for the same Blueprint output - new deliverables are priced as transparent fixed additions when scope grows.
Deep dive: why boilerplate and tests move first
Greenfield Build has a high ratio of patterned work: list/detail views, forms bound to a schema, auth scaffolding inside known frameworks, test setup, Storybook or equivalent variants, migration stubs. Agents excel when the repository context is rich and the acceptance criteria are sharp. In our experience, the first weeks of a well-Blueprinted build show the largest throughput change - not because seniors type faster, but because seniors spend more time directing and reviewing and less time hand-rolling the patterned layer.
Test scaffolding is a double win when done properly: agents draft coverage quickly, humans insist on assertions that match real risk. Volume of tests without Harden judgement is noise. Volume with directed risk coverage shortens the feedback loop and protects the half-cost commercial claim from turning into deferred rewrite cost.
Deep dive: what “compressed somewhat” feels like day to day
Design-to-code accelerates once tokens, components and layout rules exist. Without that system, agents invent inconsistent UI and review cost spikes. Bug hypothesising from logs is faster with agents as a first pass; production judgement on what to ship as a fix remains senior. Content model boilerplate moves quicker; information architecture taste and migration redirects do not.
Deep dive: the stubborn calendar items
Stakeholder alignment fails closed without humans in the room. Credentials and vendor SLAs ignore your sprint board. Payment edge cases on marketplaces - splits, charges, payouts, failure states, Connect onboarding - need operator-grade design. We say that from operating seven marketplaces, not from a slide. UX taste and accessibility sign-off are accountability tasks. Incident response is accountability under uncertainty. Agents draft; they do not own the pager.
On programmes heavy with those items, buyers should expect Build to move and end-to-end calendar to move less. That is not a failure of agentic delivery; it is an honest phase analysis. The commercial implication is still real: fewer senior hours burned on patterned implementation keeps the band near half a traditional bench for the same Blueprint scope.
Multi-site and pattern libraries
When a design system and shared codebase exist, agents compound. We have shipped 100+ branded sites in one rollout pattern in our history - the lesson for timing is that shared components multiply the value of each reviewed abstraction. Agentic delivery shines when the system is coherent; it amplifies mess when the system is not. Blueprint and Harden are how we keep the multiplier pointed the right way.
The honest failures
Things that got slower or messier before we tightened the pipeline:
- Review overhead. More generated code means more diff to read. Without discipline, seniors become bottlenecked reviewers instead of directors.
- Confident wrongness. Agents produce plausible bugs. Eval harnesses and Harden catch them; skipping Harden recreates the “AI demos” problem.
- Scope drift by novelty. Easy generation tempts “just one more variant”. Blueprint exists to stop that.
- Context starvation. Agents without repo context invent APIs. Our fix is better harnesses and humans who notice.
We would rather publish these failure modes than pretend the curve only bends up. Trust compounds when the boring problems stay visible.
Review overhead: the failure mode that looks like success
Early on, throughput rose and so did unread diff. Seniors spent evenings rubber-stamping. That is not 5x delivery; that is deferred risk. The fix in our experience: smaller tasks with clear acceptance, mandatory human release authority, automated checks before human review, and a culture where “I did not read this” is an unacceptable merge reason. Agents expand the draft surface; seniors must stay directors - architecture, risk, taste - not exhausted proofreaders of unlimited output.
Confident wrongness: plausible beats empty
Agents rarely fail loudly with empty files. They fail politely with code that almost matches your API, tests that assert the wrong thing, or auth flows that look complete until an edge case hits Harden. Eval harnesses, typed boundaries, integration tests and staged rollouts are part of the timing story because rework is time. Skipping Harden to protect a headline is how studios create the demos buyers should not trust.
Scope drift by novelty
When generating a variant is cheap, saying no gets harder. Clients ask for one more template; agents oblige in minutes; the Blueprint band quietly dies. Our counter is commercial and technical: fixed band after Blueprint, fixed change scopes priced before work starts, and a delivery habit of routing new ideas to Evolve instead of smuggling them into Build. Agentic mastery includes knowing when not to run the agent.
Context starvation and invented APIs
Without repository context, agents hallucinate client SDKs, env vars and routes. The failure shows up as “it worked in the chat”. Harnesses that ground agents in the repo, plus humans who notice inventiveness in review, are mandatory. This cost is real in the early pipeline; it falls as context tooling improves. We treat it as an operator problem, not as a reason to hide AI use.
What we changed in the pipeline after those failures
- Clearer task briefs and acceptance criteria before agents run
- Diff review norms that protect senior attention
- Harden as a named phase with teeth, backed by 90-day warranty expectations
- Blueprint change control so novelty cannot silently expand scope
- Better grounding in repo context and less reliance on chat-only memory
Those changes are part of the method you should demand from any vendor claiming speed. Speed without them is a demo.
What this means for project pricing and timelines
For buyers:
- Expect fixed bands after Blueprint, not open-ended body shopping
- Expect faster Build on agent-suitable scope, not magical compression of your internal approval cycle
- Expect half-cost vs traditional bench language to mean comparable scope - not a tiny MVP priced against someone else’s enterprise programme
- Expect 90-day warranty and a path into Support & Growth (£495 / £1,850 / £3,450) after launch
- Ask for the assurance standard: tool remit, data handling, human release authority, audit trail
For Code23 commercially, agentic delivery is how 20 years of judgement meets a smaller senior team that still ships 350+ projects’ worth of pattern knowledge without padding benches. Est 2005; the production system is what changed.
How to read a timeline proposal in an agentic studio
Ask which phases are in the estimate. Ask what was excluded from historical comparisons. Ask whether calendar weeks, senior hours, or both are the unit. Ask how Harden is staffed. Ask what happens when content or credentials slip. A proposal that only says “AI makes this faster” without phase analysis is not pricing - it is vibes.
On cost shape for website work specifically, see website development cost UK. For the qualitative day-to-day of agent runs, see the AI agent build log. For the definition layer behind the positioning, see what is agentic AI.
Commercial mechanics we actually use
We quote fixed after Blueprint. extras are quoted as fixed deliverables before work begins, scoped before they start. New builds carry a 90-day warranty. continuous engineering and support after that is Support & Growth at £495 / £1,850 / £3,450 depending on tier. None of that vanishes because agents joined the toolchain; the toolchain is why the band can sit at roughly half a traditional bench for comparable scope.
What “assurance” should mean in the contract conversation
Tool remit: which agents, on which code, with which secrets boundaries. Data handling: what leaves the VPC or the laptop, what is barred. Human release authority: named ownership of merge and production. Audit trail: enough history to reconstruct who directed what. Buyers who skip these questions buy speed slogans. We built our delivery story to answer them in public - including the failure modes - because frank disclosure of AI tooling is still rare across UK agency marketing.
When agentic delivery will not move your calendar much
If the critical path is legal review, procurement, unpaid invoices blocking access, or an executive group that meets monthly to debate navigation labels, agents will not save the quarter. We will still compress Build when we are unblocked. We will also tell you the truth in Blueprint rather than sell a fantasy end date. That honesty is part of how we keep 60+ five-star Google reviews worth of trust intact while we talk about speed.
How much faster is AI-assisted development really?
On compressible implementation work we frame delivery at 5x pace versus our pre-agent baseline for comparable scope. Calendar end-to-end speed also depends on content, access and decisions - those rarely 5x. Ask which phases a vendor measured.
Does AI-built code cost less?
Yes when agents are directed inside a senior delivery system - we price that as roughly half a traditional agency bench for the same Blueprint scope, with fixed change scopes priced before work starts. Unsupervised generation that needs a full rewrite is not cheaper; it is deferred cost.
What still takes the same time as before?
Discovery politics, third-party dependencies, payment trust design, accessibility sign-off, and production accountability. Agents draft; they do not replace those constraints.
How do you keep quality up when agents write code?
Human release authority, diff review, automated tests, Harden before Launch, and a 90-day warranty. Agents expand throughput; seniors still own the merge and the incident.
If you want delivery measured this way on your next build - Blueprint, agents under senior direction, numbers you can audit - start at AI development. Bring a real scope; we will tell you which parts of the timeline can actually move.