Can an AI Agent Do QA? Our Overnight Audit Pipeline
We run AI agents through client sites overnight: what they catch that manual QA can miss, what they still get wrong, and how human review keeps the bar up.
Can an AI agent do QA? Yes, and we run agent QA overnight. Code23 sends AI agents through studio sites and the client sites we support while everyone’s asleep: walking real journeys, capturing screenshots and console notes, and building a report ready for a person to read over coffee. What it can’t do is decide what ships. An agent can flag a broken form or a missing label. It isn’t reliable at judging whether a noisy page still converts, or whether an auth wall is there on purpose. That judgement stays with a person, every time.
We started running this pipeline because manual QA runs out of hours before it runs out of pages. Overnight coverage buys back the corners a tired afternoon pass tends to skip: legal pages, empty states, the locale nobody got round to. It also produces its share of false positives, so a person still triages every report before anything reaches a sprint. This piece covers how the pipeline works, what it catches, where it gets it wrong, and how it fits into Support & Growth.
What the overnight pipeline actually does
Can AI do QA testing?
AI can perform structured exploratory and checklist QA - walking flows, checking states, flagging accessibility and content issues - but we keep production approval with a person.
What can an AI agent test on a website?
An agent can test primary journeys, form validation, broken links and assets, responsive breakpoints, basic accessibility smells, console errors, and charter-defined edge paths - especially overnight when humans are offline.
Our pipeline in plain language:
- Charters. Written missions (“complete enquiry on mobile”, “add to basket as guest”, “check footer legal links on key templates”) - not vibes.
- Walkthroughs. Agents navigate the live or staging site, following the charter and branching into likely failure states.
- Evidence capture. Screenshots, short clips, URLs, console notes - we treat a finding without evidence as a rumour.
- Structured report. Issues grouped by journey and template, ready for triage.
- Human review next morning. Severity, duplicates, taste, and “ship / defer / ignore”.
That is the same senior-directed agent pattern as what is agentic AI and the build-time discipline in our agent build log: in our pipeline, agents execute and people decide. We describe capability here, not internal tooling names.
What agents catch that manual passes miss
Does AI QA replace human testers?
No. AI QA widens coverage and stamina; humans still own prioritisation, domain judgement and whether a release ships.
In our experience overnight runs surface classes of issues manual afternoon passes under-sample:
- Template dark corners. Legal pages, empty states, logged-out vs logged-in variants, and “rare” locales that are easy to skip late in the day.
- Regression drift. A component change that breaks a secondary template the feature owner didn’t check.
- Form edge paths. Wrong keyboards, validation order, double-submit, and error copy that only appears on one field combination.
- Accessibility smells. Missing labels and many contrast failures are cheap to catch mechanically; focus traps need a keyboard walk-through, which an agent can do on every run and a person hunting a visual bug often skips.
- Console and network noise. Failures that don’t always show on screen but can corrupt analytics or break payments.
Agents can cover the breadth people skip when time is short; people catch the meaning agents get wrong. We don’t publish issue counts from client sites without the client’s approval.
When those findings feed a productised review, they land in a UX audit fix list rather than a Slack dump.
What agents still get wrong
False confidence is the main failure mode.
False positives. Agents can flag “broken” states that are intentional (auth walls, geo blocks, feature flags). Triage time is real cost.
Context blindness. A visually noisy page may convert fine; a clean page may hide a trust problem only a practitioner spots. Taste is not a checkbox.
Flaky environments. Staging data resets mid-run; third-party widgets time out; the agent retries and writes a novel. In our experience, harness discipline matters more than model fashion.
Encoded bugs in tests. Agents sometimes assert the broken behaviour as correct - we learned that the hard way in build review; a person needs to check test intent, not only pass/fail.
Not a security test. Automated clicking isn’t a penetration test. Don’t market overnight QA as a security audit.
Mastery is knowing when to distrust a tidy report.
The human layer: severity, triage and approval
After each run, usually the next morning, a person:
- Deduplicates and merges template-level causes
- Scores severity by user and business impact, reach, reproducibility and the critical journey affected
- Separates ship-this-week patches from redesign-class issues
- Assigns owners and release windows
- Decides whether the run is “clean enough” for the agreed bar
Without that layer, you’re likely to end up with an expensive screenshot generator. With it, overnight coverage can make a small team’s testing go much further - especially when paired with the judgement model in a UX audit.
We don’t hand release authority to the agent. If a finding is contested, a senior decides, on client work and on our own properties.
What this means for support clients
How much does automated QA cost?
Tool licences are the cheap part. Published prices are in US dollars: Cypress Cloud’s Team plan is $67 a month billed annually, BrowserStack Automate starts at $59 a month for one parallel test, and BrowserStack’s AI-agent Test & Monitor plan starts at $1,099 a month billed annually (checked September 2026). What costs more is the human time to write charters and triage findings, which is why we include it in retained support rather than selling it as a tool.
For most of our clients, overnight agent QA isn’t a separate product. It’s part of how we run Support & Growth:
| Tier | Monthly | QA-relevant reality |
|---|---|---|
| Cyber Shield & SLA | £495 | Monitoring, security posture, investigation route; agent assists triage where useful |
| Unlimited Growth | £1,850 | Continuous development capacity - room for scheduled charter runs alongside product work |
| Pro Pod / Scale | £3,450 | Higher change volume and tighter ops - two concurrent tasks, private Slack, four-hour priority, weekly architecture syncs and monthly tech-debt reviews |
Extras outside the tier are priced before we do them. On build projects we test failure paths before launch; the overnight pipeline is how we keep catching drift after go-live instead of waiting for a customer tweet. For agentic systems work, see AI development; for retained care, see Support & Growth.
The lesson is boring: coverage compounds when someone owns triage. Agents give you the overnight shift; they don’t give you accountability.
"We have used Code23 for many of our website projects and they have been absolutely fantastic every time. They ensure that meetings are held to discuss the particular project needs, they set clear timelines and stick to them, they are always on hand for client support...their work is outstanding!"
Frequently asked questions
Is AI QA testing safe to run against a live production site?
It can be, if the run is scoped first. Anything that submits a form, creates an account or goes near checkout creates real records on a live site, so those journeys run on staging or with test accounts you’ve agreed. Screenshots and console notes can also capture personal data, so before the first run we agree with you which journeys are covered, what gets recorded and where the evidence is kept.
How is AI QA different from a scripted testing tool like Selenium or Cypress?
A scripted test checks the steps someone wrote for it, and it’s the right tool for regression checks you want to repeat exactly. Tools like Cypress and Playwright now add AI help for writing and repairing those tests, but they still run from a test plan. Our overnight agents work from a charter and explore around it, branching into likely failure states nobody wrote a step for, and a person triages what they find. They’re not a replacement for your existing test suite. They cover ground a script usually doesn’t know to look for.
Can AI QA catch accessibility problems?
It catches the mechanical smells a scanner would flag, such as missing labels and many contrast failures, and an agent that tabs through a journey can also spot some focus traps that scanners miss. It won’t tell you whether a journey actually feels usable to someone relying on assistive technology. Treat agent findings as the first pass, not the accessibility audit itself.
Do I need to be on Support & Growth to get overnight QA?
Usually, yes. For most clients overnight agent QA is part of how we run Support & Growth rather than a separate product, and how often it runs depends on your tier and how much your site changes. If you’re not with us for support yet, ask and we’ll show you where it fits.
Next step
Overnight AI QA only earns its keep when someone owns triage the next morning. That’s the part we build in, not the part we skip.
If you want that coverage on your own site or product, it usually sits inside Support & Growth rather than as a separate line item, and we agree how often the runs happen and what they cover with you before they start. We’ve delivered 350+ projects since 2005. This pipeline runs on our own studio work as well as on client sites.
Build with AI
Ship the product faster without cutting the quality.
Code23 combines senior engineering, AI-assisted delivery and proper testing to move complex products from idea to release.