Skip to content

From AI agent pilot to production: the operating checklist behind the firms that scale

Most AI agent pilots stall between demo and production. The owners, cost tracking, logs, exception queues and kill criteria that separate firms that scale.

By Published Updated 12 min read
AI agents in operations, 12 min read — A small glowing glass block sits in front of a long, orderly row of larger glass blocks, each lit by a faint ring of blue light.

The short answer

An AI agent reaches production when it has an operating model, not when the demo works. Before go-live, name one owner, agree a baseline and acceptance criteria, track cost per task, log every action, route exceptions to people, set a review cadence and write kill criteria. Surveys from Gartner, McKinsey and KPMG suggest most firms skip several.

Key takeaways

  • In Gartner's survey of 1,303 leaders at organisations with revenue of $50 million or more (January–April 2026), only 22% had scaled AI across multiple business units or adopted an AI-first approach.
  • In McKinsey's survey of 1,719 respondents (May–June 2026), 37% attributed at least some EBIT impact to AI, about the same share as a year earlier.
  • In KPMG's survey of 204 US leaders at $1 billion-plus firms (April–May 2026), 53% were deploying AI agents but only 26% had full, real-time visibility of AI operating costs.
  • McKinsey's high performers define how they measure impact and decide in advance when outputs need a person to validate them.
  • Write kill criteria before go-live: the exception rate, cost per task or error rate at which the agent is paused, and who decides.
  • A 30-60-90-day plan moves one agent from shadow mode to supervised production to a signed result.

Why most agent pilots stall in the same place

Most agent pilots stall at the handover from the team that built them to the team that has to run them. The demo works, a few users like it, and then nobody owns the exceptions, nobody knows what each task costs and nobody can say whether it paid. The pilot keeps running without a decision.

The survey data points the same way. In Gartner's survey of 1,303 leaders at organisations with revenue of $50 million or more, fielded January–April 2026 (geography not stated), only 22% had scaled AI across multiple business units or adopted an AI-first approach.1 We looked at what the other four in five are missing in our note on the Gartner survey.

Value is just as scarce. In McKinsey's survey of 1,719 respondents in 97 countries (4 May–8 June 2026), 37% attributed at least some EBIT impact to AI, about the same share as the year before.2 Use keeps spreading while the share reporting profit stays flat.

Our view: the gap between pilot and production is mostly operational, not technical. The model is rarely what stops an agent. What stops it is the missing owner, the unknown cost and the absence of a rule for when to switch it off. This post sets out the checklist we use at Sigzen AI to close that gap, and a 30-60-90-day plan to run it.

AI in large organisations, 2026

Plenty of agents, few scaled results

22%have scaled AI across multiple business unitsGartner, Jan–Apr 2026 1
37%attribute any EBIT impact to AIMcKinsey, May–Jun 2026 2
26%of large US firms see AI operating costs in full, in real timeKPMG, Apr–May 2026 4
Sources: Gartner, September 2026; McKinsey, The State of AI 2026; KPMG, AI Quarterly Pulse Q2 2026.

What do the firms that scale do differently?

They decide how they will measure an agent and when a person checks it before they scale it. McKinsey identifies a small group of high performers, 6% of respondents in the 2026 survey.2 These organisations report defined processes to measure the impact of their AI initiatives. They have also determined how and when model outputs need a human in the loop.2 Nearly three-quarters of them report fundamentally redesigning workflows because of AI.

Size helps, but it is not the whole story. In the same survey, 40% of respondents from organisations with revenue of more than $1 billion reported scaling AI agents, up from 27% a year earlier, while the share at smaller organisations stayed flat at 22%.3 Larger firms tend to have platform teams and named owners. A smaller firm can copy the operating habits without the headcount.

Gartner's survey adds a measurement angle. Its high performers reported positive returns on 81% of their AI initiatives, while low performers did not know the rate of return for 29% of theirs.1 Gartner's recommendation is to tie AI spending directly to business outcomes and track it by outcome category. That is a reporting habit, not a technology.

Our view: each of these findings is a correlation in survey data, not proof that the habits cause the results. But they are cheap habits, and the downside of adopting them is small.

The operating checklist

Seven items separate a supervised production agent from a long-running pilot. Each one is a document or a setting, and each one needs a name next to it. If you cannot fill one in, the agent is not ready to leave the pilot.

Seven items before an agent leaves the pilot

  1. One owner

    A named person accountable for results and exceptions.

  2. Baseline and acceptance criteria

    Today's cost, time and error rate, and the targets, signed before go-live.

  3. Cost per task

    Model, platform and people time, tracked weekly.

  4. Action log

    Every input, decision, tool call and approval, kept where auditors can read it.

  5. Exception queue

    Items the agent cannot finish go to people with a reason attached.

  6. Review cadence

    Weekly at first, then monthly, with the same report each time.

  7. Kill criteria

    The thresholds at which the agent is paused, and who decides.

Illustrative. Our checklist; the order follows how we run an automation sprint.

One owner, with authority

The owner is the person who signs the acceptance criteria before the build and the measured result after it. They also decide what happens to exceptions. In most firms that is the head of the team whose work the agent does, not someone in IT. IT owns the platform; the business owns the outcome.

Give the owner real authority over the agent's permissions. If they cannot pause it, change its scope or ask for its access to be cut, ownership is decorative.

A baseline and acceptance criteria

Measure the process before the agent touches it: volume, time per item, cost per item, exception rate and error rate. Write the targets next to them and have the owner sign. Without that, the final report becomes an argument about what "before" looked like.

Keep the criteria few and countable. "Exceptions under one in six, errors caught before posting under one in two hundred, and hours saved signed off by the finance lead" is a good set. "Improves efficiency" is not a criterion.

Cost per task, tracked weekly

Most organisations cannot see what their AI costs. In KPMG's Q2 2026 pulse survey of 204 US leaders at firms with revenue of $1 billion or more (28 April–25 May 2026), only 26% had full, real-time visibility of AI operating costs.4 In the same survey, 66% had monitoring dashboards and 36% had direct token or usage controls.4

The gap is wider further down. Gartner found that roughly 11% of organisations were entirely unaware of what their function spent on AI in 2025 (survey fielded January–April 2026).1 If the function cannot see its total, nobody can see the cost of one task.

Track cost per completed task, not monthly spend. Include model usage, platform fees and the people time spent on exceptions and reviews. Our breakdown of monthly running costs shows how those parts add up, and why exception handling is often the largest.

An action log

Every action the agent takes should be recorded with its inputs, the decision, the tools it called and who approved the result. Finance and compliance will ask for this the first time something goes wrong, and rebuilding it after the fact is close to impossible. Our audit log field list sets out what to record.

An exception queue staffed by people

Exceptions are where an agent's value is decided. Every item the agent cannot finish under its rules should land in a queue with the agent's reason attached, owned by named people with time set aside for it. If exceptions pile up in an inbox nobody watches, the agent has only moved the work.

Design the handover deliberately. Our piece on when a support agent should hand over to a person applies to back-office agents too: clear triggers, context passed with the item and no dead ends.

A narrow bridge of blue light spans a dark gap between a small glowing platform and a larger structure of translucent glass panels, with a few points of light moving across it.
The step from pilot to production is a short bridge, built from owners, logs and rules rather than from a better model.

A review cadence

Review weekly for the first two months and monthly after that, with the same one-page report each time. The report should show volume handled, exception rate, error rate, cost per task and hours saved against the baseline. When the numbers drift, the owner decides whether to retrain, rescope or pause.

Models and prompts change underneath you, so treat the review as maintenance, not a formality. A vendor model update can shift an agent's behaviour overnight.

Line in the reportHow it is countedWho acts on it
Volume handledItems the agent completed or drafted this periodOwner
Exception rateItems sent to the queue, per hundred handledOwner and queue lead
Error rateDrafts corrected by a person, per hundred approvedOwner
Cost per taskModel, platform and people time, per completed itemFinance
Hours savedBaseline time per item less current time, times volumeOwner signs
Changes this periodPrompt, model, permission or scope changes, datedTechnical lead

Illustrative. The one-page report we use; keep the same lines every time so trends are visible.

The last line is the one most reports leave out. When a number moves, the first question is what changed, and a dated change list answers it in seconds. Without it, a team can spend a week debating whether a jump in exceptions came from the supplier mix, a prompt edit or a model update.

Kill criteria

Kill criteria are the thresholds at which the agent is paused, written down before go-live. They might be an exception rate above an agreed level for two weeks, a cost per task above the manual baseline, or any error that reaches a customer or a ledger. They also name who decides.

Our view: kill criteria are the item teams most often skip, and the one that most protects the programme. A pilot with no exit rule never ends. It keeps consuming budget and goodwill until someone senior cancels the whole programme, including the parts that worked.

Agent counts are not results

Counting agents tells you about activity, not value. Microsoft said on its earnings call of 29 July 2026 that nearly 40 million agents were registered on its platform, a count of registered agents, not active ones.5 We explained why the two differ in our note on that figure.

Survey counts move the same way. In KPMG's Q2 survey, 53% of leaders said their organisations were deploying AI agents, down from 55% the quarter before, while the share orchestrating multiple agents across workflows rose from 9% to 18%.4 More complex set-ups make the operating checklist more important, not less, because a failure in one agent can pass to the next.

Results are harder to find. Gartner's analysis of customer-service AI use cases, reported by CX Dive in August 2026, found that only a minority showed a positive return and that the largest group could not show value at all.6 Our note on making customer-service AI measurable covers what that means for support teams. The common thread is that value was never defined in a form anyone could count.

Pilot habits that block production

Some habits that help a pilot get started work against it later. Recognising them early saves a rebuild.

Pilot habits and their production replacements

In the pilot

  • Builder runs it and fixes it by hand
  • Success judged by user enthusiasm
  • Cost absorbed in an innovation budget
  • Exceptions handled ad hoc by the builder
  • Broad access to make the demo work

In production

  • Business owner runs it; builder supports
  • Success judged against a signed baseline
  • Cost per task charged to the process
  • Exceptions in a staffed queue with reasons
  • Least-privilege access, drafts not actions
Illustrative. Our judgement from scoping automation work; not survey data.

Access is the habit that bites hardest. Pilots often run with an admin account because it is quick. In production the agent should hold the narrowest rights that let it do its job, usually creating drafts rather than submitting them. Our least-privilege checklist covers how to set that up in an ERP or CRM.

A 30-60-90-day plan

Ninety days is enough to move one agent from pilot to a signed result if the process was well chosen. If you are still choosing, start with our scoring model for the first process. The plan below assumes the pilot already works on real data.

From pilot to signed result in 90 days

  1. Name the owner, measure the baseline, sign acceptance and kill criteria, cut access to least privilege, switch on the log. Run in shadow mode: the agent drafts, people do the work as before.

  2. Supervised production: people approve every draft. Staff the exception queue. Weekly review of exceptions, errors and cost per task.

  3. Approve by sample where the error rate allows. Monthly review. At day 90 the owner signs the result against the baseline, or applies the kill criteria.

Illustrative. Our plan for one workflow; regulated processes may keep full approval longer.

Shadow mode is the step most teams want to skip. It is also the cheapest month of evidence you will ever collect, because the agent's errors cost nothing while the people still do the work. Compare its drafts with what people actually did, and you have your first error rate before anything is at risk.

Our earlier piece on how long a pilot should run and what it should cost covers the stage before this one. The two together give you a route from first idea to a production decision.

Who runs the agent after day 90?

Someone has to keep running the checklist after the launch: the weekly report, the model and prompt updates, the cost watch and the exception queue. Larger firms give that to a platform team. Smaller firms usually give it to the business owner plus a part-time technical person, or buy it in.

On our published ladder (see pricing), the automation sprint, from $10,000 (₹4L) over 4–8 weeks, takes one workflow live on the systems you already run, with acceptance criteria agreed before the build and hours saved signed off at the end. Managed AI-ops, from $2,000 (₹1L) a month, covers monitoring, cost per task, prompt and model upgrades and a monthly improvement backlog. A disclosure: we sell both, so weigh our view accordingly. The checklist works the same whoever runs it.

The short version

An agent leaves the pilot when seven things exist. They are an owner, a baseline with signed criteria, cost per task, a log, a staffed exception queue, a review cadence and kill criteria. Most firms in the 2026 surveys are missing at least one, most often cost visibility. Fill the gaps in shadow mode, supervise for a month and let the owner sign the result at day 90. More detail on how we scope this is on our AI automation page.

Sources

  1. Gartner, survey of 1,303 leaders at organisations with $50M+ revenue, Jan–Apr 2026; geography not stated (Sep 2026)
  2. McKinsey, The State of AI 2026: n=1,719, 97 countries, 4 May–8 Jun 2026 (Aug 2026)
  3. McKinsey, The State of AI 2026: organisations with more than $1 billion in revenue vs $1 billion or less (Aug 2026)
  4. KPMG, AI Quarterly Pulse Q2 2026: 204 US leaders at $1bn+ firms, 28 Apr–25 May 2026 (Jun 2026)
  5. Microsoft, FY26 Q4 earnings, 29 Jul 2026: registered agents, not active ones
  6. CX Dive, only one quarter of AI customer service use cases produce ROI (Aug 2026)

Questions readers ask

  • When the seven checklist items exist in writing: a named owner, a measured baseline with signed acceptance criteria, cost per task tracking, an action log, a staffed exception queue, a review cadence and kill criteria. A good error rate in a demo is not enough. If any item has no name next to it, keep the agent in shadow mode until it does.

Keep reading

Free, in two minutes. Enter your domain and we'll score it against three competitors across six engines.

No account needed. The report is emailed within 24 hours.