Skip to content

How long should an AI agent pilot run, and what should it cost?

Set a pilot's length by the cases it must handle, not by the calendar. The budget lines to plan for, and the stop-or-go rules to agree before the build.

By Published Updated 7 min read
Costs and buying, 7 min read — A glass hourglass on a dark surface, its falling grains small glowing cubes that collect into neat stacks below.

The short answer

An AI agent pilot should run until it has handled enough live cases to measure its exception rate and cost per task, usually a few hundred, which sets its length from your volume. Budget for the build, platform usage, people handling exceptions and measurement. Agree the stop-or-go thresholds before anyone builds.

Key takeaways

  • A pilot's length should follow from the number of live cases needed to measure it, not from a calendar slot.
  • In Gartner's survey of 1,303 leaders (January to April 2026), only 22% of organisations had scaled AI across multiple business units or adopted an AI-first approach.
  • In KPMG's survey of 204 US leaders at large firms (April to May 2026), 53% were deploying AI agents but only 26% had full, real-time visibility of AI operating costs.
  • A pilot budget has five lines: build and integration, platform and usage, people time for exceptions, measurement, and contingency.
  • Write the stop, extend and go thresholds before the build, and do not move them once results start arriving.

Set the length by cases, not by weeks

A pilot is long enough when it has handled enough real cases to measure what you care about. For most first agents that means the exception rate and the cost per task, and both need a few hundred live cases before the numbers settle. Divide that by your weekly volume and you have the pilot's length.

Calendar-based pilots get this backwards. A "six-week pilot" on a queue that sees forty items a week ends with too few cases to judge. The same six weeks on a queue of two thousand a week is far longer than needed, and the extra weeks are spent arguing about a result that was clear in week two.

Our view: write the case count into the pilot plan, not the end date. The date then follows from volume, and nobody can declare victory early on a handful of easy cases.

Why so few pilots become production systems

Most organisations still struggle to move AI from trials to wide use. In Gartner's survey of 1,303 leaders at organisations with revenue of $50 million or more, run from January to April 2026 (geography not stated), only 22% had scaled AI across multiple business units or adopted an AI-first approach.1 About 11% did not know what their function had spent on AI in 2025.1

Value is just as uneven. In McKinsey's survey of 1,719 respondents in 97 countries (4 May to 8 June 2026), 37% attributed any EBIT impact to AI.2 In the same survey, 40% of respondents at organisations with revenue of $1 billion or more said they were scaling AI agents, against 22% below that size.3 Our note on the McKinsey survey looks at that gap between scaling and profit.

Cost visibility is part of the problem. In KPMG's survey of 204 US leaders at firms with $1 billion or more in revenue (28 April to 25 May 2026), 53% were deploying AI agents, but only 26% had full, real-time visibility of what their AI systems cost to operate.4 A pilot that does not measure its own cost cannot make the case to scale.

AI pilots in 2026

Plenty of trials, few scaled and measured

22%of organisations scaled AI across multiple business unitsGartner, Jan–Apr 2026 1
37%of respondents attribute any EBIT impact to AIMcKinsey, May–Jun 2026 2
26%of large US firms see AI operating costs in full, in real timeKPMG, Apr–May 2026 4
Sources: Gartner, McKinsey, KPMG. Different samples and questions; read each on its own.

How many cases does a pilot need?

Enough to estimate the exception rate within a margin you can act on. If you expect about one item in five to need a person and want the estimate within five percentage points either way, a standard sample-size calculation gives roughly 250 cases. That is a rule of thumb, not a law; rarer exceptions or tighter margins need more.

Count only live cases that reached the agent, not test data. And spread them over at least two full business cycles, such as two month-ends for invoices, so that unusual weeks are included.

Items reaching the agent each weekWeeks to about 250 live casesSuggested pilot length
Around 50About 56 to 8 weeks, or widen the scope
Around 150About 24 weeks, to cover a month-end
Around 500Under 13 to 4 weeks, to cover a full cycle
Under 20Over 12Rethink the process choice

Illustrative. Assumes an exception rate near one in five and a margin of about five points; add time for build and a short shadow run.

Very low volume is a signal in itself. If the pilot needs a quarter to see enough cases, the agent will take years to earn back its build. Our scoring model for choosing the first process treats volume as the first filter for that reason.

What should a pilot cost?

A pilot's cost has five lines, and the build is often not the largest. Teams that budget only for the build are surprised by the people time an agent needs while it learns your exceptions.

  • Build and integration. Connecting the agent to your ERP, CRM, helpdesk or inbox, and writing its rules and prompts.
  • Platform and usage. Licences, model calls and hosting for the pilot period.
  • People time. The staff handling exceptions, the owner who approves and the reviewer who checks samples.
  • Measurement. Setting the baseline before launch, logging every case and producing the final comparison.
  • Contingency. Data clean-up or an integration that takes longer than planned.

Our breakdown of monthly running costs covers platform, usage and exception handling after launch. A pilot is the first month of those costs plus the build.

Stop, extend or go: agree the rules first

Write the thresholds into the pilot plan, signed by the process owner, before the build starts. The three outcomes are go to production, extend once with a named fix, or stop.

Stop-or-go rules for a first agent pilot

Go

  • Exception rate at or below the agreed target
  • Cost per task below the measured baseline
  • No serious error reached a customer or the ledger
  • The owner signs the result

Extend once

  • Close to target with one known cause
  • Too few cases to judge
  • A named fix and a new end date

Stop

  • Exceptions far above target
  • Serious errors not caught by review
  • No owner willing to sign
Illustrative. Our recommended criteria; set your own numbers before the build.

Our view: allow one extension, never two. A pilot that needs a second extension has usually picked the wrong process or found a data problem, and both are better fixed outside the pilot.

Measure the right thing too. Gartner found that productivity was the target for 75% of the functional leaders it surveyed.1 Productivity is a fine goal, but turn it into hours saved per case and cost per task, measured the same way before and after. Our note on making service AI measurable shows how.

Where our sprint fits

The Sigzen AI automation sprint is structured as a pilot that ends in production. On our pricing page it is listed from $10,000 (₹4L) over 4–8 weeks for one workflow, with acceptance criteria agreed before the build and hours saved signed off at the end. The case count and thresholds above are what those acceptance criteria contain.

If you are not sure which process to pilot, the AI readiness assessment, from $6,000 (₹2.5L) over three weeks, scores your candidates first; our AI automation page describes both. After launch, managed AI-ops, listed on the same pricing page from $2,000 (₹1L) a month, covers monitoring, cost per task and model upgrades. That is the full cost view that only about a quarter of the firms in KPMG's survey had.4

Sources

  1. Gartner, survey of 1,303 leaders at organisations with $50M+ revenue, Jan–Apr 2026; geography not stated (Sep 2026)
  2. McKinsey, The State of AI 2026: n=1,719, 97 countries, 4 May–8 Jun 2026 (Aug 2026)
  3. McKinsey, The State of AI 2026: organisations with $1 billion or more in revenue vs below $1 billion (Aug 2026)
  4. KPMG, AI Quarterly Pulse Q2 2026: 204 US leaders at $1bn+ firms, 28 Apr–25 May 2026 (Jun 2026)

Questions readers ask

  • Only when volume is high. A queue that sends several hundred live items a week to the agent can reach a measurable sample in two weeks, but the pilot should still cover one full business cycle, such as a month-end. For lower-volume processes, two weeks shows whether the agent works at all, not whether it pays.

Keep reading

Free, in two minutes. Enter your domain and we'll score it against three competitors across six engines.

No account needed. The report is emailed within 24 hours.