Skip to content

Designing the exception queue: where an AI agent's hard cases should go

An agent's hard cases decide its ROI. Set the routing thresholds, name a queue owner, give exceptions a service level and feed every fix back into the rules.

By Published Updated 7 min read
AI agents in operations, 7 min read — Glowing glass tiles flow along a dark channel while a few violet tiles divert into a softly lit side tray.

The short answer

Send an AI agent's hard cases to one exception queue with a named owner, written routing thresholds and a service level. Route on rule triggers first and model confidence second. Log why each item failed, and turn repeat reasons into new rules. The queue is where you measure what the agent really saves, so design it before the agent.

Key takeaways

  • An exception queue needs four things in writing: routing thresholds, one named owner, a service level and a feedback loop into the rules.
  • Route on hard rules first, such as amount limits, missing references and regulated topics, and on model confidence only after that.
  • The owner is a person with authority in the system of record, not a shared inbox or a team name.
  • In KPMG's survey of 204 US leaders at $1bn+ firms (28 April to 25 May 2026), 26% had full real-time visibility of their AI operating costs.
  • Cost per resolved item, including the minutes people spend on exceptions, is the number that shows whether the agent pays.

Why the exception queue decides the return

The exception queue decides whether an AI agent pays, because every item the agent cannot finish lands on a person's desk. If those items arrive unsorted, without an owner or a deadline, the agent has moved work rather than removed it. If they arrive labelled, prioritised and owned, you can count exactly what the agent saved.

Agents are now common in large firms, and cost control is not. In KPMG's quarterly survey of 204 US leaders at companies with $1bn or more in revenue (28 April to 25 May 2026), 53% said they were deploying AI agents.1 In the same survey, 26% said they had full real-time visibility of their AI operating costs.1 The minutes your people spend on exceptions are part of that operating cost, and they rarely appear on an invoice.

Our view: design the queue before the agent. The routing rules, the owner and the service level are the acceptance criteria in disguise. Our scoring model for a first process already asks for an exception rate; this post covers what happens to those exceptions once the agent is live.

Which cases should go to the queue?

Route a case to the queue when a written rule says so, and only then when the model's own confidence is low. Rules are testable and auditable. Confidence scores drift when the model, the prompt or the input mix changes, so they make a poor first gate.

Four kinds of trigger cover most processes. Hard limits come first: an invoice above an amount, a refund above a value, a discount above a percentage the business has set. Missing references come next: no purchase order, an unknown customer, a product code the master data does not hold. Then sensitive topics, such as complaints, legal threats, health questions or anything regulated. Last comes low confidence, which catches the cases the rules did not foresee.

TriggerExample ruleWho sets it
Hard limitAny credit note above a set value goes to financeProcess owner
Missing referenceNo matching order or customer recordData owner
Sensitive topicComplaint, legal threat or regulated subjectCompliance
Low confidenceModel score below the threshold set at launchAgent owner

Illustrative. Example triggers; your limits come from your own approval policy.

Write each trigger as a sentence a person could check by hand. "Route if the amount is above the approval limit" can be tested against last month's data. "Route if the case feels risky" cannot. Our escalation design for support agents covers the customer-facing version of these triggers, including when a customer asks for a person.

Who should own the queue?

One named person owns the queue, with the authority to approve, reject or correct each item in the system of record. A shared inbox is not an owner. A team name is not an owner either, because when everyone can pick up an item, the hard ones wait.

The owner is usually the person who approved this work before the agent existed. In accounts payable that is the AP lead; in support, the team lead for the tier the agent covers. They need the system role to submit what the agent drafted, as our least-privilege checklist describes, and a deputy for holidays.

Ownership also means the owner can change the rules. If the same exception arrives forty times a week, the owner should be able to raise it at a weekly review and get a rule changed within days. Our view: an owner who can only clear items, and never change why they arrive, will be clearing the same items in a year.

Where a hard case goes

  1. Agent drafts
  2. Rule and confidence check
  3. Exception queue with reason code
  4. Owner resolves
  5. Weekly rule review
Illustrative. Every exception carries a reason code, so the weekly review can change the rules that create it.

What service level should exceptions get?

Exceptions need a written service level that is at least as fast as the process ran before the agent. Otherwise the agent makes the routine cases quicker and the hard ones slower, and the hard ones are often the customers or suppliers who matter most.

Set two clocks. The first is time to first touch: how long before a person opens the item. The second is time to resolution. For a supplier invoice the clocks might be a day and three days; for a customer complaint, an hour and a day. Publish both to the owner's manager, and report breaches weekly.

Sort the queue by consequence, not by arrival. A payment due tomorrow goes above a duplicate invoice from last month. Each item should show why it was routed, what the agent proposed and what evidence it used, so the owner decides in minutes rather than redoing the work. The fields belong in the agent's log; our audit log field list sets them out.

How exceptions should feed back into the rules

Every resolved exception should leave a reason code, and repeat reasons should become rules or data fixes. That loop is what lowers the exception rate over time. Without it, the rate stays where it was in week one, and so does the cost.

Keep the reason codes short and fixed: missing reference, out of tolerance, new supplier, ambiguous request, policy question, model error. Review the counts weekly with the owner. A reason that keeps growing points to one of three fixes: clean the master data, add or adjust a rule, or change the prompt and test it on last month's cases.

Some exceptions should stay exceptions. A new supplier's first invoice, a legal threat or a contract change deserves a person every time. Our view: a falling exception rate is good news only if the cases leaving the queue are ones a person no longer needs to see. Check a sample of auto-resolved items each month to confirm it.

Measuring ROI through the queue

The queue is where the return becomes measurable, because it holds the cost the agent did not remove. Count resolved items, count exceptions, and time a sample of exception handling. Cost per resolved item is the platform and usage cost plus exception minutes, divided by items completed.

Few companies can show that return at company level yet. In McKinsey's survey of 1,719 respondents in 97 countries (4 May to 8 June 2026), 37% said AI had contributed positively to their organisation's EBIT.2 Gartner's analysis of customer-service AI use cases, reported by CX Dive in August 2026, includes a large group whose leaders could not say what value they produced.3 A queue that logs reason codes and handling time is what keeps a use case out of that group. Our note on that analysis covers what made the measurable cases different.

Staffing is a weak proxy for value. In a Gartner survey of 321 customer service leaders in October 2025, 20% had reduced staffing because of AI, while 55% reported stable staffing while handling more volume.4 Track volume handled per person and cost per resolved item instead. Our pilot-to-production checklist puts these numbers into the go-live review, and our AI automation page lists the exception queue as a deliverable of every sprint.

Sources

  1. KPMG, AI Quarterly Pulse Q2 2026: 204 US leaders at $1bn+ firms, 28 Apr–25 May 2026 (Jun 2026)
  2. McKinsey, The State of AI 2026: n=1,719, 97 countries, 4 May–8 Jun 2026 (Aug 2026)
  3. CX Dive, only one quarter of AI customer service use cases produce ROI, reporting a Gartner analysis (Aug 2026)
  4. Gartner, survey of 321 customer service leaders, Oct 2025 (Dec 2025)

Questions readers ask

  • Use confidence as the last gate, not the first. Hard rules such as amount limits, missing references and sensitive topics are testable and auditable, so route on those first. A confidence threshold then catches cases the rules did not foresee. Re-check the threshold whenever the model, the prompt or the input mix changes, because scores drift.

Keep reading

Free, in two minutes. Enter your domain and we'll score it against three competitors across six engines.

No account needed. The report is emailed within 24 hours.