Skip to content

When should an AI support agent hand over to a person? An escalation design for 2026

Escalation is the design, not the fallback. Triggers, ownership, measurement and a pre-launch test script for AI support agents.

By Published Updated 13 min read
AI agents in operations, 13 min read — A path of blue-violet light on a dark plane splitting at a glass junction, one branch curving towards a steady point of light.

The short answer

Hand over when the customer asks, when the agent cannot ground its answer in your data, when the action is costly or hard to undo, when the customer is upset or vulnerable, and when the agent has failed twice. Write those triggers down before launch, name who owns each handover, and measure resolution rather than containment.

Key takeaways

  • Escalation rules decide whether a support agent helps or harms, so design them before the happy path, not after launch.
  • Five triggers cover most handovers: the customer asks, the answer is not grounded, the stakes are high, the customer is upset or vulnerable, or the agent has failed twice.
  • In Gartner's survey of 3,566 customers in February and March 2026, people were about three times as likely to use third-party GenAI as a company chatbot for service.
  • In KPMG's survey of 204 US leaders at firms with $1 billion or more in revenue (April to May 2026), 53% were deploying AI agents and 26% had full real-time visibility of AI operating costs.
  • Measure confirmed resolution, repeat contact and time to a person by trigger; containment alone rewards an agent for keeping customers away from help.
  • Every handover passes a written summary, so the customer never repeats the story.

Why escalation is the design, not the fallback

An AI support agent is judged by its worst conversations, and those are the ones it should have handed over. Most projects design the happy path first: the order status, the password reset, the delivery date. The handover gets a single rule at the end, "escalate if unsure", and nobody defines unsure.

The cost of that shows up in customer behaviour. Gartner surveyed 3,566 B2B and B2C customers in February and March 2026. People were about three times as likely to use a third-party GenAI tool, such as ChatGPT, as a company's own chatbot for service.1 Gartner reports that company chatbot use has barely moved since 2022. Our note on that survey covers what it means for scoping.

In our judgement, a customer stuck in a bot loop once is slow to try it again. The handover is the moment the agent either earns that second visit or loses it. Our view: escalation rules are the most important part of a support agent's specification, and they should be written, tested and signed off before the first happy-path prompt.

This post sets out the triggers we would use and a matrix for what the agent may do alone. It then covers who owns a conversation after the handover, what to measure and a test script to run before launch. It builds on our scoring model for a first agent, which covers whether support should be the first process at all.

Where support agents stand in mid-2026

Agents are moving from pilots into production, and the controls are behind. KPMG's AI Quarterly Pulse surveyed 204 US leaders at companies with $1 billion or more in revenue, from 28 April to 25 May 2026. Of those, 53% said they were deploying AI agents.5 In the same survey, 18% orchestrated multiple agents, up from 9% the quarter before.5

Visibility has not kept pace. Only 26% of those leaders had full real-time visibility of their AI operating costs.5 If a company cannot see what its agents cost, it is unlikely to see which conversations they mishandle. Our note on the KPMG survey covers the cost side.

Customers, meanwhile, have moved ahead of company bots. Among GenAI users in Gartner's customer survey (February to March 2026), 58% had used it to act on their behalf, and in B2B settings the figure was 74%.1 Customers expect an agent to finish a task. When it cannot, they expect a person who can.

Support agents in 2026

Customers expect action; companies are still catching up

About 3×more likely to use outside GenAI than a company chatbotGartner, Feb–Mar 2026 1
74%of B2B GenAI users have had it act on their behalfGartner, Feb–Mar 2026 1
26%of large US firms see AI operating costs in full, in real timeKPMG, Apr–May 2026 5
Sources: Gartner customer survey; KPMG, AI Quarterly Pulse Q2 2026.

What the staffing research says

The staffing evidence says support agents change the work more than they remove it. In a Gartner survey of 321 customer service leaders in October 2025, 20% had reduced staffing because of AI, while 55% kept staffing stable and handled more volume.3 Gartner also predicted in February 2026 that half of the companies that cut service staff because of AI will rehire by 2027, under different job titles.4

The return on the spending is thin so far. A separate Gartner survey of 1,303 senior leaders (January to April 2026) found service and support teams put a median 12% of their 2025 budget into AI.2 Only 24% of service and support leaders showed a positive financial return across their AI use cases.2

Our view: those rehired roles are, in large part, the people who take the handover. If the business case assumes the agent removes them, the escalation design will be starved of staff from day one. Plan the human side of the queue as a permanent function with its own hours, skills and targets.

Five triggers for a handover

Five triggers cover most of the cases where an agent should stop and pass the conversation on. Each needs a signal the system can detect and a defined action, or it is a hope rather than a rule.

1. The customer asks for a person

When a customer asks for a person, hand over. Do not ask why, do not offer one more answer first, and do not make the request depend on exact wording. "Human", "agent", "someone real", "speak to a person" and their equivalents in every language you serve should all work. Tell the customer what happens next and how long it is likely to take.

2. The answer is not grounded

Hand over when the agent cannot find the answer in your approved sources or the customer's own records. A well-built agent knows whether it retrieved a matching policy or order; it does not need to guess. If retrieval returns nothing relevant, or returns policies that conflict, the agent should say so and route the conversation rather than compose a plausible answer.

3. The stakes are high

Some actions go to a person whatever the agent's confidence: refunds above a set amount, contract changes, account closures, security changes such as a new email or phone number, anything legal or medical. These are defined by category before launch, not judged in the moment. The agent can gather the details and prepare the case; a person approves.

4. The customer is upset or vulnerable

Hand over on clear signs of distress, a formal complaint, a threat to leave, or signals of vulnerability such as bereavement, illness or financial hardship. Sentiment detection is imperfect, so pair it with keywords and with a lower bar: a false handover costs a few minutes of staff time, while a missed one can cost the customer.

5. The agent has failed twice

Hand over after two failed attempts at the same request: the customer rephrases, says "that's not what I asked", or the same intent repeats. Two is a judgement, not a law. The point is a fixed number written into the specification, so loops end by rule rather than when the customer gives up.

TriggerSignal the system detectsWhat the agent does
Customer asksRequest for a person, any phrasingHands over at once, states next step
Not groundedNo matching source or recordSays it cannot confirm, routes with summary
High stakesAction category on the listPrepares the case, a person approves
Upset or vulnerableSentiment, keywords, complaintAcknowledges, hands over with priority
Repeated failureSame intent fails twiceStops, apologises, hands over

Our trigger set. Illustrative; the thresholds and categories are yours to set.

What may the agent do on its own?

The agent may act alone where a mistake is cheap and its answer is grounded in your data. Everywhere else, it prepares and a person decides. Two questions place any request: what does a wrong action cost, and how well can the agent ground it?

The handover matrix

Cost of a wrong action →

Hand over at onceCostly and not grounded: account security, disputes, anything legal.
Agent prepares, person approvesCostly but grounded: refunds above the limit, contract changes.
Ask once, then hand overCheap but unclear: one clarifying question, then a person.
Agent completesCheap and grounded: order status, delivery slots, simple returns.

Grounding in your data →

Illustrative. Our framework; place each of your request types in a cell before launch.

The useful work is filling the matrix with your own request types. Pull three months of tickets, group them by intent and place the top 20 intents in a cell each. Most teams find the "agent completes" cell is smaller than the vendor demo suggested, and the "prepares, person approves" cell is larger and more valuable.

Our view: the "prepares, person approves" cell is where most of the savings in a first support agent come from. The agent does the reading, the lookup and the drafting; a person spends moments approving instead of minutes assembling. That is the same pattern as an invoice agent that drafts and leaves submission to a person, described in our guide to ERP controls for invoice agents.

Who owns the conversation after the handover?

A named queue owns it, and a named person within the queue picks it up. "The team" is not an owner. Each trigger should route to a queue with stated hours, a target time to first human reply and a lead who answers for its results.

Two streams of soft blue light meeting at a glass relay point on a dark surface, one stream passing smoothly into the other without a break.
A good handover passes the whole context along, so nothing is dropped between the agent and the person.

What the person receives

The person taking over should never need to ask the customer to repeat anything. The agent passes a short written summary: who the customer is and how they were verified, what they asked, what the agent checked, what it tried, which trigger fired and what the customer was told to expect. The full transcript sits behind the summary.

What the customer is told

The customer hears three things: that a person is taking over, roughly when, and through which channel. If the queue is closed, the agent says so and offers a callback or a ticket with a reference, rather than pretending someone is coming. Out-of-hours honesty prevents most of the anger that lands on the person the next morning.

A handover that keeps the context

  1. Trigger fires
  2. Agent writes summary
  3. Named queue
  4. Person takes over
  5. Outcome logged
  6. Rule reviewed
Illustrative. Every handover ends in a logged outcome that feeds the next review of the rules.

Tell customers they are talking to an AI

Tell customers plainly that they are talking to an AI agent, and how to reach a person. It is good practice everywhere, and in the European Union it is becoming a legal duty. Article 50 of the EU AI Act requires providers to make sure people are informed when they interact with an AI system, unless that is obvious from the context.6 Article 50 is due to apply from 2 August 2026. The EU has been changing the Act's timetable for other obligations, so confirm the current dates before you rely on them.

Disclosure also helps the handover. A customer who knows it is an agent asks for a person sooner and with less frustration. We are not lawyers, and this is not legal advice: if you serve customers in the EU, have counsel confirm what Article 50 means for your channels.

What should you measure?

Measure whether customers' problems were solved, not whether they stayed with the bot. Containment, the share of conversations that end without a person, is the metric most dashboards lead with. It rewards an agent for ending conversations, including the ones where the customer gave up.

  • Confirmed resolution. The customer confirmed the issue was solved, or the order system shows the action completed.
  • Repeat contact within seven days. The same customer, same intent, back again on any channel.
  • Escalation rate by trigger. A rising "not grounded" rate points to missing content; a rising "failed twice" rate points to a broken intent.
  • Time to a person. From trigger to first human reply, by queue and hour.
  • Handover quality. Whether the person had to ask the customer to repeat anything, sampled weekly.
  • Cost per resolved contact. Agent running cost plus staff time, divided by confirmed resolutions.

Report these by intent, not in one blended number. An agent can be excellent at order status and poor at returns, and a blended figure hides both.

Where the forecasts fit

The long-range forecasts are ambitious, and they are forecasts. Gartner predicted in March 2025, now well over a year ago, that by 2029 agentic AI will autonomously resolve 80% of common customer service issues without human intervention.7 It linked that to a 30% cut in operating costs.7 That forecast concerns common issues, and 2029 is a long way off for a system you launch this year.

Our view: plan for the escalation volume you will actually have in the first year, not the share a forecast implies for the end of the decade. If the agent improves, the queue shrinks and you move staff to harder work. If you staff for the forecast and the agent does not get there, customers wait.

A pre-launch test script

Test the handovers before customers do. The script below takes a few days with two people and a staging copy of the agent connected to test accounts. Run it in full before launch and again after any change to prompts, models or sources.

Five rounds before launch

  1. Ask for a person

    Twenty phrasings, several languages, polite and rude. Every one should hand over at once.

  2. Ask what it cannot know

    Questions outside your sources. The agent should say it cannot confirm and route.

  3. Request every high-stakes action

    Each category on the list. None should complete without a person.

  4. Play the upset customer

    Complaints, threats to leave, signs of hardship. Each should hand over with priority.

  5. Loop it

    Rephrase the same request three times. The second failure should end in a handover.

Illustrative. Our test rounds; record a pass or fail for each script line and fix every fail before launch.

Read the summaries too. For each handover in the test, check that the receiving person could act without asking the customer anything. A handover that passes the trigger test but delivers a useless summary is still a failure.

What to do this month

Start with paper, not software. List your top 20 support intents from three months of tickets, place each in the matrix, and write the five triggers with your own thresholds. Name an owner and hours for every queue a trigger routes to. Then decide the metrics and how you will collect them.

If you already run a support agent, run the test script against it this week. Most teams find at least one trigger that does not fire as specified. Fixing that costs less than a month of customers learning that your bot is a dead end.

For the build itself, our AI automation service puts one workflow live on the systems you already run, with the handover rules and acceptance criteria agreed before the build and the results signed off at the end.

Sources

  1. Gartner, survey of 3,566 B2B and B2C customers, Feb–Mar 2026: third-party GenAI vs company chatbots; geography not stated (Jul 2026)
  2. Gartner, survey of 1,303 senior leaders, Jan–Apr 2026: service and support AI budget share and financial return (Jul 2026)
  3. Gartner, survey of 321 customer service leaders, Oct 2025 (Dec 2025)
  4. Gartner prediction (Feb 2026): half of companies that cut service staff due to AI will rehire, under different job titles
  5. KPMG, AI Quarterly Pulse Q2 2026: 204 US leaders at $1bn+ firms, 28 Apr–25 May 2026 (Jun 2026)
  6. European Union, Regulation (EU) 2024/1689 (AI Act), Article 50 (Jul 2024)
  7. Gartner prediction (Mar 2025): by 2029 agentic AI will autonomously resolve 80% of common customer service issues(dated)

Questions readers ask

  • There is no universal target, because it depends on the mix of requests. Set the target per intent: an order-status intent might rarely escalate, while a billing dispute should almost always reach a person. Our view: watch the trend by trigger rather than one blended rate. A falling rate is only good news if confirmed resolution and repeat contact hold steady.

Keep reading

Free, in two minutes. Enter your domain and we'll score it against three competitors across six engines.

No account needed. The report is emailed within 24 hours.