Which process should you hand to an AI agent first? A scoring model you can run in an afternoon
Pick the wrong first process and an agent pilot stalls. Score yours on volume, exception rate, cost per task and reversibility before you build anything.

The short answer
Start with the process that combines high volume, a known cost per transaction, a low exception rate and mistakes you can reverse. In our view that is usually a document or request queue, such as invoice matching or tier-1 support. Score each candidate on those four inputs before you build anything, and stop if none clears the bar.
Key takeaways
- Score candidate processes on four inputs: volume, cost per transaction, exception rate and reversibility.
- Two vetoes stop a build whatever the score: data you cannot trust, and approvals nobody owns.
- In McKinsey's survey of 1,719 respondents (May–June 2026), 37% attributed at least some EBIT impact to AI, so measurable value is not yet the norm.
- In the same survey, 40% of respondents from organisations with revenue of $1 billion or more reported scaling AI agents, against 22% at organisations below $1 billion.
- If no process clears the bar, stop: an assessment that says "not yet" costs less than a stalled pilot.
In this article
Why the first process decides the pilot
The first process you hand to an AI agent decides whether the pilot produces a number anyone believes. Pick a queue with steady volume, a known cost and mistakes you can undo, and you get a clean before-and-after. Pick a messy, judgement-heavy process and you get a demo that never leaves the sandbox.
Measurable value is still the exception. In McKinsey's survey of 1,719 respondents in 97 countries (4 May–8 June 2026), 37% attributed at least some EBIT impact to AI.1 Yet nearly nine in ten used AI regularly in at least one business function. Use is common. Profit you can point to is not.
Our view: most of that gap is decided before anyone writes a prompt. Teams choose the process that is most visible, or most annoying to a senior person, rather than the one whose results can be counted. This post gives you a scoring sheet with four inputs and two vetoes, a worked example and a rule for when to stop. You can run it in an afternoon with a spreadsheet and two people who know the work.
Where are companies actually running agents?
In fewer places than the headlines suggest. Survey adoption, registered agents and agents in production are three different numbers, and most coverage blurs them.
Adoption is not production
Deployment figures are still modest. According to Gartner's 2026 CIO and Technology Executive Survey (published April 2026; data period not stated), 17% of organisations have deployed AI agents.2 Deployed can mean one agent in one team.
Platform counts run far higher, and they measure something else. Microsoft said on its earnings call of 29 July 2026 that nearly 40 million agents were registered on its Agent 365 platform, which counts registered agents, not active ones.3 A registered agent may be a test someone built on a Friday and never opened again.
Neither figure tells you whether agents pay. A count of agents is a count of attempts. For your first process the useful question is narrower: has anyone in your industry run this exact workflow in production, with an exception queue and a measured result? If not, you are the pilot, and the choice of process carries more weight.
Organisations with $1 billion or more in revenue, and below
The split by company size is wide. In the same McKinsey survey (May–June 2026), 40% of respondents from organisations with revenue of $1 billion or more reported scaling AI agents, against 22% at organisations below $1 billion.4 McKinsey describes the smaller group's share as flat.
The survey does not split the smaller group any further. A 100-person distributor and a 3,000-person manufacturer sit in the same group, so the lower figure tells you little about firms of any particular size.
What the split does suggest is that larger organisations have things smaller ones often lack: platform teams, cleaner data and people whose job is to own an agent after launch. A smaller company can borrow the discipline without the headcount. That is what the scoring sheet is for.
AI agents in 2026
Use is common; measured value and scale are not
The four inputs that predict a good first agent
Four inputs predict a good first agent: volume, cost per transaction, exception rate and reversibility. Score each from 1 to 5, weight them and add. AWS's guide to agent frameworks for small and medium businesses (September 2025) already advises starting with one high-value workflow and keeping the scope small.5 The sheet below adds the inputs that tell you which workflow that should be.
Volume
Volume is the first filter, because an agent's fixed costs are high and its running costs are low. Building, testing and owning a workflow costs about the same whether it runs fifty times a month or five thousand. Only volume spreads that cost thin enough to show up in a quarterly review.
Count items, not hours. Invoices received, orders keyed, tickets opened, quotes sent: whatever unit the process moves. Pull three months from the system of record rather than asking the team, because people remember the bad weeks.
Our view: below a few hundred items a month, a first agent rarely earns back its build within a year, however painful the work feels. Low-volume, high-stakes work such as pricing approvals makes a better second project than a first.
Cost per transaction
Cost per transaction turns volume into money, and it is the input most teams have never measured. Ardent Partners' State of ePayables 2025 (published January 2026; sample, period and geography not stated) puts the average cost to process an invoice at $9.84.6
APQC's benchmarks put accounts payable cost at about $0.38 per $1,000 of revenue for top performers, against about $0.92 for bottom performers (published March 2026; sample and period not stated).7
Use benchmarks to sanity-check your own figure, never to replace it. Time a sample of items, multiply by the loaded hourly cost of the people doing them, and add the cost of errors you already fix downstream. A process that looks cheap per item can still be expensive in rework.
Exception rate
Exception rate decides how much work the agent actually removes. An exception is any item the agent cannot finish under the rules you set: a missing purchase order, a price outside tolerance, a customer asking for something off the menu. Each one goes to a person, so the savings come only from the items the agent clears.
Measure it before you build. Take the last few hundred items and mark each one: routine, or did someone have to look something up, chase a colleague or make a call? That share is your floor, because an agent adds exceptions of its own while it learns your edge cases.
Our view: when exceptions are the majority, the agent only re-routes work. It reads every item, fails on most and hands them to the same people with an extra step in front. Set the threshold before you build, write it into the acceptance criteria, and do not move it once the build is under way.
Reversibility
Reversibility is the input that keeps a bad week from becoming a bad quarter. Ask what happens when the agent is wrong, and how fast you find out. A draft invoice held in the ERP for a person to submit is fully reversible. An email to a customer, a payment run or a price change on a live store is not.
Score the worst realistic mistake, not the typical one. Then find the step where a person can catch it before it leaves the building. Most first agents should end in a draft, not an action.
This is also the input you can change by design. Order entry that confirms straight to the customer scores low. The same process with confirmations held for one person's approval scores high, at the cost of some speed. In the first ninety days that trade is usually worth making.
| Input | How to measure it | Weight | Scores 5 when / 1 when |
|---|---|---|---|
| Volume | Items a month, from three months of system data | 30 | Over 2,000 a month / under 200 |
| Cost per transaction | Timed sample × loaded hourly cost, plus rework | 25 | Over USD 10 an item / under USD 1 |
| Exception rate | Share of a recent sample that needed a person | 25 | Under 1 in 20 / over 1 in 3 |
| Reversibility | Worst realistic mistake, and where it is caught | 20 | Ends in a draft / money or messages leave unchecked |
The scoring sheet. Illustrative weights and thresholds: score each input 1 to 5, multiply by its weight and divide the total by 5 for a score out of 100. Our rule of thumb is to build at 70 or above with no veto.
Run the sheet in an afternoon
- List the candidates
Five to eight, named by the people who do the work.
- Pull the numbers
Three months of volume, a timed sample for cost, a marked sample for exceptions.
- Score and weight
One to five on each input, weighted, for a score out of 100.
- Apply the vetoes
Untrusted data or no approval owner removes a process, whatever it scored.
- Decide
Build the top process if it clears the bar, or stop and fix the veto.
Two vetoes: data readiness and approvals
Two conditions stop a build regardless of score: data you cannot trust, and approvals nobody owns. A process can score 90 and still fail on either, because both turn the agent's errors into your errors.

Data you cannot trust
An agent is only as good as the records it matches against. Suppose the supplier master has duplicates, or item codes differ between the ERP and the warehouse. The agent will then match against the wrong record with complete confidence.
The test is quick. Pull fifty recent items and check whether the reference data the agent would need was present and correct for each one. Where it often fails, fix the data first. That project is cheaper than an agent, and it pays off whether you build one or not.
Our view: a data clean-up that takes six weeks is a better first step than an agent that takes six weeks to prove it cannot work.
Approvals nobody owns
Every agent in production needs a named person who owns its exceptions and its results. That person approves what the agent drafts, signs the acceptance criteria before the build and signs the measured result after it. If nobody will put their name to those three things, the agent drifts, and nobody notices until a supplier or a customer does.
Ownership also needs authority built into the system. On an ERP, the agent's user should be able to create and save drafts but not submit them, so submission stays with a role a person holds. ERPNext, for example, sets read, write and submit rights per role.8 It is the simplest control you will ever add, and it is a design rule behind every workflow on our AI automation page.
Where the person sits in a first agent
- Trigger arrives
- Agent drafts
- Exception queue
- Person approves
- System of record
Scoring three processes: a worked example
Here is the sheet run on three processes at an invented company. Fabrikam is a 120-person manufacturer on an ERP, with a finance team of four and a customer service desk of three. The sales director wants order entry automated; the finance lead wants invoice matching.
| Input | Invoice matching | Email order entry | Tier-1 support |
|---|---|---|---|
| Volume | 3,000 a month: 5 | 900 a month: 3 | 1,500 a month: 4 |
| Cost per item | About USD 7: 4 | About USD 12: 5 | About USD 3: 2 |
| Exception rate | 1 in 8: 3 | 1 in 4: 2 | 1 in 10: 4 |
| Reversibility | Draft held for submit: 5 | Held for approval: 3 | Replies sent live: 2 |
| Score out of 100 | 85 | 65 | 62 |
| Veto | None | Item codes differ across systems | None |
Illustrative. Fabrikam is fictional; volumes, costs and exception rates are invented round numbers, not benchmarks.
Invoice matching wins clearly. It has the volume, a cost the finance team can already state and a natural approval step, because invoices sit as drafts until someone submits them. Its exception rate is the weak point, so the acceptance criteria would set a target for it. Our guide to the controls an invoice agent needs in the ERP covers that build.
Order entry is the interesting loser. It has the highest cost per item, which is why sales wanted it first. But its exceptions are custom specifications that need a person, and item codes differ between the ERP and customers' purchase orders. That is a data veto. Fabrikam's better move is to fix the code mapping now and rescore order entry next quarter.
Tier-1 support scores lower than its volume suggests, because each ticket is cheap and replies go straight to customers. Hold replies for approval and its reversibility score rises. Our page for manufacturers describes the document flows we would usually score first on an ERP.
Value vs reversibility: five common first candidates
Value →
Reversibility →
What about customer support agents?
Support agents can be a good first project for tier-1 questions with drafted or tightly scoped replies. They are a poor one when the business case rests on cutting the team. The evidence on staffing is more mixed than most pitches suggest.
In a Gartner survey of 321 customer service leaders in October 2025 (data now a year old), 20% had reduced agent staffing because of AI.9 In the same survey, 55% reported stable staffing while handling more volume.9 Gartner also predicted in February 2026 that by 2027, 50% of companies that attributed headcount reduction to AI will rehire staff to perform similar functions, but under different job titles.10
Our view: score support on the same sheet and treat headcount as an outcome you measure later, not the business case. Where-is-my-order questions grounded in order data score well, and our playbook for e-commerce brands sets out when that agent pays for itself. Refunds, complaints and anything touching a contract do not, because the worst realistic mistake there is expensive and public.
Build it yourself, use a vendor's agent, or hire a studio?
Use your vendor's built-in agent when it covers the whole process and its controls. Build it yourself when you have engineers who will own it after launch. Hire a studio when you have neither and want one workflow live with a measured result. The process matters more than the route, which is why the sheet comes first.
Run each option against the same four inputs and two vetoes. A helpdesk's native agent may handle tier-1 tickets well and still lack the approval step your reversibility score needs. An in-house build gives full control, but someone has to monitor it, upgrade prompts and models, and watch cost per task every month.
A disclosure: Sigzen AI is part of Sigzen Technologies, which has implemented and supported ERPNext for a decade, and we sell the third option. Weigh our view accordingly. The right route is whichever one leaves a named owner, an exception queue and a signed number at the end.
The risk of failure is real on every route. Gartner predicted in June 2025, more than a year before this post, that over 40% of agentic AI projects will be cancelled by the end of 2027.11 It named three causes: escalating costs, unclear business value and inadequate risk controls. That is a forecast, not an outcome. For a vendor's reading of the same forecast and the failure modes behind it, see beam.ai's January 2026 piece; beam.ai sells an agent platform.12
Notice how the three causes map onto the sheet. Volume and cost per transaction speak to value; reversibility and the approval veto speak to risk controls. A high score does not protect a project from cancellation, but a low one gives the forecast a head start.
What happens after you pick one?
Once a process clears the bar, write down what success means before anyone builds. Agree the acceptance criteria: the exception-rate target, the approval step, how hours saved will be measured and who signs them off.
On our published ladder (see pricing), that work sits in two places. The AI readiness assessment, from $6,000 (₹2.5L) over three weeks, inventories your processes and scores the top ten use cases on value, feasibility and risk: this sheet, with more rigour. The automation sprint, from $10,000 (₹4L) over 4–8 weeks, puts one workflow live on the systems you already run. Acceptance criteria are agreed before the build, and hours saved are signed off at the end.
Assessments are credited 50% against a sprint booked within 30 days, as listed on our pricing page. If the right process is obvious from a process review, we skip the assessment and scope the sprint directly.
If no process clears the bar, stop. Fix the data or name the owner, then rescore next quarter. Our view: an assessment that says "not yet" costs less than a pilot that stalls in month four.
Sources
- McKinsey, The State of AI 2026: n=1,719, 97 countries, 4 May–8 Jun 2026 (Aug 2026)
- Gartner, 2026 CIO and Technology Executive Survey (Apr 2026); data period not stated
- Microsoft, FY26 Q4 earnings, 29 Jul 2026: registered agents, not active ones
- McKinsey, The State of AI 2026: organisations with $1 billion or more in revenue vs below $1 billion (Aug 2026)
- AWS, AI agent frameworks for SMB owners (Sep 2025)
- Ardent Partners, State of ePayables 2025 (Jan 2026); sample, period and geography not stated
- APQC, accounts payable cost benchmarks (Mar 2026); sample and period not stated
- Frappe, ERPNext manual: role based permissions
- Gartner, survey of 321 customer service leaders, Oct 2025 (Dec 2025)
- Gartner prediction (Feb 2026): half of companies that cut service staff due to AI will rehire, under different job titles
- Gartner prediction (Jun 2025): over 40% of agentic AI projects cancelled by end of 2027(dated)
- beam.ai, why agentic AI projects will fail and how to avoid it (Jan 2026)


