AI agents now work in the background. How do you supervise one nobody is watching?
Always-on agents act between check-ins. Set an act, ask or block rule per action, read a daily digest, sample the work and give one person a pause switch.

The short answer
Supervise a background AI agent with rules set per action, not with trust in the agent. Let it read freely, act alone only on cheap, reversible steps, ask before anything that moves money or contacts customers, and block the rest. Then read a daily activity digest, sample its work each week and give one named person a pause switch.
Key takeaways
- Background agents need supervision designed for the hours nobody is watching, on top of the permissions set at launch.
- In Gartner's survey of 297 cybersecurity leaders (Q2 2026), 54% of organisations had no defined approach to limiting AI agent access, or relied on predefined human access.
- Give every action an agent can take one of three rules, act, ask or block, and start every action at block.
- A one-screen daily digest and a random weekly sample catch the quiet errors that a permission list cannot.
- Anomaly thresholds should pause the agent, not only alert, and one named person should own the switch.
In this article
Why background agents change the supervision question
A background agent needs supervision built for the hours when nobody is looking at it. Until recently most agents ran when a person asked, and that person read the result. Always-on agents keep working between check-ins.
OpenAI's dots, announced at DevDay on 29 September 2026, are a clear example. Each dot has its own cloud computer and browser, can research in the background and works under rules about what it may do.2 Buyers will meet the same pattern inside tools they already pay for.
Most organisations have not set limits even for the agents they run today. In Gartner's survey of 297 cybersecurity leaders in the second quarter of 2026 (geography not stated), 54% said their organisation had no defined approach to limiting AI agent access, or relied on predefined human access.1 Gartner advises governing agents by the actions they may take, and using guardian agents to limit the damage one can do.1
Our view: permissions decide what an agent can touch. Supervision decides what happens when it misuses that access at three in the morning. This post covers supervision. Our guides to least-privilege access in an ERP or CRM, what an agent's audit log should record and exception queue design cover the rest.
AI agents in 2026
Agents are spreading faster than the rules for them
What does a background agent do while nobody is watching?
It reads, drafts and sometimes acts, depending on the rules you gave it. OpenAI says dots run background research with read-only tools, while actions fall under Custom Rules that allow an action, require approval for it or block it.2
The announcement also lists an Activity View of what each dot did, monitoring that can pause or stop a dot, and specialist dots with their own identity and credentials.2 OpenAI says it plans to connect dots to Microsoft's Agent 365 governance. These are a vendor's descriptions of a new product, not evidence of how well the controls work.
The design is worth copying on any platform. Background time is for reading. Acting happens under rules a person wrote in advance, and every action leaves a trace someone can review.
Scale is the reason to settle this now. Microsoft said on 29 July 2026 that nearly 40 million agents were registered on Agent 365, a count of registered agents rather than active ones.4 Even a small active share is a lot of unattended work.
Act, ask or block: write a rule for every action
Give every action the agent can take one of three rules: act alone, ask a person first, or never. Start every action at block, and move it only for a reason you would defend to an auditor.
Two questions decide the rule. Can the step be undone cheaply, and does it leave the building? A draft saved in the CRM is reversible and internal. A payment, a refund or an email to a customer is neither.
| Action | Example | Rule |
|---|---|---|
| Read records and documents | Search the CRM, open invoices | Act |
| Draft inside your systems | Draft a reply or a journal entry | Act |
| Update an internal field | Tag a ticket, set a follow-up date | Act, logged |
| Send anything outside | Email a customer or a supplier | Ask |
| Move money | Post a payment, issue a refund | Ask, two people above a limit |
| Change access or settings | Add a user, change bank details | Block |
| Delete records | Remove a contact or a document | Block |
Illustrative. A starting set for a finance or operations agent; your risk owner decides the final rules.
Large firms are already drawing this line. In KPMG's survey of 314 US leaders at firms with $1 billion or more in revenue (24 July–25 August 2026), 62% were building or deploying agents and 49% had defined high-risk uses that agents may not decide alone.3 Write your list before the agent goes live, not after the first incident.
How should someone review what the agent did?
Through a short daily digest and a random weekly sample, not by reading every log line. Logs are for investigations; supervision needs something a busy owner will actually read.
The daily digest
The digest lists counts and exceptions: actions taken under each rule, approvals requested and granted, blocked attempts, model spend and anything that crossed a threshold. Blocked attempts matter most. An agent that keeps trying a blocked action is telling you that its instructions and its rules disagree.
Keep the digest to one screen. If the owner skips it for a week, it is too long. The full record belongs in the audit log, which is where an investigation starts.
The weekly sample
Each week, pull a random sample of actions the agent took alone and check them as if a new employee had done them. Was the record right, was the rule right, and would you have done the same? Random means random, not the items the digest flagged.
Our view: sampling catches the quiet errors that thresholds miss, such as a supplier coded to the wrong account every time. Sample more in the first month, and shrink the sample only when it keeps coming back clean.
Supervising an agent between check-ins
- Read-only background work
- Act, ask or block rule
- Daily digest
- Weekly sample
- Threshold pause
Anomaly thresholds and the pause switch
Thresholds pause the agent automatically when its behaviour leaves a normal range, and a named person can pause it by hand at any time. Test both before launch.
Useful thresholds are simple: actions per hour, value moved per day, approval requests per hour, blocked attempts in a row and model spend per day. Set each from a week of normal running. When one is crossed, pause the agent rather than only sending an alert. An alert at night waits until morning; a pause stops the damage.
The pause has to be real. Check that it stops scheduled runs, ends open sessions and leaves unfinished work as drafts a person can find. Gartner's guardian-agent advice points the same way: a second control that limits what the first agent can do.1 Items the agent could not finish should land in an exception queue that someone works every day.
Who owns the agent between check-ins?
One named person owns each background agent: its rules, its digest, its thresholds and its pause switch. When ownership is shared, nobody reads the digest.
The owner should sit in the business, because they judge whether the work is right. IT owns the identity and the credentials. Give each agent its own account rather than a person's login, so every action traces to the agent and its access can be cut without locking anyone out.
Cost belongs in the same review. In KPMG's survey, 74% of leaders included cost reviews in AI approvals.3 An agent that runs all day spends all day, which is why model spend sits in the digest.
A disclosure: Sigzen AI builds agents for clients, and we design them with these rules, a digest and a pause switch, so weigh our view with that in mind. Our AI automation page sets out how we scope that work.
Sources
- Gartner, survey of 297 cybersecurity leaders, Q2 2026; geography not stated (Sep 2026)
- OpenAI, Introducing dots: company announcement at DevDay, 29 Sep 2026
- KPMG, AI Quarterly Pulse Q3 2026: 314 US leaders at $1bn+ firms, 24 Jul–25 Aug 2026 (Sep 2026)
- Microsoft, FY26 Q4 earnings, 29 Jul 2026: registered agents, not active ones


