I run a fleet of AI agents. They sit in chat channels, one per area of the business, each with a defined scope and a defined set of things it is accountable for. They're always on, they don't have a queue that empties, and they do not go home. That last part sounds like the advantage. In practice it is the thing that changed my week the most, and not in the direction I expected.
What the week actually looks like
The shape of my day shifted from doing work to specifying work and then checking it. That sounds like a promotion. It is mostly harder. Doing a task gives you continuous feedback: you can feel when something is off. Specifying a task gives you one shot at being clear, and the feedback arrives later, in the form of output that is either right or confidently wrong.
Monday is heavier than it used to be, because the front of the week is where specification happens. If the scope of what I want is fuzzy on Monday, I get fuzzy work all week and the cost of correcting it lands on Thursday, when there's no room. The single highest-leverage habit I've built is spending real time on the shape of a request before anything runs.
The middle of the week is supervision. Not reading everything, which doesn't scale and doesn't work anyway. Spot checks against ground truth. If an agent tells me a number, I check the number at the source, not in the agent's summary. If it tells me something is done, I look at the artifact, not the report of the artifact. This is the single discipline that separates a useful fleet from an expensive fiction generator.
What gets delegated
Bounded, repetitive work with a checkable output. Reconciliation. Drafting a first version of something that has an established shape. Watching a system and reporting a state change. Pulling scattered information into one place. Anything where the answer is verifiable in less time than producing it takes.
That last criterion is the real filter, and it's more restrictive than it sounds. If verifying the output costs as much as doing the work, delegation buys you nothing, it just moves the effort and adds a handoff. I've learned to ask that question before I ask whether an agent is capable of the task.
What stays human
Anything with a genuine tradeoff. Anything where the correct answer depends on context nobody wrote down: a relationship, a history, an unstated constraint, a thing we tried two years ago that didn't work for reasons that aren't in any document. Agents are excellent at reasoning from what they can see and have no way to reason about what they can't.
Anything irreversible also stays human. Not because agents are careless, but because the cost asymmetry is wrong: a reversible mistake costs an hour, an irreversible one costs a relationship or a system, and there's no supervision rhythm cheap enough to catch every instance of the second kind before it happens. So the boundary is drawn at the action, not at the confidence level.
Where it breaks
The failure mode that matters is not an agent crashing. That's visible and therefore cheap. The expensive failure is an agent doing something plausible and wrong, describing it accurately in its own terms, and me accepting the description because it reads like competence.
The second failure mode is subtler: silence that looks like success. A process that stopped running produces no output, and no output looks exactly like nothing needed doing. I've been caught by this more than once, and the fix isn't better agents. It's designing every delegated job so that "working" and "not running" produce visibly different signals.
Third, and this is the one that scales badly: agents inherit whatever context you gave them, and context drifts. A rule that was right when written becomes wrong when the situation changes, and nothing in the system notices. Humans notice, complain, and quietly stop following the rule. Agents follow it perfectly into the wall.
The supervision rhythm
Three habits do most of the work.
Check the artifact, not the report. Every claim gets verified against the thing itself at least sometimes, chosen unpredictably. The point isn't to catch every error, it's that a system where nothing is ever checked drifts without limit.
Make absence loud. Anything that runs on a schedule needs a positive signal when it runs, not just an alarm when it fails. Otherwise the quietest week is indistinguishable from the week everything stopped.
Re-read the instructions periodically. Not the output, the scope: what I told each agent it's responsible for. Most of the wrong output I find traces back to an instruction that was correct months ago and is now subtly out of date. That review is cheap and I still put it off, which tells you something about how these systems degrade.
The honest accounting
The fleet doesn't give me back the hours it removes. It converts a large volume of small, shallow tasks into a smaller volume of harder ones: specifying, verifying, and deciding where the boundary sits. That's a good trade, but it's a trade, and anyone selling it as pure time savings hasn't run one for long.
What it does give me is a floor. Routine things happen whether or not I have a good week, and that turns out to matter more than any raw hour count. The output that used to depend on my attention now depends on a system I check, which is exactly the compounding I care about.
The inventory of how work got here is in The Delegation Ledger, and the framework I use for deciding how much load to carry in a given week comes from training: Training Load as a Management System.
More on this pillar: Compounding CEO.
Ready to Transform Your AI Strategy?
Get personalized guidance from someone who's led AI initiatives at Adidas, Sweetgreen, and 50+ Fortune 500 projects.