The "digital employee" promise
Major platforms already sell the idea of "digital labor": agents handling sales, service and marketing like extra staff. And the forecasts back it: Gartner estimates that by 2029 agentic AI will autonomously resolve 80% of common customer service issues, with a 30% cut in operational costs (Gartner, March 2025).
Careful: that's a five-year prediction, not the present. The useful question is what an agent can do today, and there it pays to look at the data.
What they do well today
On bounded, repeatable tasks, agents perform. In Salesforce's CRMArena-Pro benchmark (2025) —which evaluated nine leading models across 4,280 real CRM queries— the best model reached about 83% accuracy on single-turn case routing and workflow execution. In other words: classifying, routing, preparing information and running defined processes is reasonably within reach.
And capability is growing fast. The evaluation org METR measured that the length of tasks an agent completes reliably doubles roughly every 7 months, accelerating in 2024-2025 (METR, March 2025). The trend is clear; the problem is the base it starts from.
Where they break (and why it matters)
The same Salesforce study shows the ceiling: the best agents scored only 58% on single-turn tasks and dropped to ~35% on multi-turn conversations. Worse, they showed near-zero confidentiality awareness: they leaked sensitive data unless explicitly told not to. A human employee knows what not to say; the agent, by default, doesn't.
At enterprise scale, the pattern repeats. The GenAI Divide report from MIT's NANDA initiative (2025) found that 95% of enterprise generative-AI pilots produced no measurable impact on the bottom line; only ~5% created significant value. The main cause wasn't the model but a "learning gap": most tools don't retain feedback or improve with context. (Honest caveat: that 95% refers to generative-AI pilots broadly and to "measurable P&L impact", not to technical failure.)
The Claudius case: what happens when you leave the AI alone
The most honest experiment came from Anthropic itself with Project Vend: they had Claude ("Claudius") autonomously run a real vending business —buy stock, set prices, handle Slack requests— on a budget of about $1,000. The result? It lost money. It stocked the fridge with tungsten cubes and, in a later phase, was talked into giving product away. It's the best illustration that an unsupervised agent isn't a fully autonomous employee yet, however convincing it sounds.
The honest conclusion: copilot, not autonomous employee
Today, an AI agent is an excellent copilot and a mediocre autonomous employee. It performs when the work is bounded, a human supervises the important cases, and the system leaves a trail of what it does. That's why even the vendors selling "digital labor" put the human-in-the-loop at the center, not as a plan B.
The good news is the trajectory: if capability doubles every few months, what needs constant supervision today will need less tomorrow. The sensible strategy is to build the processes with a human in the loop from the start, so you can loosen supervision as the agent proves reliable —not the other way around.
How to get value today without disappointment
- Narrow the scope: one specific, repeatable task performs far better than "do everything".
- Human in the loop: supervision where an error costs money, reputation or data.
- Measure for real: resolution rate without a human, errors, time saved. No dashboard, no proof it works.
- Protect what's sensitive: explicit rules on which data it can touch and which it can't.
- Start small: a measurable pilot before scaling, so you don't join that 95% with no return.
FAQ
Can an AI agent replace an employee today?
Not autonomously for complex tasks. In CRMArena-Pro (Salesforce, 2025), the best agents scored 58% on single-turn tasks and only 35% on multi-turn. It works better as a supervised copilot than as an independent employee.
Which tasks do AI agents do well?
Bounded, repeatable tasks: executing workflows, routing cases, classifying and preparing information. In that same benchmark, the best model reached 83% on single-turn routing and workflows. Reliability drops when the task is open-ended or multi-step.
Is agent autonomy improving?
Yes, and fast. Per METR (2025), the length of tasks an agent completes reliably doubles roughly every 7 months. But it's still measured in minutes and hours, not full unsupervised workdays.
Recommended next step
If you're weighing an agent for your operation, I'll help you do it right: in a free 20-minute diagnosis we define the narrow scope, where to put the human in the loop, and what to measure so it actually returns.