← Blog.
ENRO
AI agents

Digital Employees: Where AI Agents Break in Real Operations

The pitch is a worker that never sleeps. The published record is thinner: most of the proof comes from vendors' own case studies, and the reasons projects get cancelled have little to do with what the model can do.

A digital employee is an AI agent sold as doing a whole job, not a task. Evidence that such agents work in operations comes mostly from vendors' own case studies, and Gartner forecasts that over 40% of agentic AI projects will be cancelled by the end of 2027. Its reasons: cost, unclear value, risk controls. None is about the model.

What follows is a reading of the public record to September 2026: vendor case studies, company statements, analyst forecasts and the EU's own pages. Where a figure comes from a seller of the product, it says so.

What is a digital employee?

It is the name vendors give to an AI agent that takes on a role: answering support tickets, triaging IT requests, coding invoices, screening applicants. The difference from a chatbot is that an agent acts inside your systems instead of only answering questions; the distinction is laid out in what is an AI agent.

The label runs ahead of the product. Gartner describes a widespread "agent washing", vendors rebranding assistants, chatbots or robotic process automation as agentic, and estimates that only about 130 of the thousands of vendors claiming agentic solutions offer genuine agentic features (RCR Wireless, reporting Gartner). Menlo Ventures, by its own definition, counts only 16% of enterprise deployments and 27% of startup deployments as true agents (Menlo Ventures).

Is there evidence that AI agents work in real operations?

Some, and it is worth looking at how it is made. Take a simple bar: at least three named organisations running the product, each with a published outcome figure. Customer support clears it. Klarna says its assistant does the work of more than 853 full-time agents (CX Dive). Salesforce's CEO says he cut his own support division from 9,000 people to about 5,000 (Fortune). Intercom reports its agent resolving 67% of inquiries on average across its customers (Create With, quoting Intercom). Clinical documentation clears it too, with named health systems on two vendors, such as Reid Health on Abridge and Houston Methodist on Ambience (Abridge, Ambience).

Apply the same bar to ten more back-office jobs and nine of them clear it. That sounds like proof until you look at who publishes it. In six of the nine, every named organisation comes from one seller's own case studies. The IT service desk has seven named customers, all of them Moveworks's (Moveworks); accounts payable has three, all of them Vic.ai's (Vic.ai). A job clears the bar when one vendor has a marketing team and customers willing to be named. That is useful to know. It is not evidence of what works across buyers.

The cleanest case is medical coding, with three health systems on three independent vendors. Geisinger cut coding time from over 20 minutes to under 2.5 seconds per chart with Nym (Nym); Oregon Health and Science University automated 92% of radiology coding with CodaMetrix (CodaMetrix); Your Health reached a 95.5% automation rate at 98.3% accuracy with Fathom (Fathom). Even there, every figure is published by the vendor or the buyer, and none was independently evaluated.

Where do AI agents break once they are deployed?

In the public examples below, the model is not what had to be fixed. The breaks sit in the organisation around it.

At the metric

Klarna is the public example. Its CEO said in May 2025 that cost had been "a too predominant evaluation factor" and that the result was lower quality (Fortune). An agent measured on cost per conversation will be optimised for cost per conversation. The full story is in what Klarna's AI U-turn teaches.

At the handover to a person

The same CEO's second point was that a customer must always know there will be a human if they want one. Every agent meets cases it should not handle. Whether those cases reach a person quickly, and whether anyone counts them, decides how the agent is judged by the people it serves.

At the number of agents

In a survey of 1,900 IT leaders, 96% of organisations said they already use AI agents in some capacity, yet only 12% had a centralised platform to manage them, and 94% were concerned that this sprawl increases complexity, technical debt and security risk (OutSystems). The publisher sells agent governance, so read the concern figure with that in mind. The 12% needs no such caution: most companies running agents do not have one place that lists them.

At the business case

Gartner's three reasons for the cancellations it expects are rising costs, unclear business value and inadequate risk controls (RCR Wireless). All three are decided before the agent answers its first ticket: what it will cost to run at real volume, what number it is supposed to move, and who checks its work.

Why does trust matter more than capability when buying an AI agent?

In the interviews behind MIT NANDA's 2025 report on AI in business, executives named trust in the vendor more often than any other criterion when choosing one, ahead of understanding their workflow, minimal disruption to existing tools, clear data boundaries, the ability to improve over time and flexibility. One buyer is quoted: "We're more likely to wait for our existing partner to add AI than gamble on a startup."

Read the list again as a specification and it describes behaviour over time. None of it can be shown in a demo. It can only be shown by a record, which is why the evidence a buyer can check for themselves matters more than the feature list.

What does the EU AI Act require for AI agents?

This is not legal advice, and the timeline has already moved once. Two duties apply now and were not deferred by the Digital Omnibus: the Article 50 transparency duty, which requires telling people when they are interacting with an AI system, and the Article 4 duty to take measures supporting AI literacy among staff (Cloud Security Alliance). What Article 4 means for a small company is in AI Act Article 4: what SMEs need to do.

The rules for high-risk areas, which include employment, apply from 2 December 2027, and from 2 August 2028 for AI embedded in regulated products (European Commission). An agent that screens job applicants sits in that category. Screening is also one of the back-office jobs with published deployment evidence across more than one vendor, so a job with some of the better evidence comes with the strictest rules.

How should you evaluate an AI agent before deploying it?

  • Ask for named customers with their own outcome figures, and check whether they all come from one vendor's marketing.
  • Ask what happens when the agent is wrong or unsure: who takes over, how fast, and where that is counted.
  • Measure quality and cost against a baseline taken before the pilot, not against the vendor's benchmark.
  • Keep one inventory of every agent running, who owns it and what it can access.
  • Map the AI Act duties: disclosure to the people talking to it, literacy for the staff who supervise it, and whether the job sits in a high-risk area.

The word doing the most work in "digital employee" is the second one. A new employee gets a manager, a probation period and someone who reads their first month of work. Most agents get a launch date.

If you are deciding where an AI agent fits in your operations: vladtudor.com/consulting/operations.

Frequently asked questions

What is a digital employee in AI?

A digital employee is an AI agent marketed as taking on a whole role, such as answering support tickets, triaging IT requests or coding invoices, rather than a single task. The label often runs ahead of the product: Gartner estimates only about 130 of the thousands of vendors claiming agentic solutions offer genuine agentic features.

Do AI agents work in real business operations?

In some jobs there is published evidence, mostly from vendors. Customer support, clinical documentation and medical coding each have three or more named organisations with outcome figures. In most other back-office jobs, every named customer comes from a single seller's own case studies, and none of the figures was independently evaluated.

Why are agentic AI projects cancelled?

Gartner forecasts that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing rising costs, unclear business value and inadequate risk controls. Public corrections such as Klarna's point the same way: the agent was measured on cost, and quality suffered until people were brought back.

Does the EU AI Act apply to AI agents?

Yes. The Article 50 transparency duty (telling people they are dealing with an AI system) and the Article 4 AI literacy duty were not deferred by the Digital Omnibus. Rules for high-risk areas, including employment, apply from 2 December 2027, so an agent that screens job applicants faces the strictest requirements. This is not legal advice.

How do you evaluate an AI agent vendor?

Ask for named customers with their own outcome figures and check that they do not all come from the vendor's marketing. Ask what happens when the agent is wrong or unsure, measure quality and cost against your own baseline, and keep an inventory of every agent running and what it can access.
Work with me

Want to talk it through?

If your company is working out where AI fits, the first conversation is free. A few short questions, and the reply comes from me within 24 hours on working days.