Klarna did not abandon its AI assistant. In 2024 it handled two-thirds of support chats, by Klarna's count the work of 700 agents. In May 2025 it started hiring people again, after its CEO said cost had weighed too heavily and quality had dropped. By late 2025 the assistant did the work of 853 agents. The lesson is about measurement.
The story is usually told as a reversal, AI in and then AI out. The public record says something more useful than that, and all of it is on the record: Klarna's own press releases and its CEO's interviews.
What did Klarna's AI assistant actually do?
In February 2024 Klarna said its assistant had held 2.3 million conversations in its first month, two-thirds of its customer service chats, doing the equivalent work of 700 full-time agents. It estimated the assistant would bring a $40 million profit improvement in 2024 (Klarna). These are the company's own figures; none was independently audited.
The assistant runs on OpenAI's models, which Klarna says it combined with the product knowledge of Pricerunner and its own interface (Klarna). Klarna bought the model and built the product around it. Keep that in mind for the build-or-buy question further down.
Why did Klarna start hiring humans again?
In May 2025 the CEO, Sebastian Siemiatkowski, told Bloomberg the company was recruiting human agents again (Entrepreneur). His explanation, as reported by Fortune: "As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality." And: "it's so critical that you are clear to your customer that there will be always a human if you want" (Fortune).
Read the first sentence slowly. The assistant was not judged on what it could do. It was judged on what the organisation had chosen to count. If the scorecard is cost per conversation, a system that answers fast and cheaply wins every week, and the conversations it handles badly surface later and elsewhere: in repeat contacts, in complaints, in customers who quietly leave. None of those sits on the dashboard the launch was planned against.
Did Klarna give up on AI in customer service?
No. In November 2025 Klarna said its assistant was doing the work of more than 853 full-time agents, up from 700 at the start of the year, and had saved the company $60 million. In the same report, CX Dive noted that Klarna's customer service and operations costs were still up year on year (CX Dive). The assistant kept running and kept growing. What changed was the shape of the service around it: a customer who wants a person can reach one.
Why do so many AI agent projects get cancelled?
Gartner forecasts that more than 40% of agentic AI projects will be cancelled by the end of 2027, as organisations struggle with rising costs, unclear business value and inadequate risk controls (RCR Wireless, reporting Gartner). It also estimates that only about 130 of the thousands of vendors claiming agentic products offer genuine agentic features, and calls the rest "agent washing".
Notice that none of Gartner's three reasons is that the model could not do the task. Each is about the organisation around the model rather than the model itself. Klarna's story is the second and third reasons, played out in public by a company that had the volume to see them quickly. A forecast is not a measured failure rate, but it points at the same place.
Should you build or buy an AI agent?
The most quoted evidence is MIT NANDA's 2025 report, The GenAI Divide. In its self-reported data, external partnerships (buying customised tools and co-developing them with vendors) reached deployment about 67% of the time, against about 33% for tools built internally. MIT itself warns that the correlation does not necessarily prove causation. Companies that buy may simply be the ones that arrived with a clearer problem.
Klarna did both: it bought the model and built everything around it. For most companies the decision that matters is less build or buy than who owns the measurement. A vendor can supply the agent. Only the company running it can decide that quality counts as much as cost, and check it every week.
What should you measure before and after you launch an AI agent?
The practical version of Klarna's lesson is a short list, and most of it happens before launch.
- A baseline taken before the pilot: today's volume, cost per case and quality, so the agent is compared with something real.
- Quality next to cost from the first week: repeat contacts, complaints, reopened cases, and a sample of conversations read by a person.
- A visible route to a human, and a count of how often customers take it.
- The return on your own numbers: the volume the agent handles, times today's cost of each case, times the share it resolves without a person, set against the licence, the deployment and the running cost.
- A vendor's return figure treated as a claim until your own pilot reproduces it.
For what separates an agent from a chatbot or a scripted automation, see what is an AI agent. For where agents tend to break once they are inside operations, see digital employees. And if a pilot is already running without results, why AI is not delivering results covers the usual causes.
Klarna's first numbers were not wrong. Two-thirds of the chats were handled, and the work of 700 people got done. They were the numbers that are easiest to count, which is why every organisation reaches for them first. The expensive part of the lesson is that a customer service system is judged by the customers it fails, and those are the ones counted last.
If you are planning an AI agent for customer service or operations: vladtudor.com/consulting.