Klarna's contradiction is the useful part: it reported stable satisfaction, then its CEO said its customer-service cost-cutting had gone too far.
The company says its assistant handled most support chats with no drop in consumer satisfaction. The CEO's later admission about cost-cutting makes this more useful than another argument about whether AI won or failed.
What Klarna reported
When Klarna launched its OpenAI-powered assistant in 2024, it reported:
- 2.3 million conversations in the first month
- two-thirds of customer-service chats handled by the assistant
- average resolution time reduced from 11 minutes to under 2
- service in 35 languages
- work equivalent to 700 full-time agents
- an estimated USD 40 million profit improvement for 2024
“Work equivalent to 700 agents” does not mean Klarna replaced 700 named people with a chatbot. I would describe it as workload equivalence, not 700 people replaced.
Klarna's 2025 annual report later said the assistant handled 80% of customer-service chats with no drop in consumer satisfaction. The same report linked AI adoption and restructuring to a 49% headcount reduction and a 3.6-fold increase in revenue per employee since 2022.
The efficiency numbers are good, but they are Klarna's own. They show what the company measured and chose to publish.
Then the CEO said cost-cutting went too far
Sebastian Siemiatkowski later said that the company's customer-service cost-cutting push had gone too far. Klarna also started talking about giving customers a clearer option to reach a person.
That does not mean Klarna automated support, watched satisfaction collapse, and hired everyone back. The public evidence does not show that.
Klarna did not abandon the assistant. Its filing says usage increased. The public record also does not support precise claims about an NPS collapse, hundreds of agents returning, or a proven recovery under a new three-tier model.
So I would keep the contradiction visible:
Klarna reported stable consumer satisfaction while its CEO said the customer-service cost-cutting push had gone too far.
Both can be true if the two statements use different measures, periods, customer groups, or definitions of quality. Without the underlying measurement method, we do not know.
One support metric cannot carry the decision
Resolution time is useful, but it measures speed. It does not tell you whether the customer had to contact support again, whether the issue was solved correctly, or whether a difficult case reached the right person.
The same problem applies to a single satisfaction score. An average can stay stable while a smaller group with complex problems gets a much worse service.
I would not judge support automation on resolution time alone. Can we check repeat contacts, escalation, satisfaction by case type, and cost together?
| Measure | What it tells you |
|---|---|
| Resolution time | How quickly the interaction closes |
| First-contact resolution | Whether one interaction solved the issue |
| Repeat contact | Whether the customer had to come back |
| Escalation rate | How often the system needed a person |
| Satisfaction by case type | Which journeys improved or became worse |
| Cost per resolved case | Whether the financial gain survives rework |
If speed improves while repeat contact and escalations rise, the automation may be moving work rather than removing it.
Draw the automation boundary before launch
Routine tracking questions and simple status updates are different from disputes, fraud reports, account access, or a payment problem that has already failed twice.
I would not start by automating every case and wait for complaints to show where the boundary belongs. Define the high-risk cases first, send those to people, and use AI to help the agent gather context or prepare an answer.
The handoff must also work in practice. A human option hidden behind repeated bot replies only adds another queue.
Efficiency still matters
The “AI failed” version ignores Klarna's financial results. The company reported USD 1.082 billion in Q4 2025 revenue, 118 million active consumers, and 966,000 merchants.
Revenue, active consumers, and merchants all increased while the AI programme continued. That rules out the simple story that it was a commercial collapse, but it does not establish which part caused the growth.
Cost reduction and service improvement can overlap, but they are not the same outcome. A product leader needs to know which one a dashboard is showing.
What I would take from the case
I would make four decisions from this case:
- Describe workload equivalence as workload equivalence, not jobs replaced.
- Read speed, satisfaction, repeat contact, escalation, and cost together.
- Decide which cases require human judgment before scaling the automation.
- Keep an obvious route to a person when the automated path fails.
My takeaway is simpler: read efficiency and service quality together. A clean public story is not enough.
Sources
- Klarna, OpenAI. Reports the assistant's first-month volume, share of chats, response-time change, language coverage, and estimated workload.
- Klarna Group plc 2025 Annual Report, Klarna. Reports the assistant's service share, consumer satisfaction, headcount change, revenue per employee, customers, and merchants.
- Klarna Slows AI-Driven Job Cuts With Call for Real People, Bloomberg. Reports Siemiatkowski's view that the customer-service cost-cutting push went too far.
- Klarna CEO says company will use humans to offer VIP customer service, TechCrunch. Records the plan to restore more human support while retaining AI.
- Klarna Accelerates U.S. Growth and Delivers $1bn Revenue Driven by Rapid Banking Service Adoption, Klarna. Supports the reported revenue, customer, merchant, and efficiency figures.
Get Personalized Help
Copy this prompt to ChatGPT, Claude, or your favorite AI assistant. Fill in your details and get guidance tailored to your specific situation.
I'm applying the support-automation test at https://ivanmisic.net/blog/ways-of-working/klarna-ai-u-turn-isnt-failure-story to one real service operation. My context: - Support scope and case types: [ROUTINE QUESTIONS, DISPUTES, FRAUD, ACCESS, PAYMENTS, OR OTHER] - Current automation and human route: [WHAT THE SYSTEM HANDLES AND HOW A PERSON TAKES OVER] - Measures available by case type: [RESOLUTION TIME, FIRST-CONTACT RESOLUTION, REPEAT CONTACT, ESCALATION, SATISFACTION, AND COST] - Known failure or customer-harm signals: [COMPLAINTS, REWORK, ABANDONMENT, WRONG ANSWERS, OR UNKNOWN] - Policy and capacity constraints: [REGULATION, REQUIRED HUMAN REVIEW, LANGUAGES, STAFFING, AND HOURS] Ask for missing facts before judging success. Then: 1. build one measurement view that reads speed, resolution quality, repeat contact, escalation, satisfaction by case type, and cost per resolved case together; 2. separate cases suitable for automation, cases where AI should only assist an agent, and cases that require a person; 3. design an obvious handoff that carries the conversation and relevant context to the human route; 4. propose a staged trial with case-level monitoring, stop conditions, and an accountable decision owner; 5. state what the available data cannot prove. Do not describe workload equivalence as jobs replaced, average away a harmed case group, or recommend staffing cuts from efficiency data alone. Verify current regulatory, contractual, and service-policy constraints with primary sources or qualified owners, and identify anything you cannot verify.