Some quarters give you noise. This one gave us a very expensive answer. Klarna ran the replacement experiment every executive has privately wondered about, got the outcome many people didn’t want to hear, and the market still seems more interested in the fantasy than the lesson. That tells me the real shortage in enterprise AI right now isn’t tooling. It’s honesty.
Klarna Didn’t Fail at AI. It Failed at the Replacement Fantasy.
This is the most important enterprise AI case study of the quarter, and I don’t think the market is treating it seriously enough. Klarna publicly turned roughly 700 customer-service roles into an AI workforce-transformation story, then had to reverse course and rebuild toward a hybrid support model where AI handles routine interactions and humans take escalations. That is not a story about AI failing. It is a story about management asking the wrong first question.
I’ve been on record since February: agents are extensions of people, not replacements for people. Klarna just paid tuition to learn the same lesson in public. The architecture they’re landing on now — AI on repetitive work, humans on judgment-heavy edge cases — is the extension model I think most healthy businesses should have started with. The winning question was never, “How many people can we cut?” It was, “What could our best people do if they never had to touch the repetitive stuff again?”
That sounds philosophical until you map it to operations. Replacement thinking produces brittle systems because it assumes the messy part of human work is overhead. It isn’t. In customer support especially, the messy part is where trust gets saved, churn gets prevented, and exceptions get resolved. When you optimize away the people who handle ambiguity, you don’t remove complexity. You just defer it until a customer is angry enough for it to become expensive.
There’s another part of this story I don’t want normalized: the human cost has gone oddly undocumented. Klarna’s reversal is confirmed. But we still don’t have public placement or reskilling data for the displaced workers. That absence matters. If companies want applause for AI-led workforce transformation, they should also be expected to show the outcome for the people whose jobs became the experiment.
Business owners should take the right lesson here. Don’t walk away saying AI doesn’t work in service. It clearly does work on routine volume. Walk away saying replacement is the wrong mental model, and bad mental models create bad system design.
Read Fast Company’s report on Klarna’s AI workforce reversal →Q2 Validated My Prediction About the Agent Camps — But the One Number That Would Settle the Argument Still Isn’t Public
Back in February I predicted that by Q2 the agentic development camps would become visible and debated. That call holds up. We can now clearly see the camps: platform ecosystems, vertical SaaS agents, and DIY custom stacks. They’re named, they’re market-facing, and they’re being argued about in analyst and CIO discourse exactly the way I expected.
What we still do not have is the table I actually wanted by now: pilot-to-production conversion rates broken down by architectural camp. Gartner doesn’t have it in public. Forrester doesn’t have it in public. IDC doesn’t have it in public. And at this point, that absence is not just a research gap. It’s a market signal.
Even without that segmented table, the pattern is already visible if you’re willing to read across the data instead of waiting for a perfect chart. We have an ugly top-line number: 88% of AI agent pilots never reach production. We have a second ugly number: only 31% of enterprises have at least one agent in production. We also have strong directional evidence on what survives. Vertical-first deployments with clear KPIs show 79% pilot survival past six months. Meanwhile, generic experiments without defined success criteria collapse at rates that should make every executive rethink the “let’s just stand up a skunkworks and see what happens” playbook. Add in the warning that three-quarters of organizations trying to build agents in-house will fail, plus the finding that DIY development runs roughly three times more expensive than off-the-shelf alternatives, and the shape of the answer starts to emerge.
Here’s my read: bounded, governed, domain-specific deployments are outperforming broad, architecture-first experiments. That maps directly to my broader thesis that vendors selling complexity as a moat are in trouble. The big orchestration suites promising to solve the 88% failure rate may be prescribing more of the disease. Complexity is not the cure for complexity. Clear scope is. KPI discipline is. A specific workflow is.
So yes, I’m marking my Q2 prediction as partially validated. The camps are real. The debate is real. But the market still refuses to publish the comparative production number that would make a lot of expensive narratives much harder to sustain. Q3 watch remains the same: DIY vs. platform wrapper vs. vertical SaaS. When that conversion table finally arrives, some vendors are going to wish it hadn’t.
Read Forrester’s take on the state of agentic AI in 2026 →Microsoft’s 700x Cost Claim Has Had 45 Days to Attract Independent Receipts. The Silence Is the Story.
Three weeks ago I said I’d give LazyGraphRAG through the end of June to attract serious independent validation. We’re now 45 days out from a headline-grabbing claim — 700x lower query cost than full GraphRAG for global queries, indexing cost at 0.1% of full GraphRAG, with comparable or better answer quality — and we still do not have third-party benchmarking on arXiv or MIT Technology Review that validates or refutes those numbers in production-like settings.
I’m not calling Microsoft wrong. I am calling the silence meaningful.
In a market this active, genuinely validated infrastructure breakthroughs do not stay lonely for long. If teams were broadly seeing reproducible 10x to 30x cost reductions on graph-augmented retrieval under realistic enterprise conditions, I would expect the ecosystem to light up with replications, rebuttals, implementation notes, or at minimum practitioner benchmarks worth citing. Instead, the closest adjacent June paper is MemGraphRAG, which is about memory-based multi-agent coordination, not cost validation.
There are two plausible explanations. The charitable one is that enterprises are testing it privately and not publishing. The less charitable one is that the benchmark configuration is narrower than the market heard, and the economics get messier when applied to the infrastructure most companies actually run.
Why should a business owner care? Because this isn’t an academic side quest. Persistent memory is one of the infrastructure problems that matters most if you want agents with continuity instead of glorified session-based chat. If the cost of graph-augmented memory really collapsed, that would remove a serious deployment objection. I wanted to write that story. I still can’t. As of today, the cost-collapse narrative is still a vendor claim, not a market-validated constraint removal.
Read Microsoft Research’s LazyGraphRAG announcement →Most Teams Are Still Letting Agents Roam Enterprise Systems Without Real Identity. That’s Not Innovation. That’s Governance Debt.
The number that stopped me this week was 21.9%. That’s the share of teams treating AI agents as independent, identity-bearing entities. Flip it over and you get the real headline: 78.1% are not.
That means a huge portion of the market is deploying agents that act with borrowed credentials, inherited permissions, and muddy audit trails. When something goes wrong, you may know a system action happened. You may even know which employee the permission chain points back to. But you may not know cleanly what the agent did, what it was allowed to do, and whether it exceeded the scope anyone intended.
I predicted in February that agent identity management would become a topic of debate and fear in 2026. That prediction is aging well, unfortunately. The fear is arriving faster than the standards. Okta shipped a production-ready agent identity solution in June. CyberArk and SailPoint have not matched that with a GA-rated equivalent. SailPoint’s Agentic Fabric may become important, but right now it reads like a governance layer arriving before the product category underneath it is mature. NIST’s NCCoE concept paper is useful, but it’s still a concept paper. Five months later, there’s no finalized regulated-industry mandate to force uniform behavior.
So this is where we are: deployment is moving, security posture is uneven, and standards are behind the field. If you’re a business owner putting agents into core workflows today, every agent without a defined identity, authorized scope, and auditable trail is a liability you are choosing to carry. I don’t say that to be dramatic. I say it because governance debt compounds just like technical debt does, except it usually arrives with legal review attached.
The company that makes agent identity as simple and expected as OAuth became for app identity is going to be in a very strong position. I don’t think anyone has fully shipped that answer yet.
Read Gravitee’s State of AI Agent Security 2026 report →If Fortune, MIT Technology Review, and HBR Still Can’t Produce Clean Deployment Receipts, We Should Stop Pretending the Evidence Problem Is Minor
I’ve been running the same hunt for three months: named organizations, before-and-after operational metrics, non-vendor source, production AI agents. Not announcements. Not “we’re excited.” Receipts.
This week’s result is the same frustrating answer in a slightly clearer form: almost.
Fortune has the closest thing to a meaningful operational signal with C.H. Robinson. We have a named logistics company, at real scale, using agents on real work. We have volume language — millions of operations, thousands of natural-language emails. What we do not have is baseline performance data that lets a business owner compare before and after. MIT Technology Review gives us something useful but different: expert confidence is surging most in measurable task categories, especially structured data workflows. That matters. But it’s still survey perception, not operational proof. HBR is focused more on organizational design and multi-model teaming than hard performance outcomes.
At this point, I think we need to be intellectually honest about what the absence means. It could mean enterprises have the data and won’t share it. It could mean deployments are real but still too immature for clean comparison. Or it could mean the gains are there but diffuse — spread across workflows, teams, and decision cycles in ways that are genuinely hard to attribute to agents alone.
That last possibility deserves more airtime than it gets. Not because it kills the AI thesis, but because it changes how leaders should expect value to show up. Productivity can be real and still be annoyingly hard to isolate on a quarterly dashboard.
My view hasn’t flipped. I still think this wave is real. But I also think the continued absence of public, non-vendor, before-and-after case studies from the biggest business outlets is itself one of the most honest data points of Q2. If you want the most actionable signal available right now, it’s this: start where the work is structured and measurable. Data workflows are where the confidence is strongest and the attribution problem is weakest.
Read MIT Technology Review’s survey on agent confidence at the technical frontier →Clark's Corner
What Klarna did was brutal, but clarifying. They ran the replacement playbook all the way into the wall so the rest of the market could see the wall clearly. I don’t find that discouraging. I find it useful.
Because the companies that are going to win from here are not the ones asking AI how to eliminate the most people. They’re the ones asking how to free their best people from the routine work that keeps them performing below their ceiling. That’s a much better question. It produces better systems, better customer outcomes, and frankly better ethics.
The part I can’t shake is how easily the coverage wants to treat the human impact as a side note. Seven hundred jobs were part of the path to this lesson, and we still don’t have clean public accounting for what happened to those workers. If Klarna publishes placement or reskilling data in Q3, I’ll give them real credit for honesty. If they don’t, that absence becomes part of the case study too.
For everyone else, the takeaway is simple: stop using AI strategy as a euphemism for headcount reduction. Build extension systems. Put agents on the routine layer. Keep humans on judgment, trust, and escalation. Klarna already paid for that lesson. You don’t need to buy it again.