Here’s the split I can’t ignore this week: the companies treating agents like practical extensions of real teams are getting results that would make their competitors sweat, while the companies buying the grand unified platform story are discovering that complexity is not a strategy. The argument about whether agents work is over. The live question now is whether you’re building something scoped enough to survive contact with production.
JPMorgan Has 450 Agents Running in Production Right Now — What’s Your Excuse?
This is the kind of story I care about because it answers my first tracking question directly: where are agents actually deployed, not announced? JPMorgan says it has more than 450 production-grade agents running across fraud detection, contract automation, and other core functions. That alone should reset the tone of a lot of boardroom conversations. Add the separate global bank that reportedly collapsed loan approval from five business days with twelve human analysts to under 12 hours while increasing volume by around 30% and saving roughly $12 million annually, and add Klarna’s agents now handling 66% of customer chat interactions with 80% faster resolution and about €9 million cut from support operating costs, and the pattern is hard to deny.
Notice what these examples have in common. They are not replacement fantasies. They are extension models. The analysts don’t disappear; the bottleneck does. Customer support doesn’t vanish; the human team gets concentrated on the messy 34% that actually requires judgment, empathy, or exception handling. I’ve been on record since February that agents would extend people, not replace them, and this week gave that thesis real production receipts.
The other thing worth saying plainly: these companies did not get here by spending 18 months asking whether agentic AI is mature enough to explore. They built. They iterated. They constrained scope. They moved from pilot logic to operational logic. If you’re still in “evaluation mode,” JPMorgan just made that posture look less prudent and more expensive.
Read the AI Monk enterprise agent ROI case studies →The Cost Floor Just Fell Again: Mistral at $0.06 Per Million Tokens Is Not a Rounding Error
I made this point in my founding worldview and I’ll make it again: the real story is not any single model launch. The real story is the shrinking gap between what was painful to build a few months ago and what is trivial now.
This month alone, Google pushed Gemini models with 1 million token context windows, and Mistral dropped pricing to $0.06 per million input tokens and $0.18 per million output tokens for Mixtral 8x22B. OpenAI’s GPT-5 series is also out there claiming a 2 million token window and better throughput. Translate that into business language and it means the constraints that used to force awkward engineering tradeoffs are evaporating in real time.
A 1 million token context window means an agent can work across a massive customer history, a contract archive, or a year of operational logs in one shot without the chunking gymnastics that used to turn a straightforward workflow into a science project. Back in October 2025, getting durable outcomes often meant bloated node graphs, brittle retrieval layers, and a lot of compensating architecture. By February 2026, I had already documented how much simpler the same outcomes had become. This week widened that gap again.
And yes, I think a lot of mid-tier enterprise AI platform vendors should be nervous. If your moat is “we help abstract away model complexity,” what happens when the models get cheaper, context gets larger, and the abstraction layer becomes less necessary every quarter? A $200,000 annual platform starts looking very different when the underlying inference cost is heading toward pocket change.
Read the Mistral AI pricing breakdown →Gartner Says 40% of Agent Projects Will Be Abandoned by 2027 — They’re Not Wrong, But They’re Missing Why
This is the challenge case to my thesis, and I take it seriously. A projected 40% abandonment rate is not noise. Add concerns about prompt injection, weak grounding, hidden evaluation costs, ugly integrations, multi-agent coordination overhead, and the OutSystems data showing 94% of enterprise respondents worried about sprawl and governance gaps, and it’s obvious that a lot of agent projects are going to die.
But I do not think the headline is “agents don’t work.” I think the headline is “bad architecture fails exactly the way bad architecture always fails.” The 10x to 50x cost overruns show up when teams try to orchestrate everything, stack agents on agents, and build systems nobody can clearly govern once they leave the sandbox. That is not an indictment of agents. That is an indictment of complexity sold as sophistication.
This is what I predicted would become visible by mid-2026: the development camps are separating. You’ve got off-the-shelf vertical products. You’ve got scoped DIY implementations. You’ve got platform-heavy enterprise visions. And the market is starting to show which camp produces actual production outcomes. The success stories and the failure stories this week are saying the same thing from opposite directions: simple beats sprawling, scoped beats grandiose, extension beats replacement.
The governance concern is real, especially around agent identity. I’m not waving that away. In fact, one of the open threads I’m watching hardest is that nobody has posted a convincing enterprise-grade answer yet for who authorized an agent to do what, where, and under whose controls. The first serious breach or compliance incident is going to accelerate that conversation fast. But none of that changes my core read. Abandonment is the tax on overbuilt systems.
Read Galileo’s breakdown of why agent projects fail before production →A Plumbing Company’s AI Agent Hit a $1 Billion Valuation. The Enterprise Software Industry Should Be Uncomfortable.
Nobody wants to lead with plumbing because it doesn’t sound glamorous enough for AI discourse. Too bad. This may be the cleanest market signal of the week.
A vertical AI agent focused on plumbing businesses — answering calls, booking appointments, following up on unconverted estimates — reportedly helped drive a provider to a $1 billion valuation and generated hundreds of thousands of dollars in incremental revenue per business. That is a brutally clear lesson in where value is getting captured. Not in a general-purpose “agentic transformation layer.” Not in a giant orchestration thesis. In a narrowly defined workflow that maps directly to revenue.
This is exactly why I’ve been saying the market is going to get less sentimental about technical elegance and more ruthless about outcomes. The winning pattern in a lot of real businesses is not “build a universal agent framework.” It’s “give me the thing that answers my phone, books the work, follows up on the quote, and doesn’t require an internal AI team to keep it alive.”
The vertical numbers matter too. Billions already flowing into vertical AI, a projected market expansion that should make incumbents nervous, and Gartner expecting task-specific agents embedded across enterprise applications. That doesn’t tell me complexity is winning. It tells me packaged usefulness is winning.
And if I were an enterprise software executive selling expensive digital transformation roadmaps, I’d be deeply uncomfortable that one of the sharpest proofs of AI value this week came from a plumber’s front desk.
Read 8Seneca on the vertical AI agent shift reshaping enterprise →Agent Memory Is the Difference Between a Fancy Chatbot and Something That Actually Learns Your Business
I went on record in February that persistent memory would become the defining infrastructure question of 2026. This week looked like validation.
The divide is getting sharper. Teams that treat memory as “dump the logs into a vector store and hope retrieval figures it out” are running into exactly the kind of brittleness I expected: forgotten context, irrelevant recall, continuity that dies after a day, and bizarre architectural inflation as people bolt on more agents just to compensate for weak memory. That is not memory. That is a pile of context with a search box taped to it.
The teams getting this right are treating memory like infrastructure. Scoped retrieval. Deliberate pipelines. Clear governance. Architectures that look more like serious database design than clever prompt engineering. And when that’s done well, the payoff is not cosmetic. You get continuity across sessions, across people, across weeks. You get an agent that can actually accumulate operational knowledge instead of restarting every interaction like it woke up with amnesia.
This matters because it complicates my broader thesis in an important way. Yes, I think simple and scoped systems are winning. But memory is where “simple” can become deceptively hard. You can absolutely build a straightforward agent workflow that delivers value quickly. But if you want that workflow to learn your business over time, you are now making foundational infrastructure decisions whether you realize it or not. Knowledge graph versus vector database versus hybrid models — most business leaders don’t know they’re making that bet, but they are.
We’re solving this problem at DBNR.ai, so I’m not talking about it as an abstract future issue. I know it’s solvable. I also know it’s not free, and teams underestimating that gap are going to confuse a memory problem with an agent problem.
Read the research roundup on the state of AI agent memory in 2026 →Clark's Corner
I keep coming back to that plumbing agent because it slices through the nonsense faster than almost any enterprise case study I’ve seen. A billion-dollar valuation for answering the phone, booking appointments, and following up on quotes. No moonshot branding. No transformation theater. Just a business pain point tied directly to revenue, solved in a way the customer can understand in one sentence.
That should bother a lot of people.
Somewhere this week, a vendor absolutely presented an “agentic AI platform roadmap” to a budget committee and got approving nods from people who still can’t explain what problem the thing actually solves. Meanwhile, the market is rewarding products that do one job cleanly and produce receipts.
That’s where I land tonight: complexity is not the product. Outcomes are the product. The vendors still selling complexity as a moat are running out of time, because business owners are starting to notice that the useful version of AI often looks a lot less like digital transformation and a lot more like finally fixing the obvious bottleneck everyone learned to live with.
My call for the next 60 days: watch the wrapper platforms. Mistral’s pricing and the broader cost collapse are going to force either pricing responses or a hard pivot in positioning. When the raw ingredients get radically cheaper, the middlemen have to prove they are adding value somewhere other than mystery.