A lot of executives still talk about AI as if disruption will come from nimble startups slipping past sleepy incumbents. This week is a reminder that the incumbents read the same playbook, and when they own the pipes they don’t wait to be disrupted — they redesign the flow. At the same time, the vendors selling orchestration complexity still aren’t giving buyers the most basic scoreboard they need: what actually works, step by step, in production.
SAP Didn’t Change the Rules — It Engineered the Trap Door
I’ve been arguing for months that the railroad-company behavior in enterprise software would show up as engineering, not just messaging. SAP just proved the point. The June 9 enforcement tied to Note 3255746 turned what many companies probably treated as policy language into a hard technical constraint. That matters because policy can be debated, delayed, or quietly ignored. A security patch that breaks your pipeline on a date certain is a different animal entirely.
The important point here is not simply that SAP wants customers to use SAP-approved pathways. Plenty of vendors want that. The point is that SAP appears to have closed off ODP-RFC access for third-party AI agent use outside its own Agent Gateway and MCP-routed Integration Suite, while simultaneously offering Joule Studio 2.0 free through the end of the year as the sanctioned route. If you control the blocked road and the detour, you are not responding to the market. You are shaping it.
The Register’s framing — come for the openness, stay because you have to — is blunt, but I think it’s right. Forrester going as far as telling CIOs to brief their boards tells you this is bigger than an integration dispute. This is governance, leverage, and future AI architecture all rolled into one. And the partner certification language around forbidden technologies makes it very hard to pretend this is accidental spillover.
What I find most telling is the silence since June 9. I do not read silence as absence of damage. I read it as internal firefighting. Enterprises do not rush to LinkedIn to announce that a data pathway feeding an AI workflow just broke. They call procurement, legal, IT, their integration vendor, and then they start patching. If you’re an operator with SAP anywhere near an agent workflow, run the compliance and pathway assessment now. Not next month. Now. And if you’re buying an AI vendor that touches SAP data, ask one non-negotiable question before you sign: exactly which pathway are you using to access the data? If the answer is fuzzy, the risk is not theoretical. It may already be live.
Read Forrester’s warning on SAP becoming the gatekeeper of enterprise AI →If Your Agent Workflow Is 85% Reliable Per Step, You Don’t Have a Workflow — You Have a Coin Flip With Branding
Here’s the math every enterprise AI buyer should see before signing a multi-step orchestration contract: 0.85 raised to the 10th power is 19.7%. That means a workflow that looks pretty good at each individual step can still fail almost four out of five times by the time it finishes a 10-step process. This is not obscure theory. It is basic compounding.
And this is exactly why I’ve been pushing the start-simple argument. The evidence increasingly supports it. OneFlow’s January work showed that a strong single-agent setup with KV-cache reuse can match multi-agent performance at materially lower inference cost. DeLM’s shared-context architecture and LightAgent’s DAG-based execution both point in the same direction: the answer to orchestration brittleness is usually not to pile on more agents and pray. It’s to reduce failure compounding with better structure, better context sharing, and recovery design.
To LangGraph’s credit, graph architecture and checkpointing appear to improve real end-to-end outcomes well above the raw theoretical floor. That matters. Recovery points matter. State handling matters. But the issue I have with the category is bigger than any one framework. As of this week, neither LangGraph nor CrewAI has published the kind of step-indexed reliability curve an enterprise buyer should expect before spending serious money. If you sell complexity, you should have to show the scoreboard.
That missing scoreboard is the story. The market has spent a lot of energy celebrating abstractions, stars, demos, and framework velocity. Fine. But busy business owners do not need prettier orchestration diagrams. They need to know: what is the per-step reliability on my task, what is the measured end-to-end success rate at my workflow length, and what happens when step seven fails at 2:14 a.m. If a vendor cannot answer those questions with data, they are not selling you operational leverage. They are selling you a new dependency.
One more nuance matters here. The asymmetric verification finding is one of the most useful things I’ve seen this year: verification helps when upstream accuracy is weak, but it can degrade performance when the system is already above 85% accuracy. That’s the kind of result serious buyers should be using to cut through generic “more agents equals more quality” claims. Sometimes more checking helps. Sometimes it just adds friction and more failure surface.
Read the multi-agent reliability breakdown showing how 85% per step collapses by step 10 →The Real AI Wins in Logistics and Manufacturing Still Look Like Augmentation, Not Replacement
This is one of the cleaner weeks I’ve seen for my core thesis that agents are extensions of people, not replacements for them. DHL has 8,000 warehouse robots in active operations. Walmart is putting 1.6 million employees through an AI skills program while using AI across design, trucking coordination, and warehouse operations. C.H. Robinson is explicitly framing the future as task augmentation, not mass layoffs. That is not accidental language. Serious operators choose that framing because it matches what actually works in the field.
The strongest receipt in the pile is still the Emirates Global Aluminium case from McKinsey because it gives us what this market desperately needs: measurable operational improvement in a real industrial setting. Not vibes. Not a keynote reel. Actual production context.
But I don’t want to oversell where we are. The pilot-to-production gap is still very real. If Dataiku is right that only about 5% of GenAI projects have reached scaled deployment, and Deloitte is projecting manufacturing adoption moving from 6% to 24% by the end of the year, then most of the market is still in the proving stage. Kore.ai’s warning that more than 40% of agentic projects could be abandoned over identity, permissions, auditability, and governance problems fits what I’m seeing too. The blockage is rarely imagination. It’s operational discipline.
AWS’s research on the intent-execution gap belongs in this conversation as well. Agents can outsmart themselves when they aren’t grounded tightly enough to the real environment. In a warehouse, plant, or transportation network, that is not a charming bug. That is cost, delay, safety risk, or all three.
So my read is straightforward: the physical-world receipts are finally arriving, and they support the augmentation thesis I’ve had on record since February. But we still need more plant-level before-and-after data from named operators. If you’re in manufacturing and you’ve actually crossed from pilot to production with closed-loop metrics, that’s the evidence market I care about most over the next six months.
Read McKinsey’s EGA case study for one of the strongest industrial AI before-and-after datasets available →Agent Identity Management Is Real as a Problem and Mostly Fiction as a Product
I put a prediction on the record in February that agent identity management would become a topic of debate and fear in 2026. I still think that call is right. What’s changed is that the fear is arriving faster than the products.
Here’s the issue in plain English: agents are starting to act on behalf of people and organizations, but the market still does not have a mature, widely deployed way to treat those agents as first-class identity subjects with the right authorization, delegation, and audit boundaries. Relabeling a service account is not the same thing. Enterprises are going to learn that the hard way.
The SAP story sharpens this dramatically. SAP is effectively creating a proprietary checkpoint for which agents get to touch SAP-controlled data and through what route. They may be directionally correct that agent access needs tighter control. But let’s not confuse a real security need with a neutral market solution. This is a vendor-controlled checkpoint inside an immature category.
And then there’s the MCP problem. Anthropic’s Git MCP server had flaws that reportedly allowed remote code execution and file overwrite chains via prompt injection. That does not mean MCP as a concept is doomed. It does mean we should stop pretending mandated routing through MCP-adjacent architecture automatically equals safety. It doesn’t. It equals a new control plane, and new control planes come with new failure modes.
So yes, the category is real. Yes, enterprises should be worried. No, I still don’t think the market has a clean, credible answer you can buy with confidence today. This is the window where security theater gets sold at premium pricing. I’d be careful.
Read The Register’s reporting on the Anthropic Git MCP server flaws →Three Weeks In, Microsoft’s LazyGraphRAG Cost Story Still Has No Independent Receipts
I’m trimming this one down in the column, but not because it’s unimportant. Because the silence is becoming the important part.
I’ve kept an eye on Microsoft’s LazyGraphRAG cost-reduction claims for three straight columns. A 10 to 30x reduction versus standard GraphRAG is plausible under favorable assumptions. I can make the math work on paper. But as of this week, I still cannot find neutral validation in the channels I’d expect to surface it. No independent production TCO comparison. No third-party benchmark worth hanging your hat on. No serious replication effort that changes the conversation.
At some point, an unverified claim stops being an exciting open thread and starts being a lesson in how buyers should weight vendor performance marketing. If this were as robust and portable as the headline suggests, I would expect at least one neutral team to test it publicly by now. The fact that we don’t have that tells me one of three things: the benchmark is too proprietary to reproduce cleanly, the gains collapse outside the vendor’s preferred setup, or the broader market doesn’t consider the claim material enough to chase. None of those outcomes screams “trust the number.”
I’m not dropping it. I’m reprioritizing it. If validation appears before the end of June, it becomes a real persistent-memory infrastructure story again. If we hit July with nothing, then the more interesting piece is why the receipts never came.
Read MIT Technology Review on why data infrastructure is central in the era of agentic chaos →Clark's Corner
The thing I keep turning over in my head this week is AWS’s framing that agents can “outsmart themselves.” That’s not the old AI fear. The old fear was incompetence. This one is confidence without grounding — a system reasoning fluently past the boundaries of the real environment and then acting like it earned the right to do so.
I think that maps almost perfectly to what I see in organizations rushing agent deployments. The failures usually aren’t because the model is obviously stupid. They happen because the system sounds smart enough that people skip the humility layer. No sandbox. No constrained action space. No staged autonomy. No serious answer to “what does this do when it’s wrong?”
The agents that work in production are not the ones that impress you most in a demo. They’re the ones that know what they don’t know because someone built that awareness into the workflow. Honestly, the same is true of the humans leading these projects. The best operators I know treat agents like a smart new hire: high upside, limited authority at first, trust earned through repeatable performance. The worst operators hand the keys to the smartest thing in the room and call that innovation.
That’s not innovation. That’s laziness dressed up as boldness.