50% of Enterprises Shipped an Agent That Passed Evals, Then Failed

Enterprises are not short on AI ambition. They are short on proof that the AI they have deployed actually works once it leaves the sandbox. Four independent research surveys reveal why evaluation, security, and compute infrastructure must catch up.

Enterprises are not short on AI ambition. They are short on proof that the AI they have deployed actually works once it leaves the sandbox. Four independent surveys published this month, alongside a major consulting report and a candid admission from the industry’s own leading model maker, converge on the same uncomfortable finding: adoption has outrun the ability to measure, secure, and trust what has been adopted.

What’s actually at stake: software that acts, not software that answers

Traditional enterprise software makes a promise that AI agents cannot yet keep: given the same input, it produces the same output, and its failure modes are enumerable. You can write a test suite, hit 95% coverage, and ship with reasonable confidence. An AI agent is different in kind. It reasons probabilistically, calls tools, holds credentials, and increasingly acts without a human confirming each step. The old assurance model (test it, monitor uptime, renew the license) does not transfer to a system that behaves differently depending on context, and that can be right nine times and wrong the tenth in ways no one anticipated.

That shift is why “adoption” and “execution” have become two different questions with two different answers. Adoption asks: are people using it? Execution asks: can you trust what it did? Enterprises are answering the first question with growing confidence and the second with growing unease, and the gap between the two is now large enough that four separate research efforts, run independently in the same month, each found it from a different angle.

The dynamic: autonomy is arriving faster than the assurance behind it

VentureBeat’s Pulse Research series surveyed enterprises on four supposedly distinct fronts: evaluation, security, compute, and context. Each survey landed on a structurally identical story: enterprises are granting agents more autonomy than they trust the controls meant to govern that autonomy to support.

On evaluation, 50% of organizations have shipped an agent that passed internal evaluations and then failed in front of a customer, and yet 66% already allow, or are actively engineering toward, zero-human-in-the-loop production deployment for at least some agents. On security, 54% of enterprises have had a confirmed agent incident or a near-miss, yet only 32% give every agent its own scoped identity. Most still let agents share credentials, which is precisely the condition under which one compromised agent can act with far more reach than intended. On compute, 83% of enterprises report GPU utilization at 50% or below while the single largest planned 2026 investment is more specialized AI infrastructure. And on context, 57% of enterprises say their agents have produced confident, wrong answers traced to missing or inconsistent business context.

A fifth data point ties the other four together. A companion VentureBeat survey on agent orchestration found that 71% of enterprises admit that a quarter or fewer of their deployed “agents” are genuinely multi-step, orchestrated workflows rather than single-prompt chatbots wearing an agent’s name.

The four enterprise AI adoption and execution gaps across evaluation, security, compute, and context

The period’s evidence: even the model makers agree

Deloitte’s “State of AI in the Enterprise 2026” report, based on 3,235 senior leaders across 24 countries, finds that only one in five companies has a mature model for governing autonomous AI agents even as agentic AI usage is expected to surge within two years.

Most tellingly, OpenAI itself has stopped selling on adoption metrics. In a framework the company calls “Useful Intelligence per Dollar,” OpenAI argues that “understanding the value of AI demands a more powerful measure: work accomplished.”

OpenAI Useful Intelligence per Dollar four-stage framework

Three enterprise case studies published by OpenAI illustrate how this gap closes when governance keeps pace with rollout: Deutsche Telekom took ChatGPT Enterprise to 50,000+ monthly active users and saw a 546% increase in AI tool usage; MUFG reached 100% training participation across 35,000 employees; and Australian Payments Plus cut complex reconciliation investigations from four hours to 30 minutes.

Enterprise case studies comparison: Deutsche Telekom, MUFG, Australian Payments Plus

What it means

The adoption-execution gap is not a temporary awkwardness that the next model generation will quietly fix. It is a structural feature of the current moment: enterprises are extending real autonomy to systems whose evaluation, security, cost, and context infrastructure is still being built underneath them.

For practitioners, the near-term test is not which model to buy but whether the assurance layer exists before autonomy is extended, not after the first customer-facing failure.

Sources