Key Takeaways

  • Enterprises knowingly deployed AI agents before building the controls to govern them — and now 57 to 68% plan to rip and replace vendors within a year
  • Most "agents" are chatbots in disguise: 71% of firms say a quarter or fewer can actually complete multi-step work autonomously
  • Two-thirds of enterprises already let agents push code to production on automated evaluations alone, yet only 5% fully trust those evaluations
  • Companies sharing agent credentials suffer security incidents at 63.5% versus 40.9% for those enforcing scoped identity

Enterprises didn't stumble into the agentic era. They sprinted into it with eyes open. VentureBeat Research's five parallel surveys, fielded in June across every layer of the agentic stack, reveal a pattern too deliberate for accident: organizations deployed agents ahead of the controls needed to manage them, and they did it knowingly. Now they are retrofitting at speed. In each of the five control layers measured — identity, evaluation, cost telemetry, context, orchestration — 57 to 68% of enterprises plan to switch or add vendors within twelve months. Roughly a third intend to move within the quarter. That churn rate doesn't signal healthy iteration. It signals a market that sold promises before it built primitives.

The label "agent" is doing heavy lifting. Seventy-one percent of respondents — 81% of whom recommend or decide AI purchases at their companies — say a quarter or fewer of their deployed agents can complete multi-step work on their own. Only 10% report that true agents constitute the majority of what they run. The rest are chatbots wearing a trendier badge. This distinction matters because a single-prompt chatbot with a human reading every answer needs none of the controls the other four reports measure. A genuine multi-step agent needs all of them. Most enterprises cannot say which one they have deployed. That ignorance is not benign; it is the condition under which risk compounds.

Autonomy is outrunning trust in the evaluations that gate it. Two-thirds of enterprises either already allow an agent to push a code or system change to production on automated evaluation results alone, with no human review, or are actively engineering toward that within twelve months. Only 5% fully trust the evaluations that would make that call. Half of enterprises shipped an agent that passed internal evaluations and then caused a customer-facing failure in the past year. The lesson is blunt: before removing human review from any workflow, test evaluations against production outcomes rather than internal benchmarks. Benchmarks lie. Production does not.

Credential sharing is the quiet force multiplier for breach. Sixty-nine percent of companies let at least some agents share credentials — multiple agents operating under one API key or service account. Organizations that allow credential sharing anywhere experienced a security incident or near-miss at a 63.5% rate, against 40.9% at companies where every agent has its own scoped identity. The denominator is small but the signal is loud. The fix is scoped identity for every agent, starting with the ones that touch production systems. Anything less is an invitation.

The context layer remains the most neglected control. Agents draw on business data and definitions when they answer, yet most enterprises treat context as an afterthought — a vector store here, a document dump there — rather than a governed supply chain. Cost telemetry is equally primitive. Few organizations can tell you what a given agent costs per task, per outcome, per day. Without that visibility, the economics of agentic workflows reduce to hope. Orchestration, the control plane that coordinates multi-step work, is where the rubber meets the road. But you cannot orchestrate what you cannot identify, evaluate, or afford.

Vendors recognize the gap. The 57 to 68% switching intent across every layer reflects a buyer's market that knows it was sold vapor. The winners in the next cycle will not be the loudest agent frameworks. They will be the boring infrastructure that makes agents auditable, billable, and containable. Identity providers that issue scoped credentials per agent. Evaluation platforms that correlate benchmarks with production outcomes. Cost meters that attribute spend to autonomous action. Context pipelines that version and govern the knowledge agents consume. Orchestration layers that enforce policy at every handoff.

Enterprises that treat these as sequential projects will lose. The controls are interdependent. An agent with scoped identity but no evaluation telemetry is a loaded gun with no safety. An agent with evaluation but no cost visibility is a budget leak with a report card. The retrofit must be parallel. Boards should demand a control maturity model alongside every agent roadmap. If the model shows gaps, the roadmap pauses. That discipline is the only thing separating an agentic enterprise from a chatbot farm with better marketing.