1
1
A new study by VentureBeat Pulse Research reveals a critical "evaluation gap" in the deployment of AI agents across enterprises. This gap represents the growing disparity between the increasing autonomy organizations are affording their AI agents and the diminishing trust they place in the evaluation systems designed to ensure their reliability and prevent failures. The research indicates a concerning trend where half of all organizations have already deployed an agent that passed internal evaluations only to fail a customer in production. Despite this, a significant two-thirds of enterprises are either already allowing or actively engineering towards deploying agent changes to production based solely on automated evaluation, with no human intervention.
The study, part of VentureBeat’s ongoing Pulse Research series, specifically focused on the "Agentic Reliability & Evals tracker" to examine how technical leaders measure agent performance. It delved into the reliability and evaluation platforms they utilize, their selection criteria, the level of trust they place in these systems, common production failures, and the extent to which agents operate without human oversight. The findings paint a picture of an industry grappling with the rapid advancement of AI agent capabilities without a correspondingly mature and trusted assurance framework.
Methodology and Scope of the Research
VentureBeat conducted this survey in June 2026, gathering responses from 157 qualified enterprise respondents with 100 or more employees. As a single-wave study, the report provides a cross-sectional view rather than month-over-month trends. Where questions allowed for multiple selections, shares may sum to more than 100%.
The respondent pool was senior and buyer-credible, with 38% identified as final decision-makers for AI purchases and an additional 34% as recommenders or influencers. Key roles included product and program managers (15%), consultants and advisors (10%), directors of engineering/IT (8%), and CIOs/CTOs/CISOs (8%), alongside a 37% "Other" category. The sample leaned towards mid-market organizations, with 37% having 100-499 employees and 27% having 500-2,499. Larger enterprises (2,500-9,999 employees at 20%, 10,000-49,999 at 10%, and 50,000+ at 6%) were also represented. Technology/Software dominated the industry representation at 23%, followed by Retail/Consumer (15%), Healthcare/Life Sciences (12%), and Manufacturing (10%).
The study emphasizes that with 157 respondents, the sample provides a directional signal rather than a precise measurement, being self-selected and not a probability sample. Its mid-market skew means it offers insights primarily from organizations actively establishing agent evaluation practices. It’s important to note that this survey was rebuilt for the June wave from an earlier "LLM observability and evaluations" survey, thus no comparisons are made to previous data due to differing questions and sample.
Key Findings Unveiling the Evaluation Gap:
1. A Passing Evaluation Is Not a Working Agent
The most defining revelation of the report is that half of organizations (50%) have, in the past year, deployed an AI agent or Large Language Model (LLM) feature that successfully passed their internal evaluations, only to subsequently cause a customer-facing failure. This includes incorrect outputs, broken workflows, or quality incidents. A quarter of these organizations (24%) reported experiencing such failures more than once, highlighting a recurring disconnect between internal assessment and real-world performance. Only 36% reported no such production failures, while 8% don’t run pre-deployment evaluations at all, and 6% don’t track the root cause sufficiently. This demonstrates a precise and costly problem: evaluations are certifying agents that are, in fact, not ready for production.
2. Near-Universal Distrust in Automated Evaluation
Trust in automated evaluation systems is remarkably low. Only 5% of organizations fully trust automated evaluation today, meaning a staggering 95% identify significant limitations. The most frequently cited weakness, by 29% of respondents, is the poor alignment of evaluations with real-world outcomes, directly explaining the failures highlighted in Finding 1. Other significant concerns include evaluation bias or inconsistency (21%), a lack of explainability (18%) – hindering understanding of evaluation verdicts – and data leakage or privacy concerns within the evaluation process itself (17%). Tooling immaturity was cited by 11%. This lack of trust indicates that the very tests meant to certify AI agents are not yet deemed reliable enough to do so.
3. Autonomy’s Ascent Outpaces Assurance
Despite the widespread distrust in automated evaluation, the trajectory towards greater agent autonomy is undeniable. Two-thirds of organizations (66%) are already moving towards or implementing zero-human-in-the-loop deployment. Specifically, 34% currently permit fully automated deployment for specific low-risk agents or changes, while another 33% are actively engineering their pipelines to allow it within the next 12 months. Only 22% rule out fully automated deployment for the foreseeable future. This presents a striking paradox: enterprises are removing human checks and allowing automated evaluations to gate production autonomously, even as they acknowledge these evaluations don’t reliably reflect reality. This accelerating autonomy, unsupported by robust assurance, risks scaling the false-confidence failures identified in Finding 1. Interestingly, larger enterprises appear slightly more advanced in this zero-human review path (70% vs. 64% for smaller companies) and more prone to evaluation-passing agent failures (54% vs. 48%), challenging the assumption that larger, regulated organizations maintain human oversight longer.
4. A Fragmented and Nascent Evaluation Stack
The market for agent reliability and evaluation platforms is characterized by immaturity and fragmentation, lacking a clear leader. The most common primary tools are the model providers’ native evaluations, with OpenAI Developer Platform native evals/traces leading at 17%, closely followed by Anthropic Claude Console native evals at 13%. Strikingly, tied with OpenAI at 17%, is the response that organizations use no dedicated agent reliability or evaluation tooling at all. Specialist evaluation vendors such as Confident AI (DeepEval) account for 12%, Braintrust for 8%, while others like LangSmith, Weave, Promptfoo, Langfuse, and Arize register in low single digits. Additionally, 11% rely on custom in-house tooling. This landscape suggests that many enterprises are evaluating agents with basic provider-native tools, self-built scripts, or, concerningly, nothing specific.
5. Production Monitoring’s Critical Blind Spot
The study highlights a significant blind spot in how organizations monitor AI agents in production. There’s a crucial distinction between monitoring whether a system is functioning (e.g., uptime, response speed, cost, errors) and whether its output is correct (e.g., accurate answers, correct actions, policy adherence). While a confidently wrong answer is invisible to ‘functioning’ metrics, the research shows that 51% of organizations primarily monitor only functionality (25% for transaction trace logging, 25% for gateway infrastructure tracking). Only 23% run inline quality assertions for real-time output correctness checks on live traffic. Another 10% rely on ad-hoc review, and 17% don’t track this. This means roughly three-quarters of organizations lack automated, real-time evaluation of output correctness in production, taking agent answers on faith. This blind spot is the runtime equivalent of the pre-deployment gap, allowing deployed agents to err unseen.
6. Pragmatic Choices: Cost, Integration, and Consistency
Enterprises are pragmatic in their approach to selecting and measuring the success of evaluation tooling. The primary factors influencing vendor choice are the cost of evaluations (28%), ease of integration (27%), and evaluation accuracy (24%). Broader observability (13%) and vendor roadmap (4%) are less significant. When it comes to defining success, evaluation consistency – achieving the same verdict for the same behavior every time – is the most cited goal (36%). This far outweighs speed of experimentation (19%), reduction in failures or regressions (18%), production visibility (13%), or compliance (11%). The emphasis on consistency directly addresses one of the top trust limitations (bias and inconsistency) identified in Finding 2. Overall satisfaction with current tooling averages a moderate 3.8 out of five across key metrics.
7. A Dual Investment Strategy: Humans and Observability
Despite the push towards automation, investment trends reveal a hedging strategy. The largest planned investment over the next year for reliability and evaluation is in production observability tooling (30%). Crucially, the second-largest planned investment is in human review workflows (26%). This outpaces investment in automated evaluation pipelines (16%), safety and policy evaluation (20%), and only 8% report no budget increase. This suggests a subtle contradiction: while two-thirds of enterprises are engineering humans out of the deployment decision, more plan to increase spending on human reviewers than on the automated systems meant to replace them. Organizations are simultaneously building towards autonomy and investing in closer oversight, maintaining human availability for critical judgments that automated evaluations cannot yet reliably make.
8. Impending Tooling Reshuffle
The evaluation market is poised for significant change. A clear majority (64%) of organizations intend to adopt a new, additional, or replacement evaluation platform within the next 12 months, with 31% planning to do so within the next three months. Only 36% have no plans to change. The platforms currently under consideration include Confident AI’s DeepEval (leading at 20%), OpenAI’s native evals (13%), and Braintrust (9%). Given the current reliance on provider-native tools or no dedicated tooling at all (Finding 4), this signals a nascent wave of dedicated tooling adoption rather than mere vendor defection. The open question remains which platforms will earn trust in a market where automated evaluation is currently viewed with such skepticism.
The Bottom Line: An Evaluation Gap That Autonomy Will Widen, Not Close
In conclusion, the research unequivocally shows that organizations with 100 or more employees are granting AI agents a degree of independence that their current evaluation systems are ill-equipped to support. Half of these organizations have experienced agent failures in production after passing internal checks, and almost none fully trust automated evaluation, primarily due to its misalignment with real-world outcomes. Furthermore, most production monitoring focuses on system functionality rather than the correctness of agent outputs. Yet, two-thirds of enterprises are moving towards fully automated deployments.
The vendor market for evaluation tools is early and unconsolidated, with many relying on basic native tools or nothing dedicated. Encouragingly, future investment is directed towards enhanced observability and, notably, human review, suggesting enterprises recognize this gap even as they push for automation. This study, while directional and skewed towards the mid-market, clearly indicates that autonomy is being granted based on evaluations that the very people granting it do not yet trust. The evaluation gap is not merely a matter of needing more tests; it is a fundamental challenge of creating evaluations that accurately reflect reality and can be trusted to govern AI agent behavior. The critical question for the future is whether assurance mechanisms will catch up to the pace of autonomy, or if false-confidence failures will escalate from customer incidents to autonomously deployed, systemic issues.