Enterprise confidence in automated evaluations for AI agents appears to be rising faster than the reliability signal itself, according to VentureBeat. The outlet reports that, across 108 enterprises, the share of organizations that fully trust automated agent evaluation nearly tripled in one month, from 5% in June to 13% in July. At the same time, VentureBeat says the failure rate those evaluations are supposed to predict did not move. That is the central tension for operators: automated evals are meant to provide a scalable way to decide when agents can act with less oversight. But in the figures VentureBeat cites, trust in the evaluation layer increased even though the measured failure rate did not improve. VentureBeat’s headline adds a sharper claim: enterprises that were “burned by a bad eval” are reportedly the most likely to remove humans from the loop, not the least. The provided summary does not include the breakdown behind that conclusion, so the result should be treated as an early reported finding rather than a fully corroborated benchmark. A separate item in the cluster, a first-person essay by Brent Fitzgerald surfaced via Hacker News, approaches the same human-in-the-loop question from an individual rather than enterprise angle. Fitzgerald writes about returning from a break and finding numerous paused agent sessions and AI chats, concluding that some of his AI use had become a habit-forming crutch rather than a clear productivity gain. The two items do not substantiate each other. VentureBeat provides the enterprise data point; the essay supplies personal reflection about AI habits and cognitive delegation. Taken together, they point to a live operating question rather than a settled answer: when does adding or removing a human improve the system, and when does it simply make automation feel more controlled than it is? Who benefits: The reported shift most directly benefits organizations willing to rely more on automated evaluations to scale agent deployments. The supplied evidence is too thin to identify specific vendor winners. Who's exposed: Enterprises that remove humans from the loop despite unchanged failure rates are exposed to reliability and governance risk. Teams using eval results as a proxy for production readiness should be careful not to treat rising confidence as the same thing as better performance.