According to a survey of 157 large enterprises, companies are increasingly trusting AI agents while placing less trust in the evaluations meant to constrain their autonomy. Half have already deployed agents to production that passed internal evaluations but failed in real-world conditions with customers. Only one in twenty fully trusts automated evaluations.

Despite this, two-thirds of organizations are allowing or actively working to deploy changes to agents into production based on automated evaluations without human involvement. This creates risk, as evaluations may not reflect real-world consequences, potentially leading to errors in agent performance.