AI Agents in Production
Despite research enthusiasm for sophisticated multi-agent orchestration, real-world production deployments of AI agents are characterised by simplicity, controllability, and conservative design — with evaluation and reliability as the dominant challenges.
What It Is
The UC Berkeley “Measuring Agents in Production” study (December 2025) provides the first large-scale systematic study of AI agents operating in real organisational environments: 306 practitioners surveyed, 20 in-depth case study interviews, across 26 industries. It documents the gap between research-context agent design and what organisations actually deploy.
The Google DeepMind “Intelligent AI Delegation” paper (February 2026) complements this with a theoretical framework for how agents should decompose and delegate tasks — addressing the orchestration challenge that production teams are grappling with.
Why It Matters (for Organizations)
AI agents — systems that can take actions autonomously over extended sequences to complete goals — are the primary delivery mechanism through which AI capability translates into organisational productivity. The Berkeley study establishes the practical baseline: what problems are organisations actually solving with agents today, what approaches work, and where the friction is. This is essential knowledge for organisations deciding where and how to deploy agents.
The delegation framework from DeepMind addresses a key unsolved problem in multi-agent systems: how to safely and effectively decompose complex tasks across agents (and between agents and humans), transfer appropriate authority at each step, and handle failures gracefully. As agents become more capable and are deployed in more complex workflows, intelligent delegation becomes the critical engineering challenge.
Evidence & Examples
- Top reasons organisations build production agents: increasing productivity, reducing human task-hours, automating routine labour, increasing client satisfaction, reducing human training requirements (
2512.04123v1.pdf) - Production agents are “typically built using simple, controllable approaches” — not the sophisticated multi-agent architectures dominant in research (
2512.04123v1.pdf) - Top development challenges cited by practitioners: evaluation and measurement of agent performance, handling edge cases, maintaining reliability across diverse inputs, and managing user trust and expectations (
2512.04123v1.pdf) - DeepMind argues existing delegation methods rely on “simple heuristics” that cannot adapt dynamically to environmental changes or handle unexpected failures — a significant limitation for high-stakes deployments (
2602.11865v1.pdf) - The DeepMind delegation framework covers: task allocation, transfer of authority and accountability, clear role and boundary specification, clarity of intent, and trust establishment between delegator and delegatee — applicable to both human-to-AI and AI-to-AI delegation (
2602.11865v1.pdf) - DeepMind’s framework is explicitly positioned for “the emerging agentic web” — an infrastructure-level framing that anticipates AI agents as participants in broader networked systems, not just internal tools (
2602.11865v1.pdf)
Tensions & Open Questions
- Evaluation as the unsolved problem: The most commonly cited challenge in the Berkeley study is measuring whether AI agents are actually working well. Unlike traditional software (where outputs are deterministic), agent outputs vary, failure modes are subtle, and success metrics require domain expertise to define. This evaluation gap is the primary barrier to confident production deployment at scale.
- Simple vs. capable trade-off: The preference for “simple, controllable” approaches in production creates a tension with the more capable agentic architectures available. Organisations are trading capability for reliability — which may be appropriate for current risk profiles but may leave significant value on the table.
- Human oversight design: As agents take on more complex and consequential tasks, the question of how humans stay appropriately involved (not over-supervising, not under-supervising) becomes critical. The delegation framework addresses this architecturally, but the organisational and cultural norms for appropriate human-agent oversight are not yet established.
- Trust calibration: Users and operators tend to either over-trust agents (accepting outputs without sufficient scrutiny) or under-trust them (re-checking every output, negating efficiency gains). Building correctly calibrated trust is as much a change management challenge as a technical one.
- Enterprise deployments and the move to sophisticated architectures (web search, April 2026): Only 2% of organisations have AI agents at full scale; 11% in production (Gartner, 2026). Gartner projects 40% of enterprise apps will embed AI agents by end of 2026 (up from <5% in 2025). The dominant pattern enabling reliable sophisticated deployment is supervisor-worker orchestration: a manager agent decomposes intent, routes to specialist agents, synthesises results. Feedback-loop patterns (a reviewer agent checking another agent’s output) significantly reduce hallucinations. Case studies: IBM — $3.5B cost savings, 50% productivity increase enterprise-wide; BCG global biopharma — content localisation 2 months → 1 day; Klarna — AI agents handling support for 85M users at 80% faster resolution; Uber — code migration agents (months → days). Anthropic internal research: multi-agent architecture outperformed single-agent Claude Opus benchmarks by 90.2%. Key transition: “human-in-the-loop” is giving way to “human-on-the-loop” — humans supervise rather than approve every decision. Prompt injection attacks succeed 84% of the time in agentic systems (OWASP Top 10 for LLMs 2025) — governance/security layer is now a distinct architectural requirement. [source needed — MachineLearningMastery 7 Agentic AI Trends 2026; AaiNova Enterprise Architecture Guide 2026]
- ⚠️ POTENTIAL CONTRADICTION — enterprise adoption figures: definitional divergence (web search, April 2026):
Gartner (2026)cited in this article: 2% of organisations have AI agents at full scale; 11% in production. Multiple 2026 enterprise surveys (Joget/Gartner IDC, MasterOfCode, SQ Magazine) now report 51% of enterprises have AI agents running in production and 85% have implemented or plan to by end of 2026. The gap (11% vs. 51%) likely reflects definitional divergence: Gartner’s 2% “full scale” and 11% “in production” may apply a strict definition (autonomous multi-step agents); the 51% figure likely includes simpler deployments (chatbots, copilots, single-step automation). Multi-agent architectures specifically have grown 327% in under 4 months, suggesting the sophisticated layer is also scaling rapidly. The practical upshot: simple agent deployments are now mainstream; sophisticated multi-agent orchestration is scaling but still minority. [source needed — web search: masterofcode.com, joget.com, sqmagazine.co.uk AI agent statistics 2026] - 🔴 TODO (SUBSTANTIALLY NARROWED): The trigger for moving from simple to sophisticated is now documented: supervisor-worker patterns solve reliability; feedback loops address the evaluation challenge; scale justifies investment. Remaining gap: rigorous comparison study between simple vs. multi-agent approaches in the same organisation over time.
Related Concepts
Workflow Redesign Around AI · AI Delegation and Multi-Agent Systems · Agentic AI Fundamentals · Skill Partnerships Human-AI · How AI Agents Actually Work in Production · When AI Teams Think Together, Not Just Together
AI Agents in de Praktijk
Ondanks de wetenschappelijke fascinatie voor geavanceerde multi-agent orkestratie worden AI agents in de praktijk gekenmerkt door eenvoud, controleerbaarheid en conservatief ontwerp — met evaluatie en betrouwbaarheid als dominante uitdagingen.
Wat Is Het
Het UC Berkeley-onderzoek “Measuring Agents in Production” (december 2025) is de eerste grootschalige systematische studie naar AI agents in reële organisatieomgevingen: 306 ondervraagde practitioners, 20 diepgaande casestudy-interviews, verspreid over 26 sectoren. Het documenteert de kloof tussen agent-ontwerp in onderzoekscontexten en wat organisaties daadwerkelijk inzetten.
Het Google DeepMind-paper “Intelligent AI Delegation” (februari 2026) vult dit aan met een theoretisch raamwerk voor hoe agents taken moeten opsplitsen en delegeren — een antwoord op de orkestratie-uitdaging waarmee productieteams worstelen.
Waarom Het Relevant Is (voor Organisaties)
AI agents — systemen die autonoom een reeks acties kunnen uitvoeren om doelen te bereiken — zijn het primaire mechanisme waarmee AI-capaciteit wordt omgezet in organisatorische productiviteit. Het Berkeley-onderzoek stelt de praktische basislijn vast: welke problemen lossen organisaties vandaag daadwerkelijk op met agents, welke aanpakken werken, en waar zit de weerstand? Dit is essentiële kennis voor organisaties die beslissen waar en hoe ze agents inzetten.
Het delegatieraamwerk van DeepMind pakt een belangrijk onopgelost vraagstuk in multi-agent systemen aan: hoe complexe taken veilig en effectief worden verdeeld over agents (en tussen agents en mensen), hoe op elk stap de juiste bevoegdheid wordt overgedragen, en hoe mislukkingen worden opgevangen. Naarmate agents capabeler worden en in complexere workflows ingezet worden, wordt intelligente delegatie de bepalende engineering-uitdaging.
Bewijs & Voorbeelden
- Voornaamste redenen waarom organisaties productie-agents bouwen: productiviteitsverhoging, vermindering van menselijke werktijd, automatisering van routinearbeid, hogere klanttevredenheid en minder behoefte aan menselijke training (
2512.04123v1.pdf) - Productie-agents worden “doorgaans gebouwd met eenvoudige, controleerbare aanpakken” — niet de geavanceerde multi-agent architecturen die in onderzoek domineren (
2512.04123v1.pdf) - Grootste ontwikkeluitdagingen volgens practitioners: evaluatie en meting van agentprestaties, het afhandelen van randgevallen, het waarborgen van betrouwbaarheid bij uiteenlopende invoer, en het managen van gebruikersvertrouwen en -verwachtingen (
2512.04123v1.pdf) - DeepMind stelt dat bestaande delegatiemethoden steunen op “eenvoudige vuistregels” die zich niet dynamisch kunnen aanpassen aan omgevingsveranderingen of onverwachte storingen — een significante beperking voor inzet met hoge risico’s (
2602.11865v1.pdf) - Het DeepMind-delegatieraamwerk omvat: taakverdeling, overdracht van bevoegdheid en verantwoordelijkheid, duidelijke afbakening van rollen en grenzen, helderheid van intentie, en vertrouwensopbouw tussen delegeerder en gedelegeerde — toepasbaar op zowel mens-naar-AI- als AI-naar-AI-delegatie (
2602.11865v1.pdf) - DeepMind’s raamwerk is expliciet gepositioneerd voor “het opkomende agentische web” — een infrastructurele framing die AI agents beschouwt als deelnemers in bredere netwerksystemen, niet slechts als interne tools (
2602.11865v1.pdf)
Spanningen & Openstaande Vragen
- Evaluatie als onopgelost vraagstuk: De meest genoemde uitdaging in het Berkeley-onderzoek is het meten of AI agents daadwerkelijk goed functioneren. Anders dan bij traditionele software (waarbij uitvoer deterministisch is), varieert de uitvoer van agents, zijn faalwijzen subtiel en vereisen succescriteria domeinexpertise om te definiëren. Dit evaluatiegat is de voornaamste drempel voor zelfverzekerd grootschalig inzetten in productie.
- Afweging eenvoud versus capaciteit: De voorkeur voor “eenvoudige, controleerbare” benaderingen in productie staat op gespannen voet met de capabelere agentische architecturen die beschikbaar zijn. Organisaties ruilen capaciteit in voor betrouwbaarheid — wat misschien passend is bij huidige risicoprofielen, maar mogelijk aanzienlijke waarde onbenut laat.
- Ontwerp van menselijk toezicht: Naarmate agents complexere en consequentere taken op zich nemen, wordt de vraag hoe mensen adequaat betrokken blijven (niet te veel, niet te weinig toezicht) cruciaal. Het delegatieraamwerk pakt dit architecturaal aan, maar de organisatorische en culturele normen voor passend menselijk toezicht op agents zijn nog niet uitgekristalliseerd.
- Kalibratie van vertrouwen: Gebruikers en operators neigen ertoe agents te veel te vertrouwen (uitvoer accepteren zonder voldoende controle) of te weinig (elke uitvoer opnieuw controleren, waardoor efficiëntiewinst teniet wordt gedaan). Het opbouwen van goed gekalibreerd vertrouwen is minstens zoveel een change management-uitdaging als een technische.
- Enterprise-inzet en de stap naar geavanceerde architecturen (webonderzoek, april 2026): Slechts 2% van de organisaties heeft AI agents op volledige schaal; 11% heeft ze in productie (Gartner, 2026). Gartner voorspelt dat 40% van de enterprise-applicaties eind 2026 AI agents zal bevatten (tegenover <5% in 2025). Het dominante patroon voor betrouwbare geavanceerde inzet is supervisor-worker orkestratie: een manager-agent breekt intentie op, routeert naar specialist-agents en synthetiseert resultaten. Feedbacklus-patronen (een reviewer-agent die de uitvoer van een andere agent controleert) verminderen hallucinaties significant. Casestudies: IBM — $3,5 miljard kostenbesparing, 50% productiviteitsstijging bedrijfsbreed; BCG wereldwijde biopharma — contentlokalisatie van 2 maanden naar 1 dag; Klarna — AI agents verwerken support voor 85 miljoen gebruikers met 80% snellere afhandeling; Uber — code-migratie-agents (maanden naar dagen). Intern Anthropic-onderzoek: multi-agent architectuur presteerde 90,2% beter dan single-agent Claude Opus-benchmarks. Sleuteltransitie: “human-in-the-loop” maakt plaats voor “human-on-the-loop” — mensen houden toezicht in plaats van elke beslissing goed te keuren. Prompt injection-aanvallen slagen 84% van de tijd in agentische systemen (OWASP Top 10 for LLMs 2025) — de governance/beveiligingslaag is nu een afzonderlijke architectuurvereiste. [source needed — MachineLearningMastery 7 Agentic AI Trends 2026; AaiNova Enterprise Architecture Guide 2026]
- ⚠️ MOGELIJKE TEGENSTELLING — enterprise-adoptiecijfers: definitieverschillen (webonderzoek, april 2026):
Gartner (2026)in dit artikel: 2% van de organisaties heeft AI agents op volledige schaal; 11% in productie. Meerdere enterprise-surveys uit 2026 (Joget/Gartner IDC, MasterOfCode, SQ Magazine) rapporteren nu dat 51% van de enterprises AI agents in productie heeft en 85% heeft ze geïmplementeerd of is dit van plan voor eind 2026. Het verschil (11% vs. 51%) weerspiegelt waarschijnlijk definitieverschillen: Gartner’s 2% “volledig op schaal” en 11% “in productie” hanteren mogelijk een strikte definitie (autonome meerstaps-agents); het cijfer van 51% omvat waarschijnlijk eenvoudigere implementaties (chatbots, copilots, enkelvoudige automatisering). Multi-agent architecturen specifiek zijn in minder dan 4 maanden met 327% gegroeid, wat suggereert dat ook de geavanceerde laag snel schaalt. De praktische conclusie: eenvoudige agent-implementaties zijn nu mainstream; geavanceerde multi-agent orkestratie schaalt maar is nog steeds een minderheid. [source needed — web search: masterofcode.com, joget.com, sqmagazine.co.uk AI agent statistics 2026] - 🔴 TODO (AANZIENLIJK VERNAUWD): De aanleiding om van eenvoudig naar geavanceerd over te stappen is nu gedocumenteerd: supervisor-worker patronen lossen betrouwbaarheid op; feedbacklussen pakken de evaluatie-uitdaging aan; schaal rechtvaardigt de investering. Resterende lacune: rigoureuze vergelijkingsstudie tussen eenvoudige en multi-agent benaderingen binnen dezelfde organisatie over tijd.
Gerelateerde Concepten
Workflow Redesign Around AI · AI Delegation and Multi-Agent Systems · Agentic AI Fundamentals · Skill Partnerships Human-AI · How AI Agents Actually Work in Production · When AI Teams Think Together, Not Just Together