STORY · FORSKNING_

IBM and UC Berkeley identify why enterprise agents fail

IBM Research and UC Berkeley have developed IT-Bench and MAST, tools for diagnosing failures in enterprise AI agents. The benchmark and methodology reveal systematic weaknesses in how agents handle IT tasks and complex workflows in business environments.

WHY IT MATTERS

Understanding where enterprise agents fail is critical to making them work reliably in production. This affects agentic AI adoption in large organizations and provides concrete indicators of where models need improvement.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.