Resources
Don't take our word for it. Take theirs.
None of this research is ours. It's where we'd start if we were pressure-testing the category.
Of silent failures in a documented production runtime were caught by a person reading the output — not by tests or health checks.
Of enterprises shipped an agent that passed internal evaluation and then failed a customer.
Fully trust the automated evaluations making their release decisions.
The research.
Silent failures in a production agent runtime
Eight weeks of postmortems. Thousands of tests prevented none of the novel incidents; some failures ran silent for sixty days.
Read →UC BerkeleyMeasuring Agents in Production
306 practitioners. Most production agents cap at ten steps before a human intervenes, and 74% still rely on human evaluation.
Read →Berkeley · MASTWhy do multi-agent systems fail?
Fourteen failure modes from 1,600+ traces, and the finding that better base models won't fix most of them.
Read →VentureBeatThe enterprise evaluation gap
Autonomy is arriving faster than companies can verify it. The top reason evals aren't trusted is poor alignment with real outcomes.
Read →arXivMeasuring long-horizon execution
Per-step error rate rises as a task progresses, because models condition on their own error-prone history.
Read →Our writingNotes from production traffic
What the signals catch, what they miss, and how often the lever call turns out to be right.
Read →Send us one week of past runs.
We'll come back with what your successful executions have in common, the few that stopped matching, and what we think changed. If we find nothing, we'll tell you that too.
Start a design partner conversation