Claude, GPT, Gemini Agents Fail 72% of U.S. Healthcare Workflows — CHI-Bench Study
AI labs have been positioning their agents as ready for complex, long-horizon workflows. A new benchmark released today puts that claim to the test in one of the highest-stakes environments possible: U.S. healthcare. The results are not reassuring. actAVA.ai released CHI-Bench, described as the world’s first long-horizon healthcare benchmark for AI agents. Testing 30 frontier agents across 75 real-world U.S. healthcare workflows, the benchmark found that the best-performing agent — Claude Code with Opus 4.6 — achieved a 28% pass rate at pass@1. That means even the top performer fails approximately 7 out of 10 real clinical cases. ...