The world’s first long-horizon healthcare benchmark for AI agents.
CHI-Bench paper accepted to NeurIPS 2026
Built with the clinicians who do the work.
Actava.ai partners with 20+ hospitals and top universities to evaluate frontier agents across prior authorization, utilization management, and care management. The best agent today resolves 28% of tasks at pass@1; end-to-end prior authorization automation drops to 0%.
Benchmarking agentic AI for healthcare administration.
See how CHI-Bench evaluates whether AI agents can reliably navigate real healthcare administrative workflows — across policies, applications, roles, evidence, and long sequences of actions.
One agent, one trial, one final scorecard.
Each animation replays an actual trajectory from the claude-code-opus-4-6 submission — stage transitions, real tool calls, real policy gates, real scorecard outcomes.
An RN drives a clinical referral from chart pull through policy lookup, evidence assembly, and a final action decision. Every stage commits — there is no retry.
Why it’s hard · One wrong site-of-service flip cascades into four scorecard failures.
Three capabilities underrepresented in current benchmarks.
Axes
Long-horizon
60–80 agent steps per trial across 4–6 distinct stages, with state that carries forward and can't be retried after commit.
Role-composed
One agent plays many seats — intake clerk, nurse, medical director, peer-to-peer reviewer, letter center — and each seat writes its own artifact.
Policy-driven
Medical-policy criteria, site-of-service rules, consent scripts, evidence-grounding requirements — checked by deterministic rubrics plus an LLM semantic judge.
Run a submission. Or read the methodology.
The benchmark code, dataset, and 1,279-page operations handbook are open. The leaderboard is live.
Leaderboard
30 agent configurations — harness × model — ranked across prior authorization, utilization management, and care management.
View the leaderboard75 tasks
25 per domain, each one a full working trial: seats, tools, policy gates, and a final scorecard.
Browse the tasksDocs
The methodology, the environment, and the submission guide — grounded in a 1,279-document operations corpus.
Read the docsFrom production workflows to customer-controlled intelligence.
Don't give your agentic future away to a single model provider. Don't mistake consumer tools for real, safe Enterprise Agentic Tools. Enable your citizen developers to create and manage the AI Agents they need to run their part of your business.
Build complex agents. Test their reliability. Learn from every workflow. Own your intelligence.