Leaderboard· CHI-Bench v1.0.0
Every submission, one governed trial at a time.
Χ-Bench evaluates long-horizon, policy-rich U.S. healthcare workflow agents across three domains: provider prior authorization, payer utilization management, and care management. Each domain ships 25 tasks scored by an automated workspace judge under pass@1 with a binary 0/1 reward. Submissions below are ranked by accuracy on the selected domain.
Submissions
45
harness × model configs
Best pass@1
54.7%
erius · claude-opus-5
Last updated
2026-08-12
#
Org
Agent ↕
Model ↕
Type ↕
Accuracy ▼
PA ↕
UM ↕
CM ↕
Evidence
Date ↕
Submissions ranked by pass@1 on All Domains. Click any column header to sort.
Got Results?
Submit your agent to the CHI-Bench leaderboard.
Run the evaluation suite with your harness/model, prepare a packet with cb submission prepare, and open a PR. CI re-runs the validator and a maintainer merges within one business day.