Leaderboard· CHI-Bench v1.0.0

Every submission, one governed trial at a time.

Χ-Bench evaluates long-horizon, policy-rich U.S. healthcare workflow agents across three domains: provider prior authorization, payer utilization management, and care management. Each domain ships 25 tasks scored by an automated workspace judge under pass@1 with a binary 0/1 reward. Submissions below are ranked by accuracy on the selected domain.

Submissions
41
harness × model configs
Best pass@1
37.3%
erius · claude-opus-4-8
Last updated
2026-07-24
most recent merged run
#
Org
Agent
Model
Type
Accuracy
PA
UM
CM
Evidence
Date
01
Humana
eriussubmitted by Michael Johnson (MJ)
claude-opus-4-8
Proprietary
37.3%
40.0%
16.0%
56.0%
2026-06-05
02
Anthropic
claude-code
claude-opus-4-8
Proprietary
33.3%
32.0%
28.0%
40.0%
2026-05-28
03
Anthropic
claude-code
claude-opus-4-6
Proprietary
28.0%
18.7%
41.3%
24.0%
2026-05-01
04
Anthropic
claude-code
claude-sonnet-4-6
Proprietary
26.2%
24.0%
34.7%
20.0%
2026-05-01
05
OpenAI
codex
gpt-5.6-sol
Proprietary
25.3%
36.0%
28.0%
12.0%
2026-07-24
06
Anthropic
claude-code
claude-opus-4-7
Proprietary
24.4%
24.0%
17.3%
32.0%
2026-05-01
07
Anthropic
claude-code
claude-fable-5
Proprietary
24.0%
24.0%
24.0%
24.0%
2026-07-22
08
MedGuard
hermessubmitted by cuilinke
MedGuard
Open-source
22.7%
4.0%
4.0%
60.0%
2026-07-06
09
OpenAI
codex
gpt-5.5
Proprietary
20.9%
29.3%
32.0%
1.3%
2026-05-01
10
Anthropic
claude-code
claude-sonnet-5
Proprietary
20.0%
24.0%
24.0%
12.0%
2026-07-06
11
OpenAIZhipu
openai-agents
glm-5.1
Open-source
18.7%
18.7%
33.3%
4.0%
2026-05-01
12
Nous ResearchZhipu
hermes
glm-5.1
Open-source
18.7%
10.7%
34.7%
10.7%
2026-05-01
13
OpenAIZhipu
openai-agents
glm-5.2
Open-source
18.7%
20.0%
32.0%
4.0%
2026-07-06
14
OpenClawAnthropic
openclaw
claude-opus-4-7
Proprietary
17.3%
18.7%
13.3%
20.0%
2026-05-01
15
OpenClawZhipu
openclaw
glm-5.1
Open-source
16.9%
13.3%
26.7%
10.7%
2026-05-01
16
Nous ResearchAlibaba
hermes
qwen-3.6-max
Open-source
16.4%
9.3%
26.7%
13.3%
2026-05-01
17
OpenAI
codex
gpt-5.4
Proprietary
16.0%
24.0%
17.3%
6.7%
2026-05-01
18
OpenAIAlibaba
openai-agents
qwen-3.6-max
Open-source
15.6%
16.0%
26.7%
4.0%
2026-05-01
19
Nous ResearchMoonshot
hermes
kimi-k2.6
Open-source
15.6%
18.7%
21.3%
6.7%
2026-05-01
20
OpenAIMoonshot
openai-agents
kimi-k2.6
Open-source
15.1%
17.3%
25.3%
2.7%
2026-05-01
21
OpenAIDeepSeek
openai-agents
deepseek-v4-pro
Open-source
14.2%
10.7%
28.0%
4.0%
2026-05-01
22
Nous ResearchDeepSeek
hermes
deepseek-v4-pro
Open-source
13.8%
8.0%
25.3%
8.0%
2026-05-01
23
OpenAI
codex
gpt-5.6-terra
Proprietary
13.3%
12.0%
20.0%
8.0%
2026-07-24
24
OpenAI
codex
gpt-5.6-luna
Proprietary
13.3%
20.0%
16.0%
4.0%
2026-07-24
25
Google
gemini-cli
gemini-3-flash
Proprietary
12.5%
18.7%
18.7%
0.0%
2026-05-01
26
OpenClawDeepSeek
openclaw
deepseek-v4-pro
Open-source
11.1%
14.7%
12.0%
6.7%
2026-05-01
27
LangChainZhipu
deepagents
glm-5.1
Open-source
11.1%
17.3%
10.7%
5.3%
2026-05-01
28
LangChainDeepSeek
deepagents
deepseek-v4-pro
Open-source
10.7%
14.7%
10.7%
6.7%
2026-05-01
29
OpenClawMoonshot
openclaw
kimi-k2.6
Open-source
10.2%
12.0%
18.7%
0.0%
2026-05-01
30
LangChainAlibaba
deepagents
qwen-3.6-max
Open-source
9.3%
12.0%
10.7%
5.3%
2026-05-01
31
OpenAI
codex
gpt-5.4-mini
Proprietary
8.4%
10.7%
13.3%
1.3%
2026-05-01
32
OpenAITML
openai-agents
TML Inkling 256K
Open-source
8.0%
4.0%
16.0%
4.0%
2026-07-24
33
Google
gemini-cli
gemini-3.1-pro
Proprietary
7.1%
14.7%
6.7%
0.0%
2026-05-01
34
Anthropic
claude-code
claude-haiku-4-5
Proprietary
6.2%
0.0%
14.7%
4.0%
2026-05-01
35
OpenAIxAI
openai-agents
grok-4.3
Open-source
5.8%
0.0%
16.0%
1.3%
2026-05-01
36
OpenClawAlibaba
openclaw
qwen-3.6-max
Open-source
4.9%
10.7%
4.0%
0.0%
2026-05-01
37
Nous ResearchxAI
hermes
grok-4.3
Open-source
4.4%
0.0%
13.3%
0.0%
2026-05-01
38
LangChainMoonshot
deepagents
kimi-k2.6
Open-source
3.1%
8.0%
1.3%
0.0%
2026-05-01
39
LangChainxAI
deepagents
grok-4.3
Open-source
2.2%
0.0%
5.3%
1.3%
2026-05-01
40
OpenClawxAI
openclaw
grok-4.3
Open-source
0.4%
1.3%
0.0%
0.0%
2026-05-01
41
MedArise
pa-codex-agentsubmitted by MedArise
gpt-5.5
Proprietary
N/A
68.0%
N/A
N/A
2026-06-09
Submissions ranked by pass@1 on All Domains. Click any column header to sort.
Got Results?

Submit your agent to the CHI-Bench leaderboard.

Run the evaluation suite with your harness/model, prepare a packet with cb submission prepare, and open a PR. CI re-runs the validator and a maintainer merges within one business day.

Submit a run