Healthcare IT News just ran the story of how Dr. T. Y. Alvin Liu, M.D. and his team at Johns Hopkins Medicine are approaching agentic AI. Before deploying anything, they built a benchmark on actAVA to find where frontier agents break first. χ-Bench: a simulation of 25 healthcare applications, 77 tools, and a 1,290-document managed-care operations handbook. Real prior authorization, real utilization management, real RN care management. The workflows that carry the highest dollar value and the highest cost when they fail.
Healthcare IT News
Health system leaders say reliable benchmarking, governance and workflow design must come before scaling agentic AI across administrative operations.

Newsweek
As health care increasingly adopts AI, experts say reliability and governance may matter more than model size.
AI Weekly
arXiv
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM trained through a human-gated self-evolution loop. In each evolution round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures. This data-centered loop improves the model through targeted synthetic and curated examples rather than a single generic medical-data update. Across the healthcare evaluation suite, Cura 1T ranks at or near the top among frontier baselines, while remaining competitive on out-of-domain reasoning and agentic benchmarks.
Hugging Face
GitHub
actAVA Cura: Specialized Model for Agentic Healthcare

Postcards From the Edge
The Intelligence Speed Race Continues

Qiita
「AIで医療が変わる」と騒いでいたあなたへ。 昨日(2026年5月20日)発表されたベンチマークが、AI業界に衝撃を与えました。 最新のClaude、GPT-5.5、Geminiを含む30のAIエージェントが、米国の医療ワークフローで72%失敗したのです。 結論から言うと...

The Daily News (Galveston)
Carroll County News
FinancialContent
arXiv