
Healthcare IT News
FeaturedHealth system leaders say reliable benchmarking, governance and workflow design must come before scaling agentic AI across administrative operations. Johns Hopkins Medicine is taking a deliberately cautious approach to agentic AI, focusing first on proving reliability and governance before expecting measurable financial returns.

Newsweek
As health care increasingly adopts AI, experts say reliability and governance may matter more than model size.

AI Weekly
actAVA AI released Cura 1T, a healthcare-specialized LLM trained through what the authors call a human-gated self-evolution loop. In each round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from failures.

Qiita
A benchmark released yesterday (May 20, 2026) sent shockwaves through the AI industry. Thirty AI agents — including the latest Claude, GPT-5.5, and Gemini — failed 72% of the time on U.S. medical workflows.

FinancialContent
AI labs position agents as ready for long workflows, but until now no public benchmark validated that claim in healthcare, where one missed policy check can mean a denied authorization, delayed treatment, or audit finding. Each trial in CHI-Bench runs an agent for 60-80 steps across four to six clinical stages, exposing 21 healthcare apps through 200+ MCP tools and a 1,279-document operations handbook. It evaluates the trajectory, every artifact, and world state using deterministic unit tests and LLM judge for evidence grounding, consent, and cross-stage consistency.

The Daily News (Galveston)
Across the 30 frontier agents tested, Anthropic's Claude Code with Opus 4.6 achieved the best overall performance at 28% pass@1, followed by OpenAI's Codex with GPT-5.5 at 21%. By domain, utilization review reached 41%, care management 32%, and prior-authorization paperwork 29%. Reliability remained a major issue, with no agent clearing 20% when the same case was run three times. Under endurance testing, where agents were asked to handle 25 cases in one session, the best system completed under 4%. In a fully end-to-end setting, where one AI submitted a prior-auth request and a second acted as the UM reviewer, no task passed successfully.

ITNewsOnline
actAVA now offers a purpose-built solution to test and monitor AI agents against real, accepted standards. CHRYSO offers radically simplified AI policy management, team training, and agent testing. An enterprise-level suite integrates policy management, virtual training, and a specialized Agent Registry helps organizations mitigate legal and ethical risks. By providing real-time monitoring and automated control evidence, CHRYSO ensures AI agents remain compliant with critical standards, including NIST AI RMF, HIPAA, CMS HEI, and ONC HT1.

Weixin Official Accounts Platform
Over the past year, the narrative around AI agents has shifted from "can it do a demo?" to "can it go into production?" Coding agents were the first to break out, because they have clearly defined tasks, strong feedback signals, and verifiable results. But in more complex enterprise scenarios, agents still run up against a core problem: the fact that an agent can answer a question doesn't mean it can get the work done. Healthcare may be the industry where this problem is most concentrated. On one hand, U.S. healthcare workflows are extremely standardized — full of policy, compliance, approvals, documentation, and cross-system coordination. On the other hand, they're also extremely complex: between the three parties of provider, payer, and patient, there are a great many processes that are long-chain, irreversible, and heavily regulated. What this means is that the real problem a healthcare agent has to solve isn't "medical Q&A" — it's whether it can execute tasks over the long term, reliably, and compliantly within a real workflow.