
Healthcare IT News
FeaturedHealth system leaders say reliable benchmarking, governance and workflow design must come before scaling agentic AI across administrative operations. Johns Hopkins Medicine is taking a deliberately cautious approach to agentic AI, focusing first on proving reliability and governance before expecting measurable financial returns.

Blog
We ran 30 frontier agents through 75 real healthcare workflows. The best finished 28% on the first attempt and zero prior authorizations end to end. The constraint isn't model intelligence, it's the seam: 60 to 80 steps across four to six stages, where one broken stage ends the run.
Resource
Regulatory submissions, medical writing, pharmacovigilance case processing, MLR review — pharma’s operational long tail is document-heavy, cross-system work that generic AI can’t reliably execute. actAVA turns non-deterministic AI into governed, repeatable workflows, witha named humanowner at every GxP-critical decisionpoint —built indays andauditedfrom the first action.

Video
Author expertise once, set the rules once, and let every agent inherit both without bloating a single prompt. Every agent needs two very different kinds of knowledge, and most platforms cram both into one system prompt. KORA separates them into first-class sections on the build screen. Skills carry the reusable how: org-wide capability bundles of instructions plus files, authored once and shared. Memory carries the always and never: the short standing rules unique to how one agent should behave. Different jobs, different homes. This video builds a real one. A nursing team writes an end-of-shift handoff report every shift, same format, different details, different rules per team. The user describes that job to Ava in plain language, and Ava proposes the shape of the agent rather than presenting a form to fill. Before drafting anything, she searches existing agents and the organization's Skills Store, so you surface expertise the org already owns instead of building a duplicate. Then the Skill goes on. The vetted shift-handoff-report skill carries the SBAR template, the section ordering, the step-by-step instructions, and a bundled blank PDF the agent fills, attached in a single action. Here is the architectural payoff: only the skill's name and description sit in the system prompt. At runtime the agent calls LoadSkill to pull the full instructions from the session workspace, and only when a task actually needs them. That is progressive disclosure, and it is what lets an organization attach many vetted skills without slowing any agent down. The standing rules land somewhere else on purpose. Always cc the on-call lead. Never include patient MRNs in the summary. Ava recognizes that always and never language and writes each one into Memory as its own visible, deletable entry. Guardrails stay legible and editable instead of buried in a prompt where nobody can find them, which is exactly what you want when a compliance officer asks what an agent was instructed to do. One more thing worth watching. Ava runs a server-verified readiness check and then submits the draft. Ava only ever drafts. A human admin reviews and approves before anything goes live. Skills are reusable expertise. Memory is standing rules. One agent, powered by both, built entirely in conversation.

Video
linical AI that reaches a patient has to survive evaluation first. The demand is real. Roughly 61.5 million U.S. adults need behavioral health care, and 40 percent of them live in a designated provider shortage area. That gap will never close on clinician headcount. Meanwhile about half of the patients who do start treatment disengage before it can work. The pressure to put agents into longitudinal care across millions of covered lives is obvious. The failure rate is the problem. Behavioral health is where general-purpose AI breaks down hardest. On χ-BENCH, frontier models fail 72 percent of complex U.S. healthcare workflows, and fewer than 8 percent of agents stay successful when the same task repeats. That second number is the one that matters here. Agents holding the thread between sessions have to perform on the thousandth interaction, not the first. Scaffolding, not model choice, decides whether an agent is deployable. χ-BENCH was reviewed by researchers from Stanford, Johns Hopkins Medicine, Yale, and Salesforce AI Research. actAVA is the factory underneath. With KORA, you evaluate before patients ever see it: simulate agent behavior against your clinical rubrics, log every call and trajectory, and benchmark against χ-BENCH before anything reaches live care, with reviewed work feeding back so behavior improves on evidence. With CURA, the intelligence stays yours: tuned on your clinical protocols and your population, running inside your own cloud at lower cost per interaction. Operate CURA, frontier models, or both, and swap as the frontier moves. And the safety scaffolding ships prebuilt, with crisis detection, consent validation, escalation routing, human-in-the-loop oversight, and audit logging as platform primitives, plus EHR, claims, and FHIR integrations in the box. The behavioral health agent library is yours to extend. The clinical core covers between-session support, multi-source crisis detection, outcome measurement against PHQ-9 and GAD-7, and predictive engagement. From there, intake-to-matched-provider, eligibility-to-payment, ABA session-to-billing, practice compliance, and workforce optimization become expansion surface. Every agent ships with HIPAA compliance, human approval gates, and full audit logging. Behavioral health agents you can govern, prove, and own. Every agent tests before it deploys, every action lands in the audit log, the clinician stays in the lead, and the intelligence stays in your cloud, under your control. Your protocols, your cloud, your model.
Resource
You're putting agents into longitudinal care across millions of covered lives. actAVA is the factory underneath — agents test against clinical rubrics before deployment, run with a clinician in the lead, and prove every action in an audit trail. Your protocols, your cloud, your model.
Resource
actAVA is the Enterprise AI Sovereignty Platform for Healthcare — the agentic innovation engine that carries agent ideas from prototype to production on one governed foundation. Your teams build, including the ones without engineers. And every agent carries a budget, a benchmark, an evaluation, and a named owner before it touches a claim, a chart, or a patient.
Resource
actAVA standardizes and accelerates how AI agents for care management are built, tested, deployed, and improved. We make it easy and safe to build governed digital co-workers, so your care teams can perform at the top of their license.

Blog
Joey Kennedy is a veteran health technology CRO and enterprise software sales leader with more than 25 years of experience. This is Joey’s fifth time leading sales at an innovative health technology company. He has held chief revenue officer (CRO) and sales leadership roles across multiple healthcare and AI technology startups. He also has an MBA from UCLA with a focus on Healthcare Marketing and a Bachelor’s degree from BYU.

Video
The case for why care management can't scale on people alone, and what to do about it. Care management revenue runs on one equation: coordinators times billable minutes. To serve more patients, you hire more people. But the people aren't there. Turnover runs past twenty-five percent, HRSA projects a shortage of roughly 141,000 physicians by the late 2030s, and millions of patients who qualify for chronic care management stay unenrolled. You cannot hire your way out of that gap. General-purpose AI doesn't close it either. Care management is a longitudinal, multi-system coordination problem, with teams working across months rather than tickets closed in a single pass. On χ-BENCH, which grounds real administrative work in care policy across multiple tools, the strongest frontier agents finished twenty-eight percent of tasks on the first try. The problem isn't intelligence. It's variability, and single-task automation breaks on work this variable. Here is the finding that changes the design. The same frontier model succeeded with strict rules in the harness and failed outright without them, on the identical task. Reliability isn't the model. It's the system around the model. And governance is what makes that system inspectable: as Johns Hopkins frames it, accountability requires knowing which agent did what, under which policy, and with what result. Until you can measure reliability, you cannot measure ROI. That's why actAVA is three parts of one platform rather than a point tool. KORA is where you build the workflows, with multimodal agents across background, voice, and chat, each passing testing before it goes live. CURA is where you own the intelligence: five models reached the same decision 93 to 100 percent of the time with rules kept in the harness. CHRYSO is where you prove it, with timestamped sign-offs and audit packages on demand. And the stakes are rising on a schedule. CMS is expanding Medicare Advantage audits from about 60 plans a year to all 550. Audit readiness is moving from exception to expectation, which means the governance you build now is the audit package you will need soon. Build the workflows. Own the intelligence. Prove the governance.

Video
CURA delivering specialist-grade reasoning, not trivia recall. Frontier models are generalists. Healthcare punishes generalists. The cost of a confident wrong answer is measured in patient harm, denied claims, and audit findings. That is why actAVA built CURA, a 1-trillion-parameter model trained for clinical and administrative healthcare work, and why we tested it where it matters. Across six of the hardest healthcare benchmarks, CURA outperformed frontier models. This example focuses on the reasoning that utilization review, clinical documentation integrity, prior authorization criteria matching, and medical affairs work all depend on. MedXpertQA Text spans board-style questions across 17 specialties and 11 body systems. The cases read like real vignettes, with history, labs, and family history feeding a diagnostic and treatment decision that requires multi-step inference. MedXpertQA Multimodal goes further, pairing a medical image from radiology, pathology, or dermatology with patient records and exam results. The image alone is never enough. The model has to integrate visual findings with clinical context across multiple stages of the diagnostic process, which is exactly how imaging-adjacent workflows behave in production, from radiology worklist support to dermatology e-consults to pathology second reads. AgentClinic is the agentic test, and it is where reasoning has to hold up under pressure. A doctor agent starts with almost nothing, interviews a simulated patient, decides which tests to order, interprets the results, and commits to a diagnosis within a limited number of turns. Some cases add medical imagery. Others inject cognitive-bias perturbations such as anchoring or patient-introduced bias, designed specifically to degrade accuracy. CURA held up. For anyone building autonomous clinical agents on KORA, that is the difference between a demo and a deployable Non-Human Resource that manages its own line of questioning. Deployed through KORA, governed by CHRYSO, and measured continuously with χ-BENCH, CURA gives healthcare and life sciences teams a workforce-grade model for the thousands of workflows the industry has never been able to automate.

Video
CURA executing inside the EHR without breaking anything. Frontier models are generalists. Healthcare punishes generalists. The work is conversational one minute and transactional the next, and the cost of a confident wrong answer is measured in patient harm, denied claims, and audit findings. That is why actAVA built CURA, a 1-trillion-parameter model trained for clinical and administrative healthcare work, and why we tested it where it matters. Across six of the hardest healthcare benchmarks, CURA outperformed frontier models. This example focuses on where most models quietly fail, and where the long tail of healthcare administration actually lives. MedAgentBench v2 requires retrieving patient data through structured FHIR queries, documenting vitals and clinical notes, and placing medication, lab, and referral orders through correctly structured write operations. Tasks average two to three chained actions, mirroring routine inpatient and outpatient work, and a single malformed JSON payload or invalid API call counts as a failure. There is no partial credit for good intentions. CURA's performance here means agents that reason not only about the record, but work the record safely and in the required format. That is the capability every administrative workflow beyond prior authorization depends on, and it is the reason actAVA focuses on the complex, document-heavy, cross-system work that generic AI cannot reliably execute: the workflows that are hard, policy dense, long running, and unique to your organization. One model now covers the full arc of healthcare work, from an uncertain patient conversation through specialist reasoning to a correctly formatted order in the EHR. Deployed through KORA, governed by CHRYSO, and measured continuously with χ-BENCH, CURA gives healthcare and life sciences teams a workforce-grade model for the thousands of workflows the industry has never been able to automate.

Video
CURA talking with patients the way clinicians actually do. Frontier models are generalists. Healthcare punishes generalists. The work is conversational one minute and transactional the next, and the cost of a confident wrong answer is measured in patient harm, denied claims, and audit findings. That is why actAVA built CURA, a 1-trillion-parameter model trained for clinical and administrative healthcare work, and why we tested it where it matters. Across six of the hardest healthcare benchmarks, CURA outperformed frontier models. This example focuses on the messy front door of healthcare. HealthBench Hard tests exactly that: a patient describes vague or conflicting symptoms over a back-and-forth chat, and the model must infer the missing context before it can answer. It has to spot red flags and escalate to urgent care. It has to hedge when the evidence is thin rather than overcommit. And it has to perform in global health and resource-limited settings, not only in a well-resourced US academic center. For virtual care teams, nurse triage lines, and member services, this is the whole job. CURA's performance here means patient-facing agents that ask before they answer, escalate when they should, and stay honest about uncertainty. HealthBench Professional narrows further to the non-negotiables, testing behaviors physicians broadly agree are mandatory or forbidden, such as correctly recommending emergency referral when warranted, with near-zero error tolerance. That is the benchmark that maps directly to governance. When a health system asks how it can know the agent won't miss the thing it must never miss, CURA's results are the answer, and CHRYSO turns that answer into an auditable control. Deployed through KORA, governed by CHRYSO, and measured continuously with χ-BENCH, CURA gives healthcare and life sciences teams a workforce-grade model for the thousands of workflows the industry has never been able to automate.