Blog

Newsweek Asked If Healthcare AI Hit a Reliability Wall. The Answer Is an Architecture Choice.

Newsweek asked this week whether health care's AI ambitions have hit a reliability wall. The gap is real: 40% of healthcare organizations are spending $50-100M a year on AI, and only 18% feel ready to implement it. But the wall isn't intelligence, it's architecture. The strongest agent framework we tested resolves just 28% of complex healthcare tasks on the first attempt, and that's a systems problem no bigger model fixes. This piece makes the case that reliability is a build-versus-rent decision: keep wiring generalized models into clinical workflows and inherit their blind spots, or own the stack. That's CURA, our healthcare-native model, KORA, the factory that builds, tests, and improves the agents around it, and CHRYSO, the governance layer that keeps every agent inside your policy. Reliability comes from the architecture around the model, not the model alone.

By Kevin Riley, Weiran Yao and Frank Wang

5 min read·July 23, 2026
Newsweek Asked If Healthcare AI Hit a Reliability Wall. The Answer Is an Architecture Choice.

Newsweek Asked If Healthcare AI Hit a Reliability Wall. The Answer Is an Architecture Choice.

Newsweek ran a piece this week asking whether health care's AI ambitions have stalled. The honest answer: they've hit the limit of a rented, generalized approach. The way through isn't a bigger model. It's owning the system.

The numbers in the article tell the story better than any pitch. 40% of healthcare organizations are committing $50-100M a year to AI. Only 18% feel ready to actually implement it. That gap is where budgets go to die.

So it's fair to ask if the ambition has outrun the technology. But the wall isn't intelligence. Models are plenty smart. The wall is what happens when you point a general-purpose model at a clinical workflow and ask it to be dependable across every step, every policy, and every regulatory line it has to respect.

The 28% that should worry everyone

actAVA built a healthcare workflow benchmark, χ-Bench, to measure this directly. The strongest agent framework we tested resolved just 28% of complex tasks on the first attempt.

Read that again. The best available setup, on real healthcare admin work, gets it right the first time barely more than a quarter of the time. That's not a model-quality problem you fix with more parameters. It's a systems problem: long-horizon work, hard policy constraints, and no room to quietly retry when a decision has already committed.

Health care has always been a domain where reliability matters, because decisions connect directly to people, processes, and regulations.

Reliability isn't a feature you sprinkle on at the end. In care settings, it's the whole product. A model that's brilliant 72% of the time and unaccountable the rest of the time is not something you deploy near a patient, a claim, or a prior authorization.

The real choice: build it, or rent it

Here's the shift most healthcare leaders are circling but few have named out loud. You can keep renting generalized AI, wiring a closed frontier model into your workflows and hoping it holds. Or you can build domain-specific systems you actually own and can hold accountable.

Renting feels faster. You get a capable model on day one. But you inherit its blind spots, its update schedule, and its indifference to your policies. When it fails the 72%, that failure is yours to explain and not yours to fix.

Building sounds slower until you look at what you get: a model tuned to your work, orchestration you control, governance your compliance team can defend, and a system that improves on your data. The reliability you need doesn't come from the model alone. It comes from the architecture around it.

Why it matters

A generalized model has no idea what your utilization management policy says, which reviewer signs off on what, or why a referral gets denied. Build the system and those rules become enforceable. Rent it, and they stay optional, right up until an auditor asks.

What "owning the system" actually means

Owning your healthcare AI isn't about training a model from scratch in a basement. It's about controlling four things the rented approach hands to a vendor: the model, the orchestration, the governance, and the improvement loop. Each one has a piece of the actAVA KORA factory behind it.

1. A model built for the work. Our answer here is CURA 1T, a healthcare-specialized model post-trained for patient communication, clinical reasoning, and agentic EHR workflows. General frontier models cover for healthcare's scattered training signal with raw scale. We trained for the actual work instead, and it runs at a fraction of the cost per token, which matters when one agentic task burns millions of tokens re-reading records. That's the difference between a model that's read about medicine and one built to do the work.

2. Orchestration you control. Complex clinical work is rarely one call to one model. It's many steps, many roles, and many tools, coordinated. Our KORA is where you build that: a natural-language agent builder, so a nurse manager or an ops lead can describe a workflow in plain English and stand up a working agent, with healthcare guardrails baked in from the start. That coordination is where the 28% breaks down, and where owning the system pays off.

3. Governance that ships on day one. You can't govern what you can't measure, and you can't deploy what you can't defend in an audit. Our KORA runs continuous red-teaming, hallucination checks, and bias detection calibrated for clinical scenarios. Our CHRYSO enforces policy across every agent: PHI redaction, role gating, and audit trails, the things your compliance team actually asks about. Together they turn "trust us" into numbers your team can stand behind before an agent ever touches a patient workflow.

4. An improvement loop on your data. The system that learns from your real interactions gets better in ways a rented model never will, because it never sees the data that would teach it. KORA is also the continual learning layer: agents adapt while they're in production, every interaction becomes signal, and you ship versioned upgrades instead of hoping a vendor's next release helps. The agent that solves your problem in January should be measurably better at it by March.

The stack, in one line

CURA is the model built for healthcare. KORA is the factory that builds, tests, and improves the agents around it. CHRYSO is the governance layer that keeps every one of them inside your policy. Reliability is what you get when all three run together, not any one of them alone.

Where actAVA fits

actAVA KORA is the AI factory that makes owning this practical. It gives you the primitives in one place: CURA as the healthcare-native model, KORA to build and orchestrate agents, KORA and CHRYSO to test and govern them against your access rules, and KORA to run the continual learning loop that keeps them improving. Reliability comes from the whole factory working together, not from any single model being clever.

We deploy on your infrastructure, or through zero-data-retention hosted services, so your workflows and your data stay yours. Our engineers work alongside your teams, transfer the knowledge, and design themselves out once the systems run on their own.

So did healthcare AI hit a reliability wall? Only if you keep renting your way into it. The organizations that treat reliability as an architecture decision, not a model upgrade, are the ones who'll get to the other side. Build the system. Own the outcome.


Kevin Riley

Authors

Kevin Riley

CEO & Co-Founder

Weiran Yao

Weiran Yao

CAIO & Co-Founder

Frank Wang

Frank Wang

CTO & Co-Founder

Share this