Our customersCase studyAcademic medicine / AI innovation
Johns Hopkins Medicine

Johns Hopkins is building the AI the market won't sell, and benchmarking every agent before it deploys.

The Gills AI Innovation Center benchmarked frontier agents on ACTAVA before deploying any. The evidence pushed it toward owning its models, not renting them.

Of workflows no vendor sells
90%
1 in 10 sold; 9 in 10 yours to build
Cheaper per token
20–100x
An owned post-trained model against a rented frontier model
Tools tested before deployment
200+
Across 21 healthcare applications in χ-BENCH
Measured first
Reliability
Token consumption and per-agent cost tracked against baselines; financial return reported when the evidence earns it
01 / Context

Overview

Institutional memory in a research center usually lives in the heads of the people who have been there longest, in email chains and the PDFs that describe how the place runs. It does not scale, and when those people leave, the memory leaves with them. Dr. T. Y. Alvin Liu's definition of success for the center he directs is to become replaceable: to capture everything the 2026 version of himself knows so it can be handed on.

A $10 million gift from Dr. James Gills, an ophthalmologist and early NVIDIA investor, launched the Gills AI Innovation Center at Johns Hopkins about two years ago. Its three pillars are ophthalmology AI research, collaboration across academia and industry, and commercialization. Dr. Liu directs it and sits on the Johns Hopkins Medicine AI Oversight Team, so the question of what an agent may do, and how anyone would know, is his to answer.

02 / The bottleneck

The challenge

Give a healthcare organization unlimited resources and ask whether it would build or buy, and Dr. Liu says the answer comes back the same every time: the workflows and data are too sensitive to hand off. Health systems want to own their AI destiny. The market does not make that easy. AI-native vendors sell revenue cycle, patient access, and prescribing, the biggest problems with the biggest markets. Of 100 workflows worth transforming, they serve perhaps the top 10. Nobody sells the other 90, because those problems run too small, too specific, too idiosyncratic to one organization, and most healthcare organizations cannot build them either.

A harder question came before any of that. Can frontier agents finish policy-dense, end-to-end healthcare work at all? Buying the application still means renting the intelligence underneath it, and a demo does not show whether an agent will finish the same task the same way tomorrow. Hopkins wanted that evidence before anything deployed.

03 / The workflow

The solution

The center put frontier agents through χ-BENCH on ACTAVA before deploying anything: 200+ tools across a simulator of 21 healthcare applications, with tasks grounded in real managed care policy. Reliability was the first measurement, not return.

The results pushed Dr. Liu toward ownership over rental. ACTAVA is built for the long tail of administrative work, the 90% of workflows everyone else skips, and a team that builds there comes out owning a post-trained model: 20 to 100 times cheaper to run than a rented frontier model, its own IP, trained on its own data. The team co-develops inside the platform with ACTAVA's engineers, pulls from the agent library to start, benchmarks its own workflows, and uses the results to fine-tune a model that fits how it works, starting from an open-source model of its choice.

How Hopkins moves an agent from idea to deployment

Swipe to see the full diagram

  1. Step 1: Benchmark before anything deploys

    Frontier agents run through χ-BENCH on ACTAVA against a simulator of 21 healthcare applications and 200+ tools, on tasks grounded in real managed care policy.

  2. Step 2: Measure reliability first

    The question is whether an agent finishes policy-dense, end-to-end work, and whether it does so repeatably. Financial return is not reported until the evidence earns it.

  3. Step 3: Track cost against a baseline

    Token consumption and per-agent cost are tracked against baselines, so the economics of an agent are known before it scales.

  4. Step 4: Keep every action accountable

    Accountability requires knowing which agent did what, under which policy, and with what result. Every run leaves that record.

  5. Step 5: Own the model

    Production workflows and their evidence fine-tune a model the center owns, starting from an open-source base of its choice, co-developed with ACTAVA's engineers.

The first workflow is the center's own AI brain. Every contract, operating procedure, and email tied to the center feeds a database layer, and an intelligence layer sits on top of it, across everything the center does, from research to collaboration to commercialization. Then the questions that used to depend on someone's memory get real answers: how many external organizations the center has engaged, how many became concrete collaborations, how long a contract takes to sign, and why that changed.

In their words
Dr. T. Y. Alvin Liu

Accountability requires knowing which agent did what, under which policy, and with what result.

Dr. T. Y. Alvin Liu
Director, Gills AI Innovation Center, Johns Hopkins Medicine
04 / The impact

The result

Hopkins deploys agents with the evidence in hand rather than the promise. Reliability is measured before financial return, token consumption and per-agent cost are tracked against baselines, and every agent that reaches a workflow has been through the same benchmark as the one before it.

The build-versus-buy math changes with it. A post-trained model the center keeps runs 20 to 100 times cheaper per token than renting a frontier model for the same work, which matters when a single agentic task can burn millions of tokens re-reading records. It is IP the center owns instead of a bill it keeps paying.

Dr. Liu's larger point is about scale. Run AI across the top 10% of workflows and you improve an organization. Run it across all 100 and you transform one, and the advantage stops being merely quantitative. Five years out, he expects the market to split into the organizations that got AI right and the ones that missed it. Hopkins intends to be in the first group, on models it owns.

Sources: The Johns Hopkins interview, ACTAVA, July 31, 2026Healthcare IT News, July 24, 2026

Our customers
See all customer stories
  • ThoroughCare
    Care management

    ThoroughCare enrolls patients 6x faster with automated eligibility intake and is targeting an enrollment-rate lift of up to 20%.

    Program enrollment, down from six months
    30 days
    More enrollments per care manager
    5x
    Less token burn from agent optimization
    11x
    Read the case study: ThoroughCare
  • Momentum Actuarial
    Actuarial / Clinical vendor evaluation

    Momentum Actuarial scaled clinical vendor evaluation 100x without adding analysts or compromising depth.

    From first agent build to production
    30 days
    Report output without more analysts
    100x
    Read the case study: Momentum Actuarial
  • Simplify
    Operational performance / Coaching

    Simplify brings its best operators' knowledge to every location, with a coach that helps people act.

    Background preparation, from prospect research to playbook upkeep
    3 agents
    One conversational coach inside ProNavigator
    4 roles
    Simplify's methodology, kept as versioned coaching skills
    Owned
    Read the case study: Simplify
Next step

Agents you can govern, prove, and own.

Build agents on a platform that tests them before they go live, keeps every action on the record, and puts the intelligence under your control. Orchestrate with KORA, own the model with CURA, and prove every action with the governance built in.

All customer stories