Blog

Meet our Advisors: Dr. Sanmi Koyejo, PhD

Dr. Sanmi Koyejo, PhD, directs Stanford's Trustworthy AI Research Lab and was named TIME's 100 Most Influential People in AI for 2026. He joins ACTAVA as a Strategic Advisor. His lab built a clinician-validated taxonomy of 121 real medical tasks and benchmarked nine frontier models against it.

8 min read·September 11, 2026
Meet our advisor: Dr. Sanmi Koyejo. He measures what AI can actually do.

Meet our advisor: Dr. Sanmi Koyejo. He measures what AI can actually do.

We're proud to welcome Dr. Sanmi Koyejo to ACTAVA's advisory board as an AI Research Advisor. He directs the Stanford Trustworthy AI Research Lab, and he has spent his career on the one question healthcare AI keeps skipping: how do you make an honest claim about what a system can and cannot do?

"Healthcare is where the gap between how an AI system performs in evaluation and how it performs in practice carries the highest stakes. ACTAVA is building at the orchestration layer, where that gap gets closed, and I'm glad to be helping them get it right."

Dr. Sanmi Koyejo

Koyejo is an assistant professor of computer science at Stanford, where he leads the STAIR lab. TIME named him one of the 100 Most Influential People in AI for 2026, in the Thinkers category. His lab builds the measurement science behind trustworthy AI: how to evaluate a model's real capabilities, hold it accountable for its failures, and protect privacy in systems that touch patients.

That's a narrow-sounding specialty. In healthcare it's the whole ballgame.

The gap he studies is the gap we build for

Last year Koyejo and a Stanford-led team published MedHELM, an evaluation framework for medical language models. They built a clinician-validated taxonomy with 29 clinicians: 5 categories, 22 subcategories, 121 real medical tasks. Then they ran 9 frontier models against 35 benchmarks covering all of it.

The paper opens with the problem in one line: large language models "achieve near-perfect scores on medical licensing exams," and those scores "inadequately reflect the complexity and diversity of real-world clinical practice."

Koyejo has put the same finding more bluntly:

"In other work, we have shown that foundation models can score at an expert level on medical licensing exams yet fail at basic diagnostic reasoning tasks. Understanding where and why systems fail is core to whether they can be trusted to help rather than mislead."

Computing Research News, January 2026

Every healthcare AI buyer has felt the consequence of that gap without having a name for it. The demo is flawless. The pilot looks great. Then the thing meets prior authorization at a real health plan, with real edge cases and a real appeals deadline, and the numbers stop matching the pitch deck.

Why did you decide to advise ACTAVA?

"The first reason is the people: I've known Weiran and Frank's work for years, and they approach hard problems with a rigor I trust. The second is the problem itself. Healthcare AI increasingly runs on orchestrated systems, and most of our evaluation tools were built for single models. The orchestration layer is exactly where questions about reliability, accountability, and real clinical value get decided. I think this is critical for real-world impact, and ACTAVA is building there."

Dr. Sanmi Koyejo

Weiran Yao is our CAIO and co-founder. Frank Wang is our CTO and co-founder.

The second half of that answer is worth sitting with. Almost every evaluation tool in circulation was designed to score one model on one prompt. Nobody deploys that way anymore. A prior auth workflow is a chain: retrieve the policy, read the chart, apply the rule, draft the determination, route for review. The failure you care about lives in the handoffs.

One finding from his group makes the risk concrete. AI systems often agree with each other even when they're wrong. Agreement feels like corroboration. It isn't, and in a multi-agent chain a confident mistake can pass down the line while everything downstream nods along.

ACTAVA KORA χ-BENCH ACTAVA Compliance CURA

χ-BENCH, our simulation and benchmarking suite, exists for exactly this. It simulates the workflow rather than quizzing the model, scoring agents on outcomes, compliance, and cost before they reach production. It's also why ACTAVA is model-independent by design. A second opinion is only worth something if it's actually independent.

What are the biggest changes you see underway in healthcare AI?

"The center of gravity is shifting from capability to accountability. For years the field asked whether models could pass medical exams; now the harder questions are operational. Does benchmark performance transfer to a specific clinical workflow? Who is responsible when a multi-step system fails, and can we even localize the failure? Evaluation is becoming the bottleneck, and I mean that as good news: it signals that these systems are close enough to real deployment that measurement finally matters more than demos."

Dr. Sanmi Koyejo

"Can we even localize the failure" is the question most healthcare AI stacks cannot answer today. It's a governance problem before it's a technical one, and it's what ACTAVA Compliance is built around: role-based access, audit trails, human-in-the-loop review, and a record of which step decided what.

His definition of the goal is the one we build against:

"In healthcare, 'trustworthy' means clinicians can understand what the AI system is doing well and where it might fail, so they can decide when to rely on it. It means the system works reliably across different patient populations, not just those well-represented in the training data."

Computing Research News, January 2026

That last clause isn't a footnote. His lab has built methods to keep chest X-ray diagnostic systems working equally well across racial groups, and has NSF-funded work on whether AI in breast cancer screening can reduce outcome disparities rather than just lift average accuracy. A system that performs well on average and badly for a subgroup is not a system that works.

A test you can run on any healthcare AI vendor this quarter. Ask them to name a task their system is bad at, and to show you the evaluation that proves it. Honest claims cut both ways.

The prediction: evaluation becomes a procurement requirement

We asked him what he expects for enterprise AI over the next 18 months.

"Evaluation will move from a research concern to a procurement requirement. Enterprises, especially healthcare systems, will stop accepting leaderboard rank as evidence and start demanding proof that a system works in their context, on their data, in their workflows. Vendors that can credibly furnish that evidence will win deals; measurement is becoming part of the product."

Dr. Sanmi Koyejo

Measurement becoming part of the product is the whole ACTAVA thesis, stated by someone with no reason to flatter us. Our position is that healthcare organizations should own their intelligence rather than rent it. Own the workflows, own the evaluation, own the models. That only means something if you can measure what you own, in your context, on your data.

It's also why the Learn pillar runs underneath everything. Every production workflow generates evidence, and that evidence turns into validated, auditable improvements instead of silent drift. Evaluation that stops at go-live is evaluation that expires.

Talk to the clinicians

His advice for people building healthcare AI is refreshingly unglamorous:

"Don't just read papers about AI for healthcare, talk to clinicians, observe clinical workflows, and understand what problems practitioners actually face. The best research questions come from those conversations, not from trying to find healthcare applications for your favorite ML technique."

Computing Research News, January 2026

This is why ACTAVA KORA is built as a no-code conversational interface. The people who know where prior auth actually breaks are utilization review nurses and clinical operations leads, not ML engineers. If building an agent requires an engineering ticket, the domain knowledge never makes it into the system.

More about our advisor

Koyejo is an assistant professor of computer science at Stanford and an adjunct associate professor at the University of Illinois Urbana-Champaign, where he previously held an associate professorship. He earned his PhD at the University of Texas at Austin and has served as a research scientist at Google DeepMind.

His honors include the Presidential Early Career Award for Scientists and Engineers, the Alfred P. Sloan Research Fellowship, the NSF CAREER Award, the CRA's Skip Ellis Early Career Award, and multiple outstanding paper awards at NeurIPS and ACL.

He co-founded Virtue AI, an enterprise AI safety and security company, alongside fellow ACTAVA advisor Bo Li. He is board president of Black in AI and sits on the boards of the NeurIPS Foundation and the Association for Health, Learning, and Inference.

We're fortunate to have him in the room. Healthcare doesn't need more systems that test well. It needs honest claims about what they can do, and a way to check.

Want to see how χ-BENCH validates an agent before it reaches production?

Read the docs or start a conversation with our team.

Sources

  1. Dr. Sanmi Koyejo, in conversation with ACTAVA, September 2026.
  2. Lashlee, L. "CRA Early Career Award Winner Spotlight: Sanmi Koyejo, Stanford University." Computing Research News, Vol. 38/No. 1, January 2026. cra.org
  3. Bedi, S. et al. "MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks." arXiv:2505.23802, May 2025. arxiv.org and medhelm.org
  4. Shanbhog, S. "Meet the AAS 248 Plenary Speakers: Dr. Sanmi Koyejo." Astrobites, 14 June 2026. astrobites.org
  5. "Sanmi Koyejo: The 100 Most Influential People in AI 2026." TIME, August 2026. time.com
  6. Koyejo, S. Faculty bio, Stanford Computer Science. cs.stanford.edu
  7. Stanford Trustworthy AI Research (STAIR) Lab. stairlab.stanford.edu
Share this