Medigram Insights · Clinical AI Governance · June 2026

What Happens When You Actually Test Whether Your Clinical AI Does What It Claims

A practitioner account of running Praxen, the open-source reference implementation for Agent Behavior Verification, on a governed clinical AI platform through early access, and what we found.

The Gap No One Wants to Talk About

Healthcare AI has a compliance theater problem. We write policies. We configure controls. We collect certifications. And then we deploy systems whose actual behavior under real and adversarial conditions we have never verified.

I have spent over a decade building the governance architecture to address this. I co-authored IEEE UL 2933 and ANSI/HSI 2800:2025 precisely because the field needed standards that required evidence of behavior, not just documentation of intent. But even with that architecture in place, I kept asking one question about Darwin, the governed clinical AI platform we built at Medigram: how do I know it does what we designed it to do, not just under ideal conditions, but under adversarial ones?

That question is what led us to Praxen. What we discovered when we ran it changed how we think about what clinical AI governance requires.

What Enterprise Clinical AI Governance Actually Requires

The gap between what pilots and vendors make clinical AI governance look like and what it requires is the central problem of this field.

Enterprise clinical AI governance layers diagram
This diagram was presented by Sherri Douville and Mitch Parker in the Maturity Model Committee update to the IEEE UL 2933 committee. Image copyright Sherri Douville, CEO Medigram, Founder and Chair, TTIC.

Most vendors present clinical AI governance as two layers: model accuracy and basic vendor documentation. Enterprise clinical AI governance requires eleven. The realization that Praxen and Darwin together covered all eleven layers is what inspired the TTIC certification pathway. Not as a product decision, but as a governance completeness realization. Praxen addresses the technical and behavioral layers through pre-deployment and ongoing verification. Darwin governs the clinical transaction layers continuously after deployment. Together they cover the full stack.

What Praxen Is and Why It Matters

Agent Behavior Verification is a new and important discipline in AI security. It is the practice of probing an AI system's actual behavior against its intended behavior under adversarial conditions, before deployment and on an ongoing basis after deployment. Praxen is the open-source reference implementation of ABV, built by Steve Wilson and released under Apache 2.0 by Exabeam.

Praxen reflects and extends Exabeam's established leadership in behavior intelligence and agent security. Where Exabeam has long led the field in detecting anomalous behavior by humans and non-human identities at runtime, Praxen extends that leadership upstream into pre-deployment verification, closing the gap between what an AI agent is configured to do and what it actually does before it ever reaches production.

Praxen is not a compliance tool in the SOC 2 sense. It does not ask what your system is configured to do. It probes what your system actually does, including under conditions designed to make it fail. The distinction matters enormously in clinical AI. A policy that says a system will not disclose protected health information is not a control. It is a wish. Behavioral verification is the difference between documentation of intent and evidence of behavior.

Traditional compliance frameworks were built for deterministic systems. AI agents are probabilistic, context-dependent, and drift over time without anyone touching the configuration. Praxen is designed for this reality. Darwin is engineered to address it from the other direction: its scoring is designed to produce deterministic, auditable results that enterprises can rely on, by design rather than by assumption.

What We Ran and What We Found

We accessed Praxen through its early access program before public launch. We ran it against Darwin across three successive assessment cycles, treating each run's findings as a remediation roadmap. The third run, using Praxen 0.8.0, produced the following result:

RAISE Score  ·  Praxen 0.8.0  ·  June 22, 2026
4.0
Strong · Zero Open Findings
TTIC Medigram Darwin Under the Praxen Early Access Program
Highest Praxen RAISE score recorded to date.
Steve Wilson  ·  Inventor, Praxen  ·  Chair AI Security, TTIC  ·  Chair, OWASP GenAI Security Project  ·  Chief AI & Product Officer, Exabeam

What that score reflects is not a clean deployment that was never tested. It reflects a platform that was probed adversarially, produced findings, remediated those findings by code evidence within the same calendar day, and was re-verified. The remediation cycle is part of what the score measures. Darwin's scoring is engineered to be deterministic and auditable so that the result is something enterprises can build governance accountability on, not just a model output.

The categories that required the most work were exactly the ones you would expect in a clinical context: zero trust implementation, supply chain governance, and adversarial testing. Each maps directly to the clinical accountability requirements that make a governed platform defensible in a regulated environment.

The Six RAISE Categories

Understanding RAISE: A Clinical AI Governance Perspective

By Sherri Douville CEO, Medigram · Founder and Chair, TTIC · Trust and Identity Subgroup Co-Chair, IEEE/UL 2933:2024 · Series Editor, Taylor & Francis · Brian Yam Chair, Pro Sports Track, TTIC · COO, Somnology

Each category below includes framing for:

Health System Executive Vendor Clinical AI Developer Pro Sports Executive Physician Investor

The Stadium and the Field

The best analogy for what Praxen and Darwin together certify is a stadium and a field. The infrastructure is certified and ready. The behavioral verification has been conducted. The governance architecture is in place. But the teams still have to coach, and the players still have to play the game.

That is the honest scope of what technical and behavioral certification provides. It verifies that the infrastructure meets the standard required for governed clinical play. It does not play the game for you. The clinical workflows, the human oversight structures, the accountability decisions made by physicians and health system leaders — those are the game. Certification makes the field worthy of it.

Engineering and best medical practice meet technologically and ethically in the IEEE UL 2933 TIPPSS process embodied by the Medigram Darwin program.
Dr. Art Douville  ·  CMO, Medigram  ·  Co-Chair Clinical Integration, TTIC

Why This Changes the Governance Conversation

The Caremark doctrine establishes that boards have an affirmative duty to monitor for material risks. AI agents in clinical settings are a material risk. The plaintiff's bar will ask three questions about any adverse event involving a clinical AI system: what did you know, when did you know it, and what did you do about it.

A Praxen run with zero findings, timestamped adversarial probe results, and a same-day remediation record answers all three questions simultaneously. That is not a compliance document. That is evidentiary infrastructure.

TTIC is currently testing a certification pathway in early access that builds on this foundation. It requires Praxen as Layer 1 because the complementarity between Praxen's behavioral verification and Darwin's governed transaction architecture was what revealed that all eleven governance layers could be covered. The certification was inspired by that realization, not built around either tool independently.

What You Should Do Now

If you are deploying AI agents in any clinical workflow, the first step is defining the governance committee that has authority to sign off on the work remit for a Praxen run. This is not an IT decision. It is a clinical accountability decision. The remit defines the scope of what is being verified, who is accountable for the findings, and what the remediation pathway looks like. Without that structure, a Praxen run produces findings with no accountable owner.

Praxen is free and open source. You can run it today. For information about the TTIC certification early access program and what the pathway requires for your specific platform and use case, contact Sherri Douville directly.

If you're touching patient care, you should be setting the Agentic AI and technical standard.
Sherri Douville  ·  CEO, Medigram  ·  Founder and Chair, TTIC  ·  Co-Chair Trust and Identity Subgroup, IEEE UL 2933

Author & Organizations

The thinking behind this work traces to an organic chemistry professor who taught that the gap between the mechanism you drew on paper and the reaction that actually occurs is where all the important work lives. That lesson is the foundation of everything built here.
Sherri Douville
CEO, Medigram · Founder and Chair, TTIC
Sherri Douville is CEO of Medigram and Founder and Chair of the Trustworthy Technology & Innovation Consortium (TTIC), a non-commercial standards consortium of practitioners and investors across medicine, cybersecurity, engineering, standards, law, operations, and finance. She co-authored ANSI/HSI 2800:2025, the Hospital AI Operations Governance standard, leading the TTIC contribution, and serves as Trust and Identity Subgroup Co-Chair for IEEE/UL 2933:2024. She is a Series Editor at Taylor & Francis, where she edited Mobile Medicine and Advanced Health Technology; six books carry her work across Taylor & Francis, Springer, and Artech House. At Medigram she directs the architecture of Darwin, the company's governed AI decision infrastructure for healthcare. In 2025 she received the AI Champion of the Year award at AIMed25.
Medigram
Clinical AI Governance Infrastructure
Medigram's mission is to deliver AI governance backed by metrics that delivers for patient care and the business. Darwin by Medigram provides governed clinical AI decision support with deterministic, auditable scoring engineered for enterprise clinical accountability. Medigram's CEO founded and chairs TTIC; TTIC maintains independent governance, standards, and decision-making boundaries to preserve the integrity of its consortium activities.

The launch also had meaningful reach. In the first few days post, the Praxen launch generated more than 20 placements across cybersecurity, AI, and enterprise technology media, including coverage in North America, Asia-Pacific, the UK, Europe, and Africa. Medigram as a technical demonstration under TTIC was the only clinical AI governance organization with a practitioner account published on launch day.

Take the Next Step
Ready to verify what your clinical AI actually does?

Praxen is free and open source. The TTIC certification early access program is open to health systems and vendors ready to demonstrate governed clinical AI accountability.

Choose Medigram as a preferred source in eligible Google experiences. Prefer Medigram in Google