AI is starting to help with decisions about people's health, in hospitals, and in pro sports, where teams use data about players' bodies to decide who plays and when. Darwin by Medigram is the system that checks the AI: it scores every decision on six kinds of trustworthiness and keeps a permanent record, so whether the stakes are a patient's care or a player's career, nothing happens unchecked and there's always proof. Designed to work on mobile, where the work happens.
Beyond CEO of Medigram, I author accredited standards. I set the architecture. I do the AI engineering: pinned models, drift canaries, adversarial input handling, and I built and designed the monitored agent fleet.
1. The question
If an AI made a care decision about someone you love, and a mistake was made, wouldn't you want to know what happened?
Most of the time today, that is not possible. Not because anyone is hiding it. Because the record was never made.
2. What Darwin is, plainly
Darwin is a system that writes down how a decision was made, at the moment it is made, and seals it so nobody can quietly change it afterwards.
That is the whole idea. What informed the call. What the AI was allowed to contribute. Who was accountable. Sealed then, not reconstructed later by someone assembling screenshots under pressure.
Two things about it are worth stating before anything else.
It has been independently tested. Darwin by Medigram achieved a Praxen RAISE score of 4.0 (Strong) as of June 22, 2026. The Open Agent and AI Security Community reported in August 2026 that Darwin "posted the highest Praxen behavioral score recorded to date."
It is built on standards, from inside them. IEEE/UL 2933:2024 lists me in its Participants section as Trust and Identity Subgroup Co-Chair, which means I co-chaired one of the subgroups that wrote the standard's technical content. I am listed among the individual Standards Association balloting group members for that standard, the body that votes it into existence.
The standard is moving from paper into the documents that shape purchasing. IEEE/UL 2933-2024 with TIPPSS is named in the Health Sector Coordinating Council's Model Contract-Language for MedTech Cybersecurity, version 2, as a framework a supplier may attest to following, and CHIME's Principles for Responsible AI describes TIPPSS as a cross-standard safety model. TIPPSS has reached the contracting layer: the language health systems put in front of suppliers, and the principles their CIOs are given.
Darwin is the implementation of the governance discipline I have been helping write, and what happens when that discipline becomes executable.
3. Why Sherri and the Medigram team built it
The precise version of the problem is this. Conventional AI governance produces documentation about a system: policies written, controls configured, certifications collected, questionnaires answered. All of that describes intent. None of it is evidence of what actually happened at the moment a specific decision was made.
The gap between the two is invisible right up until the day it is the only thing that matters. Then someone asks how the decision was made, and the honest answer is a reconstruction.
So I built the opposite: not a system that documents governance, but one that produces the evidence as a byproduct of the decision itself. If the record is created at the moment and place of the decision, it never needs to be reconstructed.
Medigram began in secure, reliable mobile clinical communication. The humans are mobile, the work is mobile, and the systems supporting that work have to be mobile too. That is the line from there to here, and it is told properly in the Medigram story. Darwin is where it arrives: the decision is made in motion, so the evidence has to be made in motion.
Darwin began as an internal training tool. I built it to explain the infrastructure behind our AIMed 2025 demonstration. At AIMed 2025 I was Lead Architect for "From Policy to Code: Executing Standards at Scale," the operational capstone of TTIC's High-Reliability AI Module Track. AIMed's Hall of Fame lists me among its 2025 winners as "AI Champion Non-clinician of the year." The training problem and the production problem turned out to be the same problem: both require you to state, in advance and in writing, what the system is permitted to contribute and who remains accountable. I have the git history.
4. What I designed
I architected Darwin as a layered system in which the governed record is the product, and the clinical recommendation is one field inside it.
1ModelClick here to learn what this means
2PolicyClick here to learn what this means
3GovernanceClick here to learn what this means
4Security controlsClick here to learn what this means
5VerificationClick here to learn what this means
6Evidence and auditClick here to learn what this means
A governed decision record is produced end to end.
Four decisions define it. Each had a cost, and I want to be explicit that I chose them rather than inherited them.
It never withholds a record. Whatever arrives, a sealed governance record is produced, and where the input does not support a clinical recommendation the record says so rather than going quiet.
The alternative was to fail closed. I rejected it. A system that goes silent under stress destroys evidence exactly when the evidence matters most.
Validation is recorded, not enforced as a gate. Every sealed record carries the outcome of its own validation, and a failed check marks the record rather than suppressing it.
A flag on a permanent record is a stronger accountability instrument than a block that leaves no trace.
The boundary sits before the model. What the model is allowed to see is decided before it sees it, not corrected afterwards.
Disagreement is engineered in. The platform is built to surface its own uncertainty rather than resolve it silently, and it does so by a deliberate deviation from the standing recommendation rather than an implementation of it. I took a controlled comparison I can reason about over a diversity claim I could not defend. How that comparison is constructed is in the diligence materials, not here.
The same discipline points outward, at the vendors. Darwin screens AI vendors across the six TIPPSS dimensions of IEEE/UL 2933:2024, trust, identity, privacy, protection, safety and security, and every screening produces one of four designations with a sealed record either way. That answers a question a buyer asks before mine: should this tool be in the building at all, and is it still doing what it was sold as doing. For organisations where that is the first question, the detail is on the pro sports page.
The operational controls around them are verified by behavioral test rather than asserted in documentation. A sealed governance record's input hash can be independently re-verified on demand. That is the property the whole system exists to provide.
5. How we all know it works
I do not ask anyone to take the architecture on description. My claim is about the loop, not about perfection.
1SpecificationClick here to learn what this means
2Adversarial testClick here to learn what this means
3FindingClick here to learn what this means
4RemediationClick here to learn what this means
5RegressionClick here to learn what this means
6EvidenceClick here to learn what this means
↻The evidence re-enters the specification.Click here to learn what this means
That loop has run. Fourteen distinct remediation identifiers appear across the commit messages of the 223 commits in the repository. Two near-miss entries are logged and tracked in the same record.
The loop's value is that it also records what it has not solved. Four controls are currently unresolved, and each one is written into my claims register with the same standing as the successes: named, dated, and owned. I will walk a counterparty through all four.
The loop is also fast, which is the part people assume governance costs you. Thirteen behavioral findings were opened and closed in the darwin-src record, twelve of them within five days, with a median cycle time of zero days. Governance did not slow the work down. It is the reason the work can be shown at all.
6. What I do technically
I run the company. This section is the technical side of that job, not the whole of it, and it is the part people find hardest to place.
I do six technical things, and in most organizations they are six different people.
I author accredited standards. I set the architecture. I do the AI engineering: pinned models, drift canaries, adversarial input handling, and I built and designed the monitored agent fleet. I do the security engineering at the boundary. I design the data architecture that holds the evidence. And I do assurance engineering for clinical AI, which is the claim, the evidence, and the reasoning that ties the two together.
I am the sole human author of all 223 commits in the repository; no commit is authored under any other person's name.
I did not set out to do all of it. I tried to get the standards implemented, first by organizations and then by individual engineers, and they could not get it done. Either the requirement was understood and the implementation never followed, or something shipped that satisfied the letter of the requirement without satisfying its intent. So I rolled up my sleeves and did it myself, with agent assistance. The working system is the demonstration that it can be done.
That work has a name, and I gave it one on stage in 2025: policy to code. Reading a standard closely enough to know what it actually demands, building the thing that satisfies it, and designing the test that tells the difference between satisfying a requirement and documenting it. Those three normally sit in three different institutions, and the handoff between them is exactly where implementation dies.
What the machine does. I work with AI tooling and agents for implementation: writing code against a specification I set, running the adversarial probes, executing migrations, drafting documentation I then correct. Medigram operates its agent fleet from a written operating playbook and per-agent charters. The judgment stays with me. Which controls are gates and which are flags, where the trust boundary sits, what the system is forbidden to contribute, what gets sealed into the record, whether a finding is closed. Every architectural decision in section 4, including the ones with costs, is mine, and I can defend each against the alternative I rejected.
How that shows up in the record. Of those, 125 of 223 (56.1%) carry a documented AI co-author trailer and 98 (43.9%) do not. That split correlates with author identity rather than date.
How this compares to current practice. AI-assisted engineering is now ordinary at every serious software company, and a 56% assisted rate would be unremarkable in 2026. What is not yet ordinary is disclosing it in the commit trailers, auditing the result, and putting the audit in the diligence pack. I would rather you read the split here than find it yourself in a repository I handed you.
How the system is watched. Behaviour in production is checked on a schedule against what it is supposed to do, and every result is written down whether or not anyone is looking.
It is also probed from outside by an independent behavioral verification instrument I do not control, which is where the Praxen score comes from. Continuous monitoring across the platform and the agent fleet runs underneath all of it.
And beyond running the company, we support the ecosystem with the digital infrastructure underneath it. I personally design, architect and ship it. I built and shipped medigram.com and trustworthytechnologyinnovation.com, alongside the governed platform and the records system that tracks what may be said and what is withheld. I authored more than 500 commits across six repositories in 2026.
The record now updates monthly. August 2026: 437 commits, bringing 2026 GitHub contributions to 651, across the repositories that carry Medigram's and TTIC's digital infrastructure. The point is not any single number. It is that the infrastructure both organizations run on, the sites, the platform, the records systems, the quality gates, is designed, shipped, and maintained continuously.
The sites carry their own enforceable standard rather than borrowing Darwin's. Both are held to a quality budget checked on every public route in mobile emulation before every deploy: accessibility 95, best practices 100, SEO 100, and performance 95 on medigram.com and 93 on the TTIC site. A route below budget fails the build. You can check that one yourself without asking me for anything.
7. The standards loop
Standards and implementation discipline each other, and I run both.
Standard → implementation → verification → evidence → standard evolution. A standard states what trustworthy behavior requires. An implementation either satisfies it or reveals that the requirement was written without a workable mechanism behind it. Verification tells you which, because a specification that cannot be behaviorally tested is an aspiration. The evidence then feeds back: where a real system cannot comply is precisely where the next revision needs work.
Why an accredited standard, and not a framework. Medicine already has a way of deciding what to trust, and it is accreditation. Accreditation does not check whether the teaching is correct; it checks how it was made: who was in the room, what their interests were, how disagreements were settled, who could object. Accredited standards development runs on the same machinery. Not every document called a standard is made that way, and accreditation is what separates the ones that are. TTIC was built as a standards consortium rather than a think tank, an advisory practice, or a research institute because that is the only form medicine already has a reason to trust. The full account is in TTIC's origin story.
That is also the answer to why Darwin records how a decision was made rather than only what was decided. It is the same test medicine already applies to everything else it trusts, turned on the software.
I am a working group member for ANSI/HSI 2800-2025, Healthcare Organization Management and Artificial Intelligence Governance in Healthcare Operations, representing Medigram. The working group is responsible for the technical content of the standard, and working group members co-author standards.
I am Series Editor of the Trustworthy Technology & Innovation in Healthcare book series, commissioned by Taylor & Francis. I founded the Trustworthy Technology & Innovation Consortium. It grew organically out of the Trustworthy Technology & Innovation in Healthcare book series commissioned by Taylor & Francis, which I created and edit, and I have chaired it since its formation.
When I cite a standard I use its exact form, and that form is IEEE/UL 2933:2024. Citation discipline is the first place governance credibility is won or lost.
8. What we deliberately do not publish
Some things are withheld, and I would rather name them than let their absence look like an oversight:
- Darwin operates against a written behavioral remit that is advisory only and preserves named human decision authority at all times.
- Darwin is evaluated against versioned benchmark question sets, with more than one version able to be active at a time.
- Darwin's verification scores are produced by shared, versioned grading machinery rather than per-run judgement.
- Darwin keeps each evaluated entity in a fully separate table set, stamped from one canonical schema template.
- Darwin's outbound egress surface and injection pattern set are documented internally and withheld from external disclosure.
Each is withheld deliberately, the disclosure decision is mine alone, and each has a review trigger recorded against it.
9. How to verify this
None of the above asks for trust on assertion.
A dated engineering chronology exists for the repository, with every figure traceable to a specific command and its raw result. A claims register records each externally usable statement against a named verification method, and carries the negative findings alongside the positive ones. A withheld register records what is not disclosed and the single sentence that may be said about each item.
These are available to serious counterparties under NDA, and I will walk any technical diligence reader through them directly.
The claims in this letter are maintained under Medigram's internal claims verification register; supporting diligence materials are available to qualified parties under NDA.