Attestation tells you what someone says exists. Adversarial behavior tells you whether it works.
AI is starting to help with decisions about people's health, in hospitals, and in pro sports, where teams use data about players' bodies to decide who plays and when. Darwin by Medigram is the system that checks the AI: it scores every decision on six kinds of trustworthiness and keeps a permanent record, so whether the stakes are a patient's care or a player's career, nothing happens unchecked and there's always proof. Designed to work on mobile, where the work happens.
For security and enterprise risk, the key distinction is between documentation about controls and evidence of actual system behavior. Consequential AI needs boundaries that can be challenged, findings that remain visible, and evidence that can be independently re-verified.
A short reader path into Building Darwin, the full CEO Letter.
The security test for governance
TTIC’s security note draws a hard line between a requirement that a policy exist and an engineering constraint on records, authorization and evidence. The latter can be tested.
Darwin’s control philosophy
The trust boundary sits before the model. Validation is recorded. A failed check marks the permanent record rather than erasing the event. Behavioral monitoring continues after deployment.
A finding is an input, not an endpoint
A security or risk finding that terminates in a report has done its job on paper but not in practice, if nothing downstream (architecture, policy, implementation, verification, or the evidence standard itself) is required to change in response. This architecture treats a finding as an input to all of those layers, not an endpoint.
Why Sherri and the Medigram team
That’s possible today because the same person who defines the security boundary also has standing in the architecture, the implementation, and the evidence design, so a finding travels rather than stalling at a departmental line. The goal is to make that traversal a property of Medigram’s process, not a dependency on one person’s ability to walk a finding across four different teams.
Independent behavioral evidence
The CEO Letter states that Darwin achieved a Praxen RAISE score of 4.0 (Strong) and was reported by the Open Agent and AI Security Community as the highest behavioral score recorded to date as of August 2026.
Risk question to carry forward
Can the organization prove which model or agent acted, what it was allowed to see and do, whether the controls were operating, what failures were found, how they were remediated and whether the evidence can be independently verified later?
Need the full architecture and evidence?
The CEO Letter contains the detailed technical rationale, verification loop, standards context and diligence path.