Behind every governed AI decision in a clinical or high-performance environment is a person whose name is on the outcome. A physician who cleared an athlete to return. A care team that acted on a recommendation. A specialist who trusted a second opinion. Governance architecture exists to protect those people and the patients who depend on them. When it fails silently, between verification cycles, after deployment, the person whose name is on the record has no defense. That is the problem continuous monitoring solves.
Why Most AI Governance Conversations Stop Too Early
Most AI teams deploying agentic systems in high-stakes environments are not doing verification. The ones that are doing verification are often not doing continuous monitoring. The result is a governance posture that looks defensible at deployment and becomes difficult to defend when something goes wrong six months later.
At Medigram we have worked through what it means to govern a production agentic AI system not just at the moment of deployment but continuously. Here is the architecture we arrived at and why it matters.
Our Overall Governance Posture
Darwin’s governance claims are verified, not self-declared. The framework crosswalk against eleven authoritative governance sources was conducted internally against primary sources. The behavioral verification score was confirmed independently by the creator of the instrument used to produce it, who has no commercial interest in the outcome. We distinguish between those two things explicitly because most AI governance claims do not. The governance model Darwin runs was built in public, through TTIC, with the field. That is not a positioning statement. It is a description of how the standards we implement were developed and by whom.
The full crosswalk is documented on our evidence page.
Verification Answers One Question. Monitoring Answers a Different One.
Darwin achieved RAISE 4.0 Strong on Praxen, independently confirmed by Steve Wilson, Praxen’s creator, who has no commercial interest in the outcome. That score answers a specific question: does this system behave the way its governance documentation claims it behaves, across six behavioral categories, at this point in time?
Praxen re-verification is required before any change affecting behavioral scores is treated as production. That discipline is built into how we operate.
But verification at a point in time is a snapshot. A governed system in production needs continuous confirmation that the behavioral standard is being maintained, not just that it was met when someone looked.
Darwin’s agent fleet uses Observra, an open source agent telemetry and observability framework created by Exabeam. Observra captures, normalizes, and routes runtime activity from AI agents into a standardized telemetry layer, including prompt injection detection, automatic PII redaction, CIM normalization for SIEM readiness, and per-session behavioral tracking. It provides the consistent record of agent behavior that governance authorities require without custom integrations for every agent framework. Steve Wilson, who independently verified Darwin’s RAISE 4.0 Strong score on Praxen and serves as Chair of AI Security at TTIC, is Chief AI & Product Officer at Exabeam, the organization that created and open-sourced Observra. Darwin uses Praxen for behavioral verification and Observra for continuous monitoring. They answer different questions. Together they close the loop. For large enterprises already running Exabeam’s security operations platform, Observra’s CIM-normalized telemetry and Praxen’s behavioral verification findings integrate directly into their existing SIEM infrastructure, making Darwin’s governance record part of the enterprise security fabric they already operate.
Why Continuous Monitoring Is a Governance Requirement, Not an Operational Preference
Four governance authorities are explicit on this point, and each maps directly to what behavioral monitoring addresses in a production agentic system:
IEEE UL 2933 is the governance architecture standard that establishes TIPPSS (Trust, Identity, Privacy, Protection, Safety, Security) as the framework for clinical AI governance. It is the standard cited in CHIME AI Principles as the governance reference for health system CIOs and CISOs nationally. Darwin was built to satisfy it. Its architect co-authored it.
The Indiana Executive Council on Cybersecurity AI Security System Architecture Layers (in.gov/cybersecurity) defines six architectural layers required of AI systems operating in high-stakes environments. It references IEEE UL 2933 explicitly as the data provenance standard and calls for continual monitoring of IT, OT, and AI/ML environments as a foundational requirement. Designed by Mitch Parker, Co-Founder of TTIC, Vice Chair of IEEE UL 2933, and CISO of Indiana University Health, it is the most operationally specific state-level AI security framework published to date. Observra addresses the continual monitoring requirement directly at the agent fleet level.
The Closed Governance Loop
A sealed governance record exists for every governed decision Darwin makes. Praxen independently verifies that Darwin’s behavior matches its governance documentation. Observra monitors behavioral signals continuously and surfaces deviations before they become findings. TTIC, the Trustworthy Technology Innovation Consortium, governs all three without commercial interest in any of them.
Why This Matters to the People Accountable for AI Decisions
This distinction matters to the D&O carrier who wants to know the system was behaving correctly before an incident, not just that it passed a test before deployment. It matters to the regulator who asks whether the governance posture was maintained continuously. It matters to the clinical team whose names are on the decisions the system supports.
What Governed AI Means for the Balance Sheet
For the organization accountable for AI decisions, the closed governance loop is not an engineering achievement. It is a financial one. Four numbers define the exposure:
A sealed governance record created before any of those conversations is a materially different asset than one reconstructed after the fact. Every deviation surfaced by continuous monitoring before it becomes a finding is a procurement conversation that does not stall, a D&O renewal that does not flag, a union grievance that does not require reconstruction under adversarial conditions.
What This Means for the Field
Most agentic AI systems being deployed today have verification or monitoring, but rarely both operating as a closed loop with independent governance oversight. The gap between a verified system and a continuously governed one is where liability accumulates quietly, and where governance claims that looked solid at deployment become difficult to defend when something goes wrong.
Governed AI that stays governed is not a product feature or a promise. It is a commitment that has to be demonstrated every day with proof after deployment.
- OWASP Agentic AI Top 10 (owasp.org)
- NIST AI RMF (nist.gov)
- ISO 42001 (iso.org/standard/81230.html)
- ANSI/HSI 2800:2025
- Indiana Executive Council on Cybersecurity AI Security Architecture Layers (in.gov/cybersecurity)
- Praxen behavioral verification (open-agent-ai-security.github.io/praxen)
- Exabeam Praxen announcement (cioinfluence.com)
- Exabeam Expands Behavior Intelligence to Secure the Agentic Enterprise, July 1, 2026 (businesswire.com)
- Darwin by Medigram evidence page (medigram.com/evidence/)
- Healthcare Analytics Statistics 2026 (knowi.com)
- Risk and Insurance AI Litigation Report, October 2025 (riskandinsurance.com)
- Cherry Bekaert 2025 CFO Survey (cbh.com)
- Menlo Ventures 2025 State of AI in Healthcare (menlovc.com)