Clinical AI rarely fails on the day it launches. It fails quietly, months later, as the patient population shifts and the model's assumptions stop matching the people it is scoring.
That is the failure mode almost no governance process is designed to catch, because validation is treated as a gate rather than a schedule. A model clears review once, goes live, and then nobody looks again until something visible goes wrong.
Two things make this urgent rather than theoretical in 2026.
The first is scale. Predictive models are no longer a pilot category. They sit inside the record system clinicians open every shift, flagging deterioration risk, ordering worklists, drafting documentation, and surfacing patients for intervention. More than 96 percent of US hospitals use certified health IT, so whatever the certification programme requires is effectively the national floor.
The second is that the floor may be moving. The HTI-1 final rule established the first federal transparency requirement for predictive algorithms in certified health IT, through the Decision Support Interventions criterion. In January 2026, HHS proposed a rule known as HTI-5 that would scale those criteria back. The comment period has closed and the rule is not final, so nothing is settled. But the direction of travel is worth planning around, because it points the governance burden away from the certification programme and toward the organisations deploying and building these systems.
So this is a runbook rather than a policy explainer. What to capture, what to watch, on what cadence, and what forces a review.
What the current floor actually requires
Worth stating precisely, because it is the baseline you inherit or lose.
The DSI criterion replaced the clinical decision support criterion that had stood since 2012. It defines a Predictive DSI as technology supporting decision-making based on algorithms or models that derive relationships from training data and produce an output resulting in prediction, classification, recommendation, evaluation, or analysis.
For those, certified health IT must surface a defined set of source attributes to clinical users: 31 for predictive interventions, alongside 13 for evidence-based ones. They cover what data trained the model, how it was validated, and how it was tested for fairness. ONC paired this with the FAVES principles, that decision support should be fair, appropriate, valid, effective, and safe.
Two details matter for anyone building in this space. The criterion was drafted deliberately to cover both FDA-regulated software and non-device predictive tools delivered through certified health IT, which is the faster growing category. And where a vendor has not configured source attribute documentation, their health system customers inherit the compliance exposure. Transparency obligations flow downstream whether or not the contract mentions them.
The runbook
Before go-live: capture what you will need later
You cannot monitor drift without a baseline, and the baseline has to be recorded before the model starts influencing care.
Record the population the model was validated on, and how it differs from yours. Record performance by subgroup, not just in aggregate, because aggregate accuracy hides the failures that matter. Record the intended use, stated narrowly enough that off-label use is recognisable when it happens. And record the source attributes the vendor provided, or note precisely which ones they declined to provide.
That last point becomes the whole exercise if the transparency floor is lowered. A record of what you asked and what you received is the difference between a governance process and a hope.
Continuously: watch the inputs, not just the outputs
Input drift precedes performance degradation, and it is far easier to detect.
Monitor the distribution of the features feeding the model against the baseline. A shift in case mix, a new referral source, a change in how a field is coded upstream, or a newly integrated site with different documentation habits will all move the inputs before anyone notices the outputs are worse.
This is ordinary data engineering rather than machine learning work, which is why it sits with backend development and the healthcare data analytics platform layer rather than with whoever owns the model.
Monthly: watch the humans
Override and acceptance rates are the most underrated signal available.
If clinicians stop acting on a model's output, that is information regardless of whether the model's accuracy has changed. Rising override rates mean either the model has degraded or trust has, and both require a response. Falling override rates on a high risk model can be equally concerning, because automation bias is real and uncritical acceptance is not the goal.
Segment this by unit and by role. Aggregate override rates conceal the specific team that has quietly stopped trusting the tool.
Quarterly: revalidate against outcomes
Performance against actual outcomes, by subgroup, on your own population.
This is the step organisations skip, because it is genuinely harder than the others and because nobody is asking for it. It is also the only one that detects the silent failure mode described at the top. A model can pass every input check while steadily becoming less useful for the patients it matters most for.
On trigger: what forces an unscheduled review
Define these in advance, because in the moment there will be pressure not to.
A change in the upstream data source or its coding practices. A new site, service line, or population added to scope. A vendor model update, including ones described as minor. A safety report or near miss involving the tool. And any use that falls outside the documented intended use.
Vendor updates are the one most often missed. A model that changes underneath you has invalidated your baseline, and the release note will rarely say so.
Standing: one named owner
Not a committee. A committee reviews. A person is accountable, and in practice the difference decides whether the quarterly revalidation actually happens.
What to ask a vendor now
If the transparency requirement is scaled back, these questions stop being answered by default and become contractual.
Ask what population the model was trained and validated on. Ask for performance broken down by subgroup rather than in aggregate. Ask what the documented intended use is, and what falls outside it. Ask how you will be notified of model updates, and whether you can decline one. And ask what monitoring the vendor performs on your deployment, as distinct from monitoring their model in general.
A vendor who can answer all five has done the work. A vendor who answers three and is straightforward about the other two is usually a better partner than one who claims all five without evidence. The answers belong in the contract rather than in a sales conversation, particularly if the regulatory backstop weakens.
Where Woltrio fits
Woltrio builds the systems clinical models run inside rather than the models themselves, and in practice that is where most of this runbook lives.
Monitoring infrastructure, input distribution tracking, override capture, and the evidence trail behind a model's outputs are all engineering. They are also the parts that get deferred, because they produce no visible feature and their value only appears when something has gone wrong. That combination is exactly how organisations end up with a model in production and no way to tell whether it still works.
Where an organisation is deploying its own models, AI development and workflow automation work carries the same requirement, and delivering the output into the clinical workflow properly usually means touching the record layer through custom EMR and EHR development.
A closing note on judgement. Regulation here is a floor, not a ceiling, and the floor is currently under review. Building to whatever the minimum turns out to be is a defensible legal position and a weak clinical one. The organisations that will be comfortable in three years are the ones instrumenting now, while the answer to what is required is still open.
Start with a scoped assessment from Woltrio.


