Almost every health system now has AI. A study published in early 2026 found that nearly two-thirds of US hospitals using Epic had adopted ambient documentation tools by mid-2025. Adoption is no longer the dividing line.
The dividing line is what happened next. In one group of organizations, AI moved into daily operations, changed a measurable outcome, and stayed. In a much larger group, it produced a portfolio of pilots, a series of encouraging presentations, and no durable change to care delivery or financial performance. RAND research published in 2025 found that across industries, roughly 80% of enterprise AI projects fail to deliver the promised value, with about a third abandoned before reaching production. Healthcare has no special immunity.
Most of the writing about this failure blames the technology or the culture. The more precise explanation is that a large number of organizations bought something the evidence did not support, measured for a benefit nobody had demonstrated, and concluded that AI does not work when what actually happened is that the business case was incorrect from the start.
So the honest place to begin is not with the blueprint, but by sorting out what is actually known.
Sort the Claims Before Evaluating the Tools
Marketing in this category makes three quite different claims, and they have varied evidence behind them. Vendors present them as a single package, which they are not.
The first claim is that clinicians will feel better. A 2025 study in JAMA Network Open covering more than 1,400 clinicians at Mass General Brigham and Emory Healthcare supports this claim. The research found a 21.2% absolute reduction in burnout prevalence at Mass General Brigham at 84 days and a 30.7% absolute increase in documentation-related well-being at Emory at 60 days. Similar findings appear repeatedly across sites and vendors. If you are buying relief from documentation misery, the evidence is on your side.
The second claim is that clinicians will save time. However, this assertion is mixed and weaker than the marketing suggests. Stanford Medicine studied AI-drafted replies to patient portal messages and discovered that clinicians used the drafts about 20% of the time, with statistically significant reductions in task load and burnout and no measurable change in time spent. A systematic review across 23 studies of the same use case concluded that there was no strong evidence of time savings, alongside reasonably consistent evidence of perceived efficiency gains. A separate analysis found meaningful time savings for support staff and close to none for physicians.
Perceived efficiency and actual efficiency diverging is not a contradiction. Doing the same work with less cognitive strain feels faster and is genuinely better, but it does not produce an hour you can fill with a patient.
The third claim is that the organization will save money. This is largely unproven. The Peterson Health Technology Institute assembled eight health systems, including Intermountain, Ochsner, Providence, Mass General Brigham and Yale New Haven, along with ten vendors, and concluded that ambient scribes appear to reduce burnout and cognitive load but have not demonstrated a financial return. The report went further and warned that given the costs and the thin ROI evidence, there is a real risk that continued adoption leads systems to implement these tools in ways that add to the overall cost of care.
Read that twice, because it is the opposite of what almost every vendor deck asserts.
None of this is an argument against the technology. The burnout finding is real and replicated. In a market where replacing a physician is commonly estimated to cost anywhere between $250,000 and $1 million, it can carry a serious business case on its own. The argument is that the three claims must be bought separately, because they are supported separately.
The Tests Worth Applying to Any Claim
A few habits that distinguish leaders who read this evidence well from those who get sold:
Ask whether the benefit was measured or reported. Burnout scores are self-reported, which is appropriate because burnout is a subjective state. Time savings are not subjective, and when a vendor supports a time claim with a satisfaction survey, it is a substitution rather than proof. The Stanford result is valuable precisely because it measured both and found them pointing in different directions.
Ask who was in the study. Pilots run on volunteers, and volunteers are the people most inclined to like the tool. Peterson observed that real-world physician adoption is heterogeneous, typically splitting into heavy users, partial users, and clinicians who try the tool and abandon it. Your enterprise economics depend on the average across all three groups, not on the enthusiasts who joined the pilot.
Ask whether the outcome is the outcome or a stand-in for it. Time in the chart is not margin. Alerts generated are not lives saved. Notes drafted are not throughput. Every step between the surrogate measure and the thing you actually care about is a place where the benefit can leak away, and the leaks are usually organizational rather than technical.
Ask who produced the number. Widely circulated statistics, such as an average return of $3 and change for every dollar invested in AI, generally originate from vendor-sponsored or industry-association surveys of self-selected respondents. Compare that with the Peterson finding, which came from an independent body convening named health systems and reached a far more cautious conclusion. Both are real data but not equally informative.
Watch for the claim that quietly creates new exposure. Peterson noted that ambient vendors are expanding into coding, promising to optimize evaluation and management levels and capture hierarchical condition categories, and that the downstream impact remains unknown. Revenue uplift from documentation change is the claim in this category that deserves the most scrutiny and usually receives the least.
Here, the distinction between accuracy and intensity matters. A tool that surfaces a code the record genuinely supports but a clinician missed is capturing earned revenue. A tool that raises the score without changing what the documentation supports is borrowing against a future audit; in a risk-bearing contract, it is also inflating a benchmark you will be measured against later. On a monthly dashboard, these two look identical, because both show the same line moving up. Under review, they are not remotely the same thing. Any evaluation of a coding tool that reports only the lift, and not the share of that lift a reviewer would uphold, has measured the wrong thing.
An HFMA survey of around a hundred healthcare organizations found that the single most common reason cited for not seeing a return on AI was difficulty measuring the return at all. It is not a technology problem but a consequence of buying before defining what success would look like.
With that established, here is what the organizations that got real results did.
They Started with the Decision, Not the Model
The most instructive comparison in health AI is between two prediction tools that did roughly the same thing and produced opposite results.
Kaiser Permanente Northern California built Advance Alert Monitor to identify ward patients likely to deteriorate, scanning records hourly and aiming to give about 12 hours of warning. Across a staggered rollout, the program was associated with mortality of 9.8% against 14.4% in comparison patients, along with shorter hospital and ICU stays. Researchers estimated more than 500 deaths prevented each year.
The Epic Sepsis Model was deployed widely across the industry. When researchers at Michigan Medicine validated it externally on more than 38,000 hospitalizations, they found an area under the curve of 0.63. It generated alerts for 18% of all hospitalized patients but failed to identify 67% of patients who actually developed sepsis. A later validation in two county emergency departments found that, among alerts for patients who did develop sepsis, half were triggered after sepsis had already occurred. The Michigan authors observed that widespread adoption of a model performing this poorly raised fundamental questions about how sepsis is managed nationally. Worth sitting with.
A tool reached hundreds of hospitals on the strength of distribution and reputation, and the evidence arrived years later from academics rather than from a planned evaluation.
The more useful difference, though, is that Kaiser did not build an alert. It built a response. The score routes to a centralized team of specially trained nurses who are not doing anything else. There are standardized rescue protocols. There is a defined pathway for conversations about goals of care, because many deteriorating patients need a different plan rather than a more aggressive one. The model is the smaller half of the system.
An alert with no named owner, no protocol, and no protected time is not a clinical intervention. It is a notification added to the workload of people who already have too much. So the first question is not how accurate the tool is. It is who acts on this output, within what window, with what authority, and what they stop doing in order to do it. If that has no answer, accuracy is irrelevant, because nothing will happen either way.
They Assumed the Tool Would Not Work Here Until Shown Otherwise
Vendor performance claims are measurements taken on someone else's population, in someone else's workflow, with someone else's documentation habits. They are evidence the tool can work, not that it will work in your hospital.
Duke Health made this operational. Its Algorithm-Based Clinical Decision Support oversight committee established that no new algorithm reaches patients without review, and Duke's published work argues for recurring local validation rather than the one-time external validation borrowed from device approval. Models do not stay accurate. Populations shift, coding practices change, upstream systems get reconfigured, and clinicians adapt to the tool in ways that quietly corrupt its inputs.
The Epic sepsis case shows how subtle this gets. One of the model's inputs was whether a clinician had ordered antibiotics, meaning the model was partly learning from the fact that a human had already suspected infection. That kind of leakage is invisible in a demonstration and obvious in a local validation, if anyone runs one.
They Decided in Advance the Currency They Were Buying
Return to the Stanford result. 20% utilization, no change in time, significant reduction in burnout.
If the business case was clinician hours recovered, that is a failure, and the tool gets canceled. If the business case was retention and cognitive load, it is a clear success. Same deployment, same data, opposite verdicts, decided entirely by a choice made months earlier in a budget meeting.
The arithmetic is worth doing explicitly. Ambient documentation costs roughly $100 to $600 per provider per month. Take a system with a thousand clinicians at $300. That is about $3.6 million a year.
Justified on productivity grounds, it requires each clinician to generate roughly $3,600 more in annual contribution. It sounds trivial until you notice that these tools mostly return time in the evening rather than during clinic. Unless someone deliberately changes the schedule template to convert recovered time into visits, no revenue appears. Many systems expected revenue, changed nothing about scheduling, received well-being instead, and recorded the program as underperforming. The tool did what the evidence said it would do. The business case had been written against a claim the evidence never supported.
Justified on retention, the same 3.6 million looks different. At a mid-range replacement cost of $500,000, preventing seven departures a year covers the entire program, and the burnout evidence for that case is strong.
The failure mode is choosing one currency, measuring another, and finding out after the money is committed.
They Ran Comparisons, Not Demonstrations
MultiCare Health System evaluated ambient documentation by putting roughly 550 physicians and advanced practice providers through a head-to-head comparison of three vendors, measuring time in chart, after-hours work, productivity, and coding outcomes. Their chief medical information officer called it a bake-off.
A demonstration tells you a tool can work somewhere. A comparison at that scale tells you which tool works here, in your specialty mix and documentation culture, and establishes a baseline. You cannot claim improvement against a measurement you never took, which is one reason so many organizations cannot answer the ROI question later.
It is also noteworthy that there are now more than sixty ambient products on the market, with vendors differentiating by expanding into adjacent workflows rather than by core capability. When a category commoditizes that fast, the marginal difference between vendors shrinks and the difference in how you configure and measure grows. The premium is better spent on integration, on tuning to your specialty mix, and on the measurement scaffolding, because those are the parts specific to you and cannot be bought pre-built.
They Picked a Few Things and Stayed with Them
The instinct in a fast-moving field is to diversify. Twenty small experiments; scale the winners. In practice, they produce twenty things individually too small to matter and collectively too large to support. Every pilot consumes integration work, security review, legal review, clinical champion time, and a slot in the change calendar, and these costs are mostly fixed regardless of pilot size. A portfolio of small pilots is the most expensive available way to buy a small amount of value.
The answer is not to stop piloting, but to pilot at the scale of the decision you are actually making, and to stop paying the setup cost twice. Most of what a pilot consumes is scaffolding, meaning integration, security review, measurement, and monitoring, and that scaffolding is largely the same from one use case to the next. Organizations that build it once and reuse it can afford to evaluate properly. Organizations that rebuild it for every pilot cannot, which is why their pilots remain small and their conclusions ambiguous.
This is easier to see when the use cases are related. Coding accuracy, denials management, and value-based care analytics look like three separate programs on an org chart, but they run on substantially the same foundation: clean longitudinal data on the same patients, a reliable view of what was documented and what was paid, and the ability to trace an outcome back to a decision. Treating them as three programs means paying for the foundation three times and getting a thinner version of it each time.
Kaiser's deterioration program began in 2013 and reached full operation across all of its Northern California hospitals by 2019. Six years on one problem. But that is not slowness. It is exactly what redesigning nurse staffing, rescue protocols, escalation pathways, and goals of care conversations across twenty-one hospitals actually takes, and why the results are real.
They Funded the Plumbing Once and the Monitoring Forever
Two cost categories reliably get underestimated.
The first is integration and data. A model that needs manual data assembly in a pilot needs automated assembly in production, and the curation that made the pilot look good is often exactly what cannot be reproduced at scale.
The second is monitoring, and almost nobody budgets for it properly. Deployed models degrade and need surveillance for accuracy drift, performance differences across patient groups, and changes in how clinicians actually use them. MultiCare's stated goal of a single view to track bias, safety, drift, and return across every deployed tool is the right instinct, and it implies a permanent operating cost with a permanent owner. It is harder to get through a capital committee than a new pilot, which is precisely why it distinguishes the lasting systems.
They Gave it to Clinical Operations
Where the AI program reports determines what it optimizes for. An innovation function is measured on novelty, partnerships, and pilots launched. A clinical operations function is measured on outcomes, throughput, and cost. Only one of those is naturally motivated to finish the unglamorous 90% of work of turning a functioning model into a changed workflow.
Every deployment described here had a clinical owner accountable for the outcome rather than for the technology. At Kaiser, it sat with hospital operations. At Duke, the governance committee is co-chaired by clinical leadership. At MultiCare, the evaluation was led by the chief medical information officer and framed around physician well-being, a problem the organization already owned.
The Real Advantage
The blueprint is short:
- Sort the claims by how much evidence each one carries and buy only what is supported.
- Define the decision and the responder before evaluating any model.
- Validate locally before deployment and continuously afterward.
- Name the currency of value in advance and instrument for it.
- Compare vendors at real scale against a real baseline.
- Concentrate on a few problems for years.
- Fund integration and monitoring as shared infrastructure with permanent owners.
- Put all of it under clinical operations.
Very little of this is about artificial intelligence. Discipline is what separates organizations that implement anything well from organizations that implement things repeatedly. Which is the finding hiding underneath all of this.
Model quality is converging quickly and is increasingly available to anyone with a contract. The apparatus around the model, meaning the workflow redesign, governance, measurement, ownership, and patience, cannot be bought off the shelf, and building it from scratch takes years. This apparatus is the durable advantage now. The systems getting results are not winning because they picked better AI. They are winning because they were already the kind of organization that could make almost any capable tool work, and the tools finally became capable.
Which suggests a more useful set of questions for the next steering meeting than the ones usually asked:
Not whether the model is accurate, but who acts on its output and what they stop doing in order to act. Not whether the tool is impressive, but which currency you are buying and who owns that number. Not whether the pilot went well, but what the baseline was before it started. And not whether it works today, but who is watching it in year three. An organization that can answer these four questions will get value out of an average tool. An organization that cannot will not get value out of an excellent one.
These are also the conversations healthcare leaders need to have as AI moves deeper into clinical and operational workflows. At HLTH USA 2026, I will be meeting industry veterans working through the same questions around ownership, governance, workflow design, measurement, and what it takes to make AI hold up beyond the pilot. If those decisions are on your agenda, meet our team in Las Vegas to discover how we are approaching them with healthcare organizations.
November 15 to 18 | Booth 4051, Level 2, Exhibit Hall, AI Zone | The Venetian Expo, Las Vegas

