News

AI Says It Saved Nurses 15 Minutes. Now Prove What Happened Next.

Part of Nurse.org’s Nursing AI Watch, our ongoing investigation into how artificial intelligence is reshaping nursing practice. This is the conclusion of The Nursing Value Project, nurse and clinical AI strategist Mark Smith’s three-part series. Part one asked how healthcare learned to value software before nursing; part two found the idea that never became infrastructure. This final installment, built on interviews with researchers, informaticists, and policy experts, shows how nurse leaders can turn nursing capacity into evidence that survives Monday’s budget meeting.

Suppose an AI vendor tells a hospital that its product will save every nurse 15 minutes per shift.

The arithmetic arrives quickly. Multiply 15 minutes by the number of nurses, shifts, units, and days in a year. Convert the total to full-time equivalents. Put a dollar sign in front of it.

The result can look substantial.

But what changed in the hospital’s budget?

Did overtime fall? Did the organization use fewer agency hours? Did a unit safely care for more patients? Did discharges occur earlier? Did nurses spend the time on surveillance, education, mobility, care coordination, or the work they had previously been unable to complete? Did the technology create new work through alerts, exceptions, corrections, escalation, or oversight?

If no one can answer those questions, the hospital has not measured value. It has multiplied a time estimate.

The first two articles in this series asked how healthcare learned to value software before nursing and why decades of nursing-intensity research never became a durable nursing valuation framework. This final article asks the question that follows:

What can a nurse leader do now?

The answer is smaller than a new national reimbursement system, but more useful on Monday morning.

Choose one decision where nursing evidence is missing. Build the evidence needed to change that decision. Record the result in a form that finance, operations, researchers, payers, and policymakers can examine.

Want to see more Nurse.org articles in your Google results? Add us as a preferred source.

Much of the current discussion begins with a reasonable frustration: healthcare is developing coding and payment pathways for some algorithmic services while the work of bedside nurses remains financially difficult to see.

That comparison exposes a real asymmetry. It can also send nursing toward the wrong solution.

Olga Yakusheva, PhD, an economist and professor of nursing at Johns Hopkins University and lead author of the published Nursing Human Capital Value Model, argues that nursing belongs in a shared-production model. Most bedside nurses are employees paid by the hour, although employment arrangements vary. A hospital stay is produced by nurses, physicians, pharmacists, therapists, facilities, equipment, and many others. Trying to isolate a separate price for every contribution can be artificial, administratively expensive, and vulnerable to the same volume incentives that have troubled fee-for-service medicine.

The problem is not hourly pay itself. The problem is that nursing compensation is visible as labor expense while nursing’s contribution to the hospital’s shared output is not.

That distinction changes the starting point. Nurse leaders do not have to prove that one nurse alone caused an outcome or attach a price to every assessment, escalation, and teaching encounter. They do need to show that nursing capacity is a necessary input to outcomes the organization already values, and that changing the input changes what the organization can reliably produce.

The practical target is not individual ownership of a patient’s outcome. It is a fair share of organizational attention, authority, and reinvestment.

Yakusheva described the first step as awareness: leaders must recognize that nursing generates organizational and financial value, not only quality and patient value. That awareness should lead to measurement, reporting, and eventually negotiation for the resources required to grow that value.

This is a subtler argument than “make nursing billable.” It may also be more achievable.

The urgency comes partly from policy developments surrounding clinical algorithms.

The American Medical Association is reviewing a proposed category called Clinically Meaningful Algorithmic Analyses, or CMAA. Separately, in the CY 2027 hospital outpatient proposed rule, the Centers for Medicare & Medicaid Services proposed the term Software as a Medical Service, or SaMS, for certain software-based technologies that support clinical decision-making through algorithmic analysis. CMS also proposed a new outpatient status indicator, O1, that would identify designated SaMS services for separate payment through an ambulatory payment classification. Nurse.org has reported in depth on both the CMAA proposal and the SaMS designation.

None of that means AI has acquired a universal Medicare revenue category. The AMA is reviewing a coding category. The CMS policy is a proposed outpatient payment pathway for designated services. It does not apply to every AI product, payer, patient, or setting.

For inpatient hospital care, Medicare generally makes a prospectively determined payment per discharge rather than separately reimbursing each itemized nursing service. Making nursing visible therefore does not necessarily require converting every nursing activity into a fee-for-service claim. It could also mean carrying patient-level nursing intensity more accurately into the bundled payment and the decisions made within it. (CMS overview of IPPS)

See also  Charity to build network of MND research nurses

Robert Longyear, a digital health and payment policy analyst and author of “A Virtual Care Blueprint,” emphasized an important sequence during an interview for this article: a code can identify a service, but identification does not establish coverage, valuation, or payment. CMS and other payers still have to decide whether a service is covered, how it is packaged, what resources it consumes, and what rate or payment status applies.

That warning is supported by a broader concern. In a 2026 Health Affairs policy analysis, Longyear and Robert Berenson argued that Medicare’s physician fee schedule was not designed to value software well. AI can change clinician time, vendor costs can be difficult to observe, and developer pricing can be mistaken for resource cost. A payment rate that exceeds the technology’s actual clinical and economic value could reward volume without improving care. A technology that reduces clinician work but receives no appropriate recognition could face the opposite problem. The authors call for clearer coverage standards, better resource valuation, and cost-effectiveness analysis rather than payment based on enthusiasm alone.

Nursing should learn from that warning before seeking an equivalent set of codes.

Financial visibility without disciplined valuation can produce overpayment, fragmentation, and incentives to generate more billable activity. The goal should not be to reproduce those defects around nursing. It should be to prevent nursing from remaining invisible when health systems decide where capacity, capital, and the gains from technology will go.

Nursing does not have to wait for a new national code to begin producing better evidence.

Researchers have measured patient-level nursing time, intensity, skill mix, cost, and variation for decades. The earlier work reviewed in part two showed that patients within the same diagnosis-related group can require very different amounts of nursing care. That research also showed that flat room charges, unit averages, and DRG averages do not carry that variation reliably into the financial record.

John Welton, PhD, RN, FAAN, Professor Emeritus at the University of Colorado College of Nursing, offered a useful way to organize what should be measured. The starting point is not staffing in the abstract, but the individual patient’s nursing-care needs and the nursing time, skill, and experience available to meet them. Leaders then need to distinguish the care a patient is expected to need from the care actually delivered.

The first protects against mistaking deprivation for efficiency. If a unit is short-staffed and care is missed, measuring only delivered activity could make the patient appear less resource-intensive precisely because needed care did not occur.

The second protects against paying for a theoretical need without examining what was provided. Expected need and delivered care answer different questions. Neither is a complete measure of value.

That distinction also reveals why common nursing measures cannot be treated as interchangeable:

  • Acuity estimates future need.
  • Documentation records selected events.
  • Workload describes demand placed on a nurse or team.
  • Staffing describes available labor.
  • Cost describes resources consumed.
  • Value connects resources with outcomes that matter.

A number from one category does not become another because it appears on a dashboard.

A 2017 Providence St. Joseph Health presentation supplied by clinical informatics nurse Matthew McCann (presentation on file with the author) shows how much can already be built inside an electronic health record. Its nursing workload acuity tool used more than 180 rules drawn from orders and nursing documentation. The score was designed to estimate work expected on the next shift, help distribute high-acuity patients, and identify unusual care needs.

The project involved more than 125 expert nurses across five states, a 45-day soft launch in six California hospitals, and more than 2,000 pieces of feedback. Nine months after implementation, the team conducted a seven-day, 14-shift validation study. The presentation was explicit that the score was one input to staffing, not a measure of work already performed or disease severity.

That may be the most important part of the example.

A score can make nursing demand visible at the patient, nurse, unit, and shift levels. It cannot decide what should happen next. It does not establish the correct staffing level, a dollar value, causal improvement, or a payment amount.

It is also inseparable from the documentation that generates it. During the Providence soft launch, many issues first classified as build problems turned out to involve documentation and training. That finding cuts both ways. Better documentation can improve the tool. It can also create the illusion that documentation completeness and patient-care completeness are the same thing.

They are not.

A 2024 mixed-methods study of nurses’ EHR use found that staffing, patient load, redundancy, navigation, and lack of time affected documentation burden across acute and critical-care units. A documentation-derived score may therefore reflect the patient, the workflow, the interface, local policy, and the nurse’s opportunity to chart.

Contemporary validation research offers another caution. In a prospective ICU study, agreement among nurses, nurse managers, and physicians about workload intensity was only partial, and predicting the highest-intensity patients was harder than predicting lower-intensity patients. The cases most likely to demand action may also be the cases a model has the most difficulty classifying.

Stephanie Witwer, PhD, RN, NEA-BC, FAAN, a consultant to the AAACN Center for Innovation and Excellence and author of a two-part Nurse.org series on nursing’s measurement and billing infrastructure, said a measure must be reliably collected, meaningful to nurses, comparable across genuinely similar practice models, and relevant to the decision being made. Large datasets do not repair a measure that does not represent the work leaders think it represents.

See also  Are VA and Federal Nurses Losing Their Union Rights? Trump's New Executive Order Explained

The safest interpretation is simple: a dashboard is a hypothesis about nursing work. Local validation determines whether leaders should trust it for a particular decision.

Most measurement projects begin by asking what data are available.

That is often how another dashboard gets built.

This article’s interviews point toward a different first question: What decision will change if the evidence crosses a defined threshold?

Consider an AI documentation tool proposed for one medical-surgical population. The hospital should identify the decision owner and the possible actions before the pilot begins. The decision might be to expand the tool, modify the workflow, renegotiate the contract, reinvest released capacity, or stop the project.

Then the leaders involved should agree on what evidence would justify each action.

Stephen A. Ferrara, DNP, FAAN, a nurse practitioner, past president of the American Association of Nurse Practitioners, and founder of an AI literacy academy for nurses, described five tests for a credible AI value claim:

  1. Confirm patient outcomes remain at least equivalent across patient subgroups. A model that performs worse for some patients can preserve inequities already present in the system. At the scale of one unit over one quarter, uncommon events such as falls or failure to rescue should be treated as harm surveillance; there will usually be too few events to prove improvement.
  2. Compare performance with an observed pre-implementation baseline. Use direct observation or work sampling across different shifts and census levels. Staffing, census, and EHR changes can occur alongside implementation, making attribution to the technology alone difficult.
  3. Measure net time after subtracting new work created by the technology. Include note editing, alert review, exception handling, and override documentation.
  4. Identify a real budget action or measurable volume change. Examples include reductions in agency hours or overtime, a vacancy deliberately left open, or more paid volume moving through the same staffing.
  5. Confirm that the tool improved clinicians’ work life. Measures might include documentation time after the shift or intent to leave among clinicians on the affected units.

Ferrara connected those tests to the quintuple aim: better patient experience, better population health, lower cost, better clinician work life, and health equity. A tool that advances only one aim while leaving the other four unexamined has not established its overall value. Cost reduction may be the easiest benefit to claim, but it is the weakest evidence on its own.

This prevents the most common accounting error in the AI discussion. The same nursing hour cannot be counted once as cash savings and again as additional clinical capacity.

If a tool reduces overtime expenditure, the hospital may have a realized financial gain. If nurses use the time for care that previously went undone, the hospital has gained capacity. If no spending changes but future agency labor is avoided, the result may be cost avoidance. If the time disappears into the normal variability of a shift, there may be no demonstrable organizational gain at all.

Each result can still matter. They are not the same result.

Nurse leaders may find that the missing value begins upstream of the workload dashboard.

Robert Wingo, BSN, RN, NI-BC, a nursing informatics specialist and founder of Perceptive Staffing Innovations, proposed a basic consistency check for nursing budgets. Wingo works on nurse staffing, nursing finance, and workforce analytics as a writer, researcher, and commercially, and has an interest in how nursing’s contribution is measured and valued. He is speaking here as a nurse and independent analyst, not on behalf of any product.

Assume a unit needs 80 direct-care FTEs and 20 percent of total paid time must cover leave, education, orientation, and other support activity. Total FTEs would be calculated as 80 divided by 0.80, which equals 100. Adding 20 percent of 80 produces 96. The second calculation leaves the budget four FTEs short of its own assumption because it treats a percentage of total time as though it were a percentage of direct-care time.

The arithmetic does not prove a unit is understaffed. It does not establish bad faith, attribute a patient outcome to the shortfall, or solve a workforce shortage. It does expose an internally inconsistent plan worth reviewing jointly with finance.

An archived 2017 Veterans Health Administration staffing directive offers a historical example of the principle. It required staffing calculations to include a replacement factor for leave, education, and systems-improvement activities, with review by nursing, finance, and human resources. It should not be represented as current VHA policy, but it shows that time away from direct assignments can be treated as a planned resource rather than a failure of productivity.

This matters because support time is often the first capacity consumed when a budget is too tight. Education, shared governance, orientation, quality improvement, breaks, and leave appear nonproductive until their absence produces turnover, agency use, overtime, brittle staffing, or safety risk.

The budget can therefore erase nursing value before a shift begins. It can fund the visible assignment while borrowing invisibly from the capacity that makes the assignment sustainable.

>> Read Robert’s full analysis in The Math Error Hiding Inside the Nursing Shortage.

A nurse leader does not need a universal nursing price to change the conversation with finance. The leader needs a decision-ready evidence chain.

Bring one page with eight fields:

  1. The decision: What will leadership approve, change, expand, reduce, or stop?
  2. The population: Which patients, unit, episode, and time period are included?
  3. Expected nursing need: What care should this population require, and how was that estimate validated?
  4. Delivered nursing care: What care, surveillance, coordination, and documentation actually occurred, including missed or delayed work?
  5. Clinical consequence: Which outcome could plausibly change within the test period?
  6. Financial consequence: Which existing expense, revenue protection, throughput measure, or capacity constraint recognizes that outcome?
  7. Action threshold: What result will be sufficient to trigger a decision, and who has authority to act?
  8. Disposition of the gain: Who controls any released capacity or financial benefit, and how will reinvestment, redeployment, or reduction be decided?
See also  Hospital Fires Staff After Allegations of Sharing Patient Photos and Using AI to Mock Them

The financial consequence should be one the organization already uses. Overtime, agency labor, avoidable days, closed beds, delayed discharge, premium shifts, turnover exposure, and lost volume can be more persuasive than an invented price for nursing activity.

Start with a problem the CEO or CFO is already trying to solve. Then show where nursing capacity sits inside it.

That approach also changes the retention argument. Turnover, agency use, overtime, and closed capacity are not merely proof that nursing is expensive. They can become a counterfactual: what does the organization spend or fail to produce when stable nursing capacity is absent?

Yakusheva described a de-identified example in which a chief nursing officer sought a one-year retention investment. The organization modestly increased pay and added transition-to-practice coaching and retention incentives. She reported no turnover among nurses hired during that year, an almost 80 percent reduction in agency staffing, and a positive first-year return. The organization and underlying figures cannot be independently audited, so this should not be treated as a published case study. Its logic is still useful. The CNO connected the request to an organizational problem, defined an intervention, and followed outcomes finance could recognize.

The first pilot should be deliberately narrow.

  • Days 1-15: Define the decision. Select one recurring decision, one patient population, one workflow, and one accountable executive. Record how the decision is made today. Set the possible actions and thresholds before examining results. Agree in advance on who controls any released capacity or financial benefit and how competing uses will be decided.
  • Days 16-30: Specify the evidence. Define expected need, delivered care, one clinical outcome, and one financial consequence. Establish the baseline. Record implementation, licensing, training, oversight, and workflow costs. For AI, name the task changed, the nurse oversight required, the exceptions generated, and where released time is expected to go.
  • Days 31-75: Test the measure. Monitor missing data, documentation burden, work transfer, safety, equity, and subgroup differences. Compare the score with frontline nursing judgment and observable demand. Preserve null and contradictory findings. A midpoint review should include bedside nurses, not only analysts and executives.
  • Days 76-90: Make the decision. Present the findings in the same format used for the original staffing, capital, or operating decision. Record whether leadership expanded, modified, stopped, or deferred the intervention. Publish the definitions, implementation cost, missing-data problem, and reasons for the decision.

Ninety days will rarely prove that nursing caused a complex outcome. It can show whether a measure is usable, whether an AI time claim survives contact with workflow, and whether leadership will act when nursing evidence becomes visible.

That is decision evidence. It is not yet a national valuation model.

A 2025 scoping review found many models for measuring or billing nursing care, but no universally accepted approach. That is not a reason to stop. It is a reason to make local evidence more transferable.

Every pilot should report enough detail for another organization to understand what was measured:

  • Patient population and exclusions
  • Expected need and delivered-care definitions
  • Nursing intensity and skill mix
  • Baseline and comparison method
  • Clinical and financial outcomes
  • Technology version and human oversight
  • Implementation and downstream costs
  • Missing data and documentation limits
  • Action threshold and decision taken
  • Negative, null, and unintended results

Without those elements, nursing produces stories. With them, nursing begins producing infrastructure.

The AI payment debate is important because it shows how economic categories are made. A service is defined. Boundaries are drawn. Evidence standards develop. A code may follow. Coverage, valuation, packaging, and payment decisions come later.

Nursing does not need to copy that pathway exactly. It does need to participate before technology purchasing and payment models decide that an algorithm’s output is economically legible while the nursing capacity required to interpret, supervise, and act on it remains an undifferentiated expense.

The most useful question for Monday morning is therefore not, “What is a nurse worth?”

It is this:

What decision is the organization making without seeing nursing clearly, and what evidence would make that decision and its consequences harder to ignore?

Start there.

The Nursing Value Project (complete series):

More from Nurse.org’s reimbursement series:

🤔What’s one decision your hospital made this year where nursing evidence was missing, and what would it have taken to change it? Share in the comments below.

If you have a nursing news story that deserves to be heard, we want to amplify it to our massive community of millions of nurses! Get your story in front of Nurse.org Editors now – click here to fill out our quick submission form today!

Nurse.org Analysis

  1. Published on

    September 8, 2026

    Written by

    Mark Smith, MBA, BSN, RN Critical-Care Nurse | Nursing Value, Clinical Judgment & AI | Research & Advisory

Source link

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button