Direct Answer: What the Savings Could Be
Clinician low-value care analytics identifies services that provide little or no health benefit for a particular patient or population, including unnecessary imaging, duplicative laboratory testing, overused medications, and procedures that lack a clear indication. Independent estimates commonly place annual U.S. spending on low-value care at roughly $75 billion to $210 billion, although no single figure is definitive because definitions, specialties, and data sources differ. The upper estimate is broader and includes potentially avoidable care, while the lower estimate is concentrated on services that professional societies or evidence-based guidelines identify as routinely overused. Measuring care at the clinician level can reduce this spending, but only when analytics are clinically valid, fair, and connected to useful feedback. The goal is not to rank every clinician by a single utilization number or reward whoever orders less. It is to identify patterns of potentially inappropriate care, verify them with clinical judgment, and support targeted improvement over time.
Also worth reading: How Can I Relieve Headaches Naturally Without Taking Medicine? · How Much Does Low-Value Healthcare Spending Cost, and Can Clinician-Level Measurement Reduce It? · How Does Predictive Analytics Drive Healthcare Cost Control in Modern Organizations?
A credible analytics program combines claims, electronic health records, order-entry data, and patient outcomes. It can compare utilization against evidence-based measures such as Choosing Wisely recommendations, appropriate-use criteria, peer-reviewed research, and local practice patterns. Savings arise when clinicians modify diagnostic or treatment decisions and patients remain equally well served. That distinction matters: eliminating a low-value service is beneficial only if a higher-value alternative, watchful waiting, or no intervention is appropriate. As of September 26, 2026, the technology is more accessible than it was a decade ago, but automated conclusions still require clinical validation and governance.
How Clinician-Level Measurement Works
The process begins with defining a specific clinical behavior rather than applying a generic notion of waste. For imaging, a measure might examine when low-back imaging is ordered without documented red flags or when uncomplicated low-back pain reaches advanced imaging too early. For medication management, it could flag long-term concurrent prescribing where the expected benefit is unclear. Claims data can estimate population-level utilization, while electronic health records are usually needed to assess symptoms, diagnoses, prior tests, contraindications, and documented clinical reasoning. Together, these sources provide a stronger signal than either claims or ordering data alone.
The system then calculates a rate using a defensible denominator. A raw count of advanced imaging orders is not an appropriate-use measure because clinicians differ in panel size, patient complexity, specialty, referral patterns, and service mix. A better denominator may be patient-years, eligible encounters, or cases that satisfy an evidence-based appropriate-use criterion. Many programs apply minimum case thresholds—often 10, 20, or more eligible cases—before displaying peer comparisons, because small denominators create unstable percentages. Results may be shown as peer distributions rather than league tables, with confidence intervals and case mix considered wherever possible. The useful output is a prioritized review queue, not an automated accusation of poor care.
Why Measurement Can Change Behavior—and Why It Sometimes Fails
Clinicians may not know that their practice differs from evidence-based norms, particularly when ordering is distributed across physicians, advanced practice clinicians, residents, and covering teams. A dashboard can expose variation hidden by an organization-wide average and make repeated ordering visible. Timely feedback can then improve selection criteria, order defaults, clinical decision support, and shared expectations. Research has also found an association between greater clinical knowledge and lower use of low-value imaging, suggesting that education and measurement work best when joined. Measurement alone, however, is usually too weak to produce durable change. Programs are more likely to succeed when they combine local data with peer review, workflow changes, and repeated follow-up.
There is also a risk of directing resources away from patients who need more care. Utilizing fewer tests is not the same as practicing conservatively, and a lower-cost outcome can coexist with delayed diagnosis or poor access. Quality measures should therefore run beside efficiency measures: adverse events, emergency visits, readmissions, patient-reported outcomes, complaints, mortality where relevant, and appropriate treatment rates. A clinician can exceed an imaging threshold because they are documenting inadequately, while another can appear efficient because coding is incomplete. The program should measure data quality before interpreting differences. Strong implementations treat analytics as a clinical safety instrument as much as a financial tool.
Practical Steps for a Health System
The first step is to choose one high-volume, well-defined target, such as unnecessary CT imaging in uncomplicated acute sinusitis, duplicate testing within 48 hours, or benzodiazepine use in older adults with fall risk. A narrow start usually produces more reliable results than attempting to score all specialties at once. The organization should appoint clinical owners, verify the underlying evidence, inspect actual workflows, and identify where inappropriate orders originate. Order-entry defaults, routing rules, or pretest reminders can sometimes reduce variation, but they should be designed with clinicians rather than imposed as black-box restrictions. Every intervention should be tested for alert burden, override rates, accessibility, and effects on clinically appropriate care.
A practical measurement cycle runs in three stages: measure, review, and reassess. During the first 1 to 3 months, the team establishes a baseline, checks coding and documentation, and tests for missing data. In months 3 to 6, clinicians review aggregated peer data, examine representative cases, and choose one workflow adjustment. Over months 6 to 12, the organization compares utilization, quality, and experience outcomes with the baseline and decides whether to maintain or revise the program. Many vendor implementations require significant setup because patient matching, clinical terminology, specialty-specific logic, and risk adjustment can be complex. A simpler dashboard using validated measures may be more valuable than an expensive platform with unproven predictions.
| Feature | Clinician Analytics Program | Broad AI Waste Model |
|---|---|---|
| Typical design | Uses defined populations, valid denominators, specialty measures, and peer review | Assigns general cost or risk scores to encounters, orders, or patients |
| Clinical interpretation | Requires clinicians to review diagnoses, indications, and exceptions | May provide a fast signal but may lack local clinical context |
| Best use | Improvement projects for imaging, diagnostics, medications, and procedures | Prioritizing records for review or generating preliminary hypotheses |
| Quality safeguards | Tracks appropriate care, patient outcomes, and documentation completeness | May omit balancing measures and false-positive review |
| Financial expectation | Savings emerge after workflow and behavior change | Apparent savings may disappear after case validation |
| Governance | Clinical leadership, data stewardship, patient privacy, and periodic validation | Greater dependence on the vendor's model and proprietary assumptions |
Organizations do not have to buy enterprise AI to begin. Manual audits of 20 to 50 records per clinician can establish whether a suspected problem is real, while EHR reports can extract high-volume order patterns. Registry data, quality dashboards, utilization-management software, and accountable-care organization shared savings data are other options. These approaches are less scalable and can suffer from inconsistent abstraction, but they are often easier to explain and may be adequate for a focused initiative. The key comparison is not simply price. Buyers should assess validation evidence, integration requirements, ability to produce specialty-specific measures, data ownership, audit access, and the vendor’s willingness to show performance across hospitals rather than only a selected customer.
Implementation costs vary widely because some products sit on top of an existing claims warehouse, while others require data engineering, terminology mapping, security review, and clinical annotation. Custom enterprise deployments can cost hundreds of thousands to several million dollars, whereas narrower dashboard or analytics projects may cost tens of thousands. Subscription, per-clinician, per-facility, and per-record pricing models all exist, and health-system discounts are not standardized. Vendors may claim savings of several million dollars, but buyers should ask whether the estimate represents gross avoided spending, net system savings, or revenue capture from reduced claims. Administrative costs, implementation labor, ongoing monitoring, and the cost of alternate care must be included.
A useful return-on-investment calculation divides verified net savings by software, implementation, and maintenance costs. If a contracted program costs $120,000 annually and produces $180,000 in verified net savings after clinical review and workflow expenses, its gross return on investment is 50%. That calculation should not be presented as guaranteed because avoidable spending is not the same as realized savings: reducing a claim may increase another claim, affect revenue, or expose a service line that was not considered avoidable. Health systems should request audited baselines, confidence intervals, sensitivity analyses, and at least 6 to 12 months of post-implementation data. A pilot with a break-even threshold of 1.0 is reasonable, but safety and evidence quality remain nonnegotiable.
Common Mistakes That Produce False Results
One common error is equating a lower order rate with higher quality. Diagnostic intensity can be appropriate in emergency, oncology, transplant, and other complex settings, and a model trained on average practice may misclassify unusual but necessary cases. Another error is using a generic prior-authorization list as though it were a comprehensive standard of care. Choosing Wisely recommendations are intended to spark conversation and improve judgment, but some apply more broadly than others and do not replace individualized assessment. Claims data can also miss unbilled services, create duplication when clinicians are linked incorrectly, and infer conditions that are absent from the coded record.
Organizations should avoid public dashboards, punitive compensation links, and rankings based on small samples. These practices can encourage gaming, coding changes, documentation inflation, or referral avoidance without improving patient care. Another mistake is installing alerts that every clinician dismisses; even a small override rate can be expensive if a high proportion of alerts are irrelevant. Measures must also account for shifts in team responsibilities, particularly as advanced practice clinicians and other health professionals assume larger roles in U.S. delivery. The safest program distinguishes suspected low-value care from confirmed inappropriate care and records the reason for every exclusion. Independent clinical review and documented model updates are necessary as guidelines, coding, workforce roles, and patient populations change.
When to Act, Pause, or Escalate
Measurement is appropriate when utilization is rising without evidence of better outcomes, a service is a major local expense, or external payment policy creates pressure for defensible efficiency. Organizations should also act when patients may be exposed to avoidable harms from unnecessary imaging, medications, or procedures. A 12-month baseline is often practical for a high-volume service, but urgent safety concerns should be reviewed immediately rather than waiting for a full analytic period. A pilot can be reconsidered if fewer than 20 eligible cases exist per monthly period, missing documentation dominates, or no clinician can explain the measure’s clinical logic. Those limitations do not prove a practice is high quality or low quality; they show that the available measurement is inadequate for a confident conclusion.
Escalation is warranted when savings are accompanied by worsening access, delayed treatment, patient complaints, adverse events, or disparities between patient groups. Health leaders should pause automated restrictions and return to case-level review if a measure repeatedly disagrees with specialty guidance. Expansion should require evidence that appropriate-use rates improved without unacceptable balancing effects, rather than merely proving that spending fell. For accountable-care organizations, shared savings should be assessed at the attributed population level so that fewer services are not mistaken for better total value. A 5% reduction in selected low-value imaging with stable quality is a more meaningful result than a 30% reduction caused by accepting fewer appropriate referrals. Governance, clinician engagement, and transparent review therefore determine whether analytics becomes a useful management system or another source of distrust.
The Bottom Line for Buyers and Clinicians
Clinician low-value care analytics can help U.S. health systems address tens of billions of dollars in potentially avoidable annual spending, and clinician-level measurement can make specific improvement opportunities visible. It cannot determine the whole truth about medical value from utilization data alone, and no estimate of $75 billion to $210 billion should be presented as money that can automatically be recovered. Real benefits depend on validated measures, local workflow, patient characteristics, and the possibility of using a better alternative. The strongest program asks three questions: is the service unlikely to help this patient, is an appropriate alternative available, and are outcomes holding steady or improving?
For a 2026 health-system buyer, a focused pilot is usually the best starting point. Select one measure with a clear evidence base, establish a baseline, involve clinicians in interpretation, and contract for transparent quality results rather than only cost reduction. Report verified utilization change, net savings, patient experience, and balancing measures side by side. Scale only when the improvement persists for at least 6 to 12 months and frontline clinicians regard the measure as clinically credible. Clinicians should not resist all measurement, but they should reject simplistic rankings and unsupported automation. Used carefully, analytics can reduce waste while strengthening appropriate care; used carelessly, it can penalize complexity, narrow access, and obscure the value of decisions that truly matter.