Low-value care is generally estimated to account for roughly 30% of U.S. health spending, or about $1 trillion annually when expressed in today’s healthcare dollars. That figure is not a precise invoice total: definitions of low-value care differ, estimates include both overuse and potentially avoidable spending, and some services classified as wasteful in one population may be appropriate in another. Nevertheless, the scale makes clinician-level measurement a serious operational opportunity rather than a minor efficiency project. Measuring individual clinicians and clinical teams can reveal persistent variation in imaging, laboratory testing, medication prescribing, preventive services, referrals, and follow-up intervals. It cannot, by itself, determine whether a clinician is providing poor care, because patient needs, local practice patterns, access constraints, and documentation quality affect utilization. Its practical value is diagnostic: it identifies where additional data, workflow redesign, clinical decision support, or performance feedback may be needed.
As of September 2026, health systems are also evaluating AI-based tools that summarize utilization patterns, surface suspected low-value orders, and recommend evidence-linked alternatives. Such systems may reduce administrative burden, but the evidence does not support treating an AI-generated flag as a final judgment about medical necessity. A defensible program combines measurement with clinical review, transparent criteria, patient context, and a clear process for overriding an alert. The following sections explain how low-value care creates waste, what clinician-level measurement can detect, and what a safe implementation plan should look like.
Also worth reading: How do organizations accurately calculate AI benefits consulting ROI measurement in healthcare settings? · How Can AI Healthcare Benefits Reduce Employee Costs and Improve Access in 2026? · What Are the Biggest Healthcare AI Privacy Risks and How Can Health Organizations Reduce Them?
What Low-Value Care Is—and What It Is Not
Low-value care includes services for which the expected benefit is small, uncertain, or lower than the associated cost, time, or risk for a particular patient. Examples can include unnecessary imaging for uncomplicated low-back pain, duplicate laboratory testing, medications that do not improve outcomes for a defined population, or screening performed outside an evidence-based interval. The category also covers potentially avoidable hospital transitions, excessive specialty referrals, and follow-up testing that would not change management. Low-value care is not synonymous with all high-cost care. An expensive treatment may be the right choice for one patient and wasteful for another, while a low-cost service can still expose a patient to unnecessary risk or delay appropriate care.
The best-known U.S. estimate is that approximately one-third of healthcare spending is low-value care. Published estimates vary because researchers use different time periods, populations, and attribution rules, and spending figures are often converted into newer dollars. State-level research in commercially insured and Medicare Advantage populations has also shown that measured low-value utilization differs substantially across regions. That variation is useful for identifying questions, but it should not be used to rank hospitals or clinicians without adjustment. A clinician who treats older, sicker, or more medically complex patients may appropriately order more diagnostic services than a clinician serving a healthier population.
A sound program therefore distinguishes three categories: care that is clearly unnecessary, care that may be appropriate but requires review, and care that is clinically justified. This classification is essential for avoiding false penalties. The goal is not simply to reduce service counts. It is to reduce ineffective or duplicative care while preserving access, patient choice, diagnostic accuracy, and clinically appropriate treatment.
How Much Does Low-Value Care Cost the U.S. System?
The commonly cited estimate of approximately $1 trillion per year provides a useful description of the opportunity size, but it should not be presented as money that could be recovered immediately. A large portion of the estimated amount consists of opportunity costs rather than discrete line items: clinician time, facility expense, patient time, adverse effects, and downstream treatment. Eliminating even a small percentage of low-value utilization would have a large effect on aggregate spending, yet a system should not assume that all identified spending is recoverable. Money saved by omitting a test may reappear in another appointment, a different diagnostic pathway, or a later episode of illness.
Cost estimates also vary depending on the service examined. Avoiding low-value imaging can reduce imaging facility charges, interpretation expense, downstream testing, and incidental findings that require follow-up. Reducing unnecessary antibiotics may avoid adverse drug events and resistance, although the immediate pharmacy-cost saving may be modest. Preventing a low-value hospitalization can produce a much larger direct saving, but the organization must distinguish a genuinely avoidable admission from one caused by disease progression or limited outpatient capacity. Likewise, shortening an unnecessary course of therapy may save relatively little while still improving patient experience and reducing risk.
A useful business case separates gross spending identified from net savings that are actually realized. Gross opportunity equals the cost of services flagged by a validated measure. Net savings subtracts replacement services, implementation costs, clinician training, software expense, and the value of any care moved rather than removed. It also should account for the possibility that a measure identifies overuse without producing a budget reduction because saved capacity is used for other demand. Even when immediate savings are limited, reducing low-value care can release clinical capacity and reduce patient burden, which are legitimate benefits but should be reported separately from cash savings.
Can Clinician-Level Measurement Actually Reduce Waste?
Yes, but measurement is an intervention only when it is connected to action. A dashboard that ranks clinicians by low-value service volume may motivate improvement, but it may also encourage underuse of necessary care, cherry-picking, or documentation changes that make performance appear better without changing practice. Evidence described in Medical Economics and other clinical literature suggests that greater clinical knowledge is associated with fewer low-value imaging orders. This association is not a simple causal proof, because knowledge interacts with specialty, setting, patient mix, and organizational culture. Still, it supports combining feedback and education rather than relying on financial penalties alone.
Clinician-level measurement has four main advantages over system-level analysis. First, it can distinguish a broad regional problem from a specific service line, such as excessive repeat imaging in one emergency department. Second, repeated patterns are more actionable than isolated orders. Third, comparing peer clinicians can support a focused discussion about criteria, workflows, and exceptions. Fourth, local analysis can reveal whether a widely cited measure performs poorly in the organization’s actual patient population. A system may select several high-burden measures, establish baseline rates, and review them monthly or quarterly rather than attempting to measure every form of low-value care.
Results should be interpreted over time and in context. One unnecessary test by a physician is not a reliable measure of quality; persistent, adjusted patterns matter more. Suitable metrics can include the percentage of uncomplicated low-back-pain patients receiving early imaging, duplicate test rates within a defined window, or the share of patients receiving a medication without a documented indication. Denominators, exclusions, and risk adjustment must be specified. For example, imaging for a patient with red flags should not count against a measure designed for uncomplicated cases. Without careful measurement design, apparent savings can be created by excluding difficult patients rather than improving care.
Clinician Measurement Versus Alternative Approaches
Organizations can address low-value care at the population, institutional, or individual-clinician level. Population approaches are cheaper and simpler, while clinician-level approaches can target variation more precisely but introduce more data, governance, and potential gaming risks. Institutional interventions—such as changing default imaging protocols, modifying electronic order sets, or restructuring discharge processes—may produce more consistent results than education directed only at physicians. The best approach usually combines levels rather than selecting one exclusively.
| Feature | Clinician-level measurement | System or facility-level intervention |
|---|---|---|
| Primary purpose | Identify meaningful variation among clinicians and teams | Standardize a service, workflow, or facility policy |
| Typical data need | Attribution, denominators, risk context, specialty, time trends | Aggregate service volume, orders, outcomes, and capacity |
| Main advantage | Precise targeting of feedback and education | Faster implementation and less risk of misranking individuals |
| Main weakness | Can penalize appropriate care or encourage gaming | May hide important differences between teams and clinicians |
| Common intervention | Peer review, coaching, decision support, selective feedback | Order-set defaults, staffing changes, protocol redesign |
| Best use case | Persistent, unexplained variation within similar teams | Broadly accepted inappropriate-use pattern |
Patient-directed alternatives include shared decision-making, evidence-based self-management, and choosing conservative treatment when outcomes are similar. These are valuable but slower and harder to scale than a default-order change. Conversely, a top-down restriction may be efficient but unsafe if it removes a clinically useful option. A mixed strategy is usually strongest: simplify the appropriate path, make the low-value path harder to select without reflection, and preserve an easy route for justified exceptions.
How to Build a Practical Clinician-Level Measurement Program
The first step is to choose a narrow clinical problem with meaningful volume, credible evidence, and an accountable owner. A health system should not begin by attempting to construct a universal quality score. Instead, it can select one service or condition, such as imaging for uncomplicated low-back pain, and define eligible patients, excluded cases, the observation period, and the expected standard. Baseline data should then be reviewed with clinicians who understand the workflow. A high rate may indicate overuse, poor documentation, coding problems, or a legitimate difference in patient risk. Interviews and sample chart reviews are necessary before conclusions are drawn.
Next, the organization should test whether existing electronic health record data can reproduce its chosen measure. Data quality, duplicate records, missing problem lists, and inconsistent order attribution can distort results. A practical review may examine the last six to twelve months, use monthly or quarterly cohorts, and require a minimum denominator before displaying a rate. It should also establish an exception pathway for patients with relevant complications. No single threshold is universally correct, but changes that are statistically unusual, persistent for two or more periods, and associated with similar patient outcomes are more credible than small month-to-month fluctuations.
Only after validating the measure should the system introduce interventions. These may include evidence-linked order sets, duplicate-order alerts, peer review, concise educational feedback, or workflow changes that make the recommended action easier. Improvement should be monitored using process measures, balancing measures, and—where feasible—patient outcomes. A program that lowers imaging by 15% but increases delayed diagnosis, emergency visits, or later treatment has not necessarily succeeded. Target reductions should therefore be interpreted as hypotheses to test rather than quotas imposed in advance.
Where AI Healthcare Benefits Consulting Can Help
AI can support clinician-level measurement by mapping data, generating utilization summaries, identifying patterns across time, and drafting educational or documentation feedback. These tasks can be time-consuming for analysts, particularly when the same suspected low-value behavior appears across several locations. AI may also help compare service-specific patterns with evidence-linked criteria and route cases for human review. That is a different function from autonomous clinical decision-making: the analyst determines the measure and population, the AI processes information, and qualified clinicians determine whether action is justified.
The technology still has important limitations. Models can misclassify symptoms, miss outside records, use outdated evidence, or produce inconsistent explanations. They can inherit biases if historical utilization reflects unequal access or inappropriate prior treatment. They may also create alert fatigue if every possible exception is presented to a busy clinician. A benefits consultant should therefore evaluate data readiness, clinical governance, integration, privacy, and expected workload—not only whether a vendor’s model reports a high accuracy score. A concise pilot with one specialty and one measure is more informative than a broad demonstration using synthetic scenarios.
Consulting services vary widely in price. A narrowly scoped workflow review may cost several thousand dollars, while a multi-site analytics implementation, clinical governance program, software integration, and ongoing monitoring can range from tens of thousands to several hundred thousand dollars. Annual support and maintenance may be added. There is no universal market price because effort, data complexity, software licensing, and required clinical review differ. Before buying, request a transparent statement of deliverables, hosting model, security terms, validation results, implementation timeline, and whether the quoted fee includes ongoing measure recalibration. A useful pilot should have predefined success criteria, such as data completeness above 95%, agreement with chart review, and measurable adoption, rather than promising a fixed percentage of national savings.
Common Mistakes That Make Measurement Backfire
One common mistake is equating low utilization with high quality. Fewer tests are not automatically better if clinicians are failing to diagnose disease or defer necessary care. Another is using a raw count rather than a rate. Comparing ten flagged orders from a large oncology practice with three from a small primary-care office tells the organization little. Rates must have defensible denominators, and complex patients should not be removed merely because they make the rate less attractive. Risk adjustment can help, but it is imperfect and should be documented rather than presented as a perfect correction.
A second mistake is launching financial penalties before clinicians understand the data. Surprise rankings damage trust and encourage gaming, especially when attribution rules are unclear. It is also easy to confuse a potential savings estimate with realized savings and to omit the cost of replacement services. A third mistake is selecting measures because they are easy to compute rather than because they matter clinically. A dashboard may accurately count duplicate tests while ignoring a much more consequential pattern, such as avoidable admissions. Leaders should prioritize burden, potential benefit, actionability, and risk of harm rather than ease of extraction alone.
Finally, measurement itself can become low-value care if it generates more documentation and review than benefit. Organizations should limit the number of measures, retire metrics that do not influence decisions, and audit whether feedback changes practice. Transparency is especially important when a consultant, health plan, or AI vendor uses proprietary scoring. Clinicians should know when their data is being analyzed, how cases are flagged, what evidence supports the measure, and who can correct errors. A program that cannot explain its output should not be used to impose penalties or restrict care.
When to Act—and How to Know Progress Is Real
Acting is most appropriate when a validated measure shows persistent overuse, the evidence-based alternative is clear, the organization can collect reliable data, and an owner is empowered to change the workflow. Many programs can begin with a three-month data review, a six-month pilot, and a twelve-month outcome assessment. That timeline is not a universal rule, but it provides enough time to establish a baseline, correct attribution issues, train users, and observe downstream effects. Immediate large-scale deployment is rarely wise when the baseline is uncertain.
Progress should be evaluated at several levels. Process measures show whether the selected service fell, duplicate orders declined, or guideline-concordant alternatives increased. Balancing measures reveal whether referrals, emergency visits, complications, diagnostic delays, or patient complaints worsened. Experience measures can assess whether patients felt their concerns were heard. Financial measures should report gross identified spending, verified avoidable spending, and net retained savings separately. For a target metric, a practical pilot might seek a 10% to 20% relative reduction in a well-defined overuse measure, but the appropriate target depends on baseline performance and should not be copied mechanically from another organization.
The measurement program should be recalibrated at least annually as evidence, coding, and care patterns change. Expansion should occur only after the first measure is trusted and produces a demonstrable result. If a measure has no actionable ownership, inconsistent data, or evidence of unintended harm, it should be revised or retired. The strongest long-term program is not the one that flags the most orders. It is the one that helps clinicians deliver more appropriate care with less waste while preserving the ability to explain every exception.
The Bottom Line for U.S. Healthcare Leaders
Low-value care likely consumes on the order of one-third of U.S. healthcare spending, but the estimated $1 trillion is a broad opportunity estimate rather than an immediately recoverable savings pool. Clinician-level measurement can help by identifying persistent variation, supporting peer review, and revealing where workflow or decision-support changes may be useful. It cannot judge medical necessity in isolation, and apparent differences may reflect patient mix, access, documentation, or local capacity. Therefore, the most defensible conclusion is not that every high-utilization clinician is wasteful or that AI can automatically remove the spending.
A successful approach begins with a focused clinical question, a validated denominator, transparent exclusions, and a baseline established with frontline clinicians. It then combines local feedback with system-level changes, such as better order sets and streamlined access to appropriate services. AI healthcare benefits consulting can reduce analysis and design effort, but independent validation, human oversight, privacy controls, and a clear appeal process remain necessary. When these elements are present, clinician-level measurement can reduce low-value care safely; without them, it risks becoming another administrative burden or an inaccurate incentive. The central test is whether spending, clinical burden, and patient outcomes improve together rather than whether a dashboard records fewer services.