From comparison to a tested change
Peer benchmarking can help a CMO identify where patient outcomes warrant investigation and choose a change to test. The comparison must distinguish a potential performance gap from differences in case mix, measurement, and care setting. External measures from CMS or clinical registries become more useful when clinical teams understand the definitions and have a routine for examining variation, assigning responsibility, and evaluating a response.
“A dashboard can identify a low percentile without establishing what the service line should do next.”
Leaders may disagree about whether the comparison reflects their patients, whether the data are reliable, or how quickly an outcome could change. A request for fewer preventable harms needs an investigation and an implementation plan, with a timeframe appropriate to the measure.
Treat benchmarking as a recurring clinical management responsibility. Agree on definitions and comparable peers, identify a clinical and operational owner, and turn a concerning result into a specific question the team can investigate. An annual report alone leaves too much time between recognizing a problem and testing a response.
If clinicians doubt the peer group's relevance, adding more results may not resolve the objection. Address definitions, peer selection, and risk adjustment early. A legitimate comparability concern should change the analysis; an assumption that does not hold up to examination should not indefinitely postpone action.
Peer selection requires clinical and analytical judgment together. Compare referral patterns, service mix, patient characteristics, and care settings; labels such as academic or safety-net are starting points, not sufficient matches. Two peer sets can help: an aspirational group for learning and an operationally comparable group for informing targets. Assess those targets against clinical standards and local improvement opportunities as well. A comparable group's median is neither guaranteed to be attainable nor necessarily good enough.
Make the risk-adjustment method open to review. Show inclusion criteria, exclusions, relevant present-on-admission logic, and sensitivity to changes in case mix. Clinical and analytical teams can then examine whether the inputs and assumptions fit the population, rather than repeatedly debating an unexplained result. CMS publishes methodology resources for hospital measures reported through Care Compare, providing a reference for the applicable specifications and adjustments.
A process measure describes whether a care activity occurred; an outcome measure describes a result such as complications, function, or survival. Report them separately so improved adherence is not mistaken for improved patient outcomes. The process change is a proposed means of improvement, and its relationship to the outcome still needs to be assessed.
Consider a service line that appears in the bottom quartile. The clinical leader challenges comparability, analysts explain the adjustment, and the meeting ends before anyone agrees on a next step. If the same discussion repeats at the next review, the problem is no longer just the disputed result. The team lacks a process for resolving the dispute and deciding what to examine in care delivery.
Separate questions about data quality from questions about clinical practice, assign an owner to each, and set a date for returning with findings. That allows the team to investigate a potential care problem while checking the measurement, without treating either explanation as settled in advance.
Share the proposed definitions and peer cohort before the results discussion. At the working session, select a manageable number of possible contributors to variation, perhaps two or three, and agree on what can be tested within 90 days. Keep the patient consequence explicit, but allow enough scrutiny of the denominator and adjustment to establish whether the comparison is valid.
“A benchmark identifies variation; it does not by itself establish the cause.”
Care-model design, staffing, discharge reliability, and coding differences may each contribute to a result. Before funding a remedy, examine which explanation is supported by chart review, workflow observation, or a focused test. Otherwise, the intervention may address a plausible cause that is not responsible for the local gap.
Tie the comparison to a defined care-delivery concern, such as complications or unreliable transitions. ACS NSQIP uses clinical data from patient records to report risk-adjusted surgical outcomes through 30 days after surgery. That provides a specified basis for comparison. It does not remove the need to examine case selection, data quality, and the local process that might explain a difference.
Mortality and readmission rates can look satisfactory while patients report poor recovery of function or quality of life. ICHOM's condition-specific outcome sets include patient-reported measures alongside clinical outcomes and relevant case-mix variables. Collecting those measures can reveal aspects of recovery that mortality and readmission measures do not capture. The team still needs a plan for interpreting responses and acting on them.
HCAHPS provides a standardized, publicly reported comparison of patients' experience of hospital care. It can give clinical leadership, operations, and the board a common reference for questions about that experience. Use it alongside measures suited to the clinical concern; it does not represent every aspect of care quality or establish why an experience score changed.
A 90-day cycle can organize a bounded improvement test. During the first 30 days, confirm definitions, peer cohorts, and the clinical and operational owners. By day 60, investigate the selected variance areas and begin testing a response. By day 90, review implementation evidence and the outcome data available, then decide whether to continue, adapt, expand, or stop the test. Some outcomes need longer follow-up. Keep definitions consistent within the comparison period, but document corrections when a genuine measurement error is found.
Make that cycle part of a scheduled quality operating review. Involve physician and nursing leaders in measure governance, review disputed cases, and keep a usable data dictionary available. These practices give people a way to challenge an error and understand a valid finding. Agreement on a measure supports action; it is not a substitute for showing whether care improved.
When an internal disagreement remains unresolved, a neutral clinical peer reviewer may help assess comparability and definitions. Bound the review to a particular measure set and an explicit question. The useful output is a documented conclusion about the disputed method and the next analytical or clinical step.
Ask peers what they changed, what they tried unsuccessfully, and what else changed during the same period. How did they assess competing explanations for the result? How did they distinguish better documentation from a change in care? What did clinicians stop doing to make room for the new practice? Removing a redundant handoff or documentation step may create capacity, while other improvements require additional work. The comparison should make that trade-off visible rather than assume improvement always comes from subtraction.
Start with one measure linked to a recognized clinical concern, establish a defensible comparison, and assign a team to investigate the variation. Keep the review active long enough to evaluate a change in practice and its effect on patients. A higher percentile is useful only when the team understands what changed, whether the result is credible, and whether the improvement can be maintained.
Confidential peer exchange can make it easier to discuss failed tests, resource constraints, and implementation choices that a published result does not explain. Connex convenes CMOs and CNOs in Think Tanks and Trusted Peer Groups for that kind of comparison.