Healthcare policy · Data analysis

The Readmission Penalty Paradox

Medicare penalises hospitals for excess readmissions at a flat percentage of payments. Joining CMS penalty performance for 2,373 US hospitals to their audited financial statements shows the penalty is unrelated to who performs badly — and imposes a 14–18× difference in burden on who can afford it.

CMS HRRP FY2026CMS HCRIS FY2023 2,373 hospitalsJul 2021 – Jun 2024 21 robustness specifications

FindingsPerformance doesn't track finances. Burden does.

The expected finding did not hold. Readmission performance is uncorrelated with financial health — r = , 95% CI , spanning zero. It survives five sets of controls and 21 robustness specifications. Reported as found.
The burden is what diverges. The penalty costs roughly the same dollars everywhere but consumes more of a struggling hospital's annual income than a healthy one's. Because performance is independent of finances, that gap is structural: a uniform rate meeting radically non-uniform balance sheets.

Analysis 01Financially weak hospitals do not readmit more

Three views of the same null, in increasing order of how hard they are to dismiss: the raw relationship, the relationship after controls, and whether the design could have detected an effect at all.

The fitted line is flat, and the confidence band is narrow enough to prove it

Each point is one hospital. The line is an OLS fit; the shaded band is its 95% interval. A 10-percentage-point improvement in margin moves the readmission ratio by — indistinguishable from zero on a measure centred at 1.0.

Adding controls does not move the estimate

The coefficient on total margin at each stage, with 95% confidence intervals. If the raw null were an artefact of confounding, the estimate would move as size, capacity, payer mix, service scope and geography are absorbed. It does not budge from zero. R² rises from 0.000 to 0.142 across the ladder — almost all of it at the state-fixed-effects stage, which is a finding in itself.

This design could have detected an effect six times smaller than nothing

At 80% power and n = 2,373, the smallest correlation this study could reliably detect is 0.058. The observed effect and its confidence interval sit entirely inside the undetectable zone — which is what separates evidence of absence from absence of evidence.

Analysis 02The same penalty costs 2% of income here, 40% there

Hospitals in five equal bands by total margin, modelled penalty 0.5% of net patient revenue. Switch to dollar cost and note how flat it is: the gap is entirely about capacity to absorb, not the size of the bill.

Bars show the median; the thin line spans the interquartile range. The weakest band has no bar because every hospital in it is already loss-making — there is no income for the penalty to be a share of.

How large the gap is depends on choices a reasonable analyst could make differently

The burden multiple under each specification tested. It holds between 14× and 18× across sample restrictions, and the penalty rate is irrelevant by construction. But banding hospitals on operating rather than total margin gives 3.7× — a different question about a different kind of weakness. The headline is therefore a range, not a point estimate.

Analysis 03One condition in six shows a real financial link

Testing all six conditions separately means a 26% chance of a false positive at conventional thresholds, so p-values are corrected for the family. Five conditions are null. Hip and knee replacement — the only elective procedure in the set — survives, in the opposite direction to the equity critique: financially stronger hospitals readmit more.

Only elective joint replacement separates from zero after FDR correction

Correlation between total margin and each condition's excess readmission ratio, with 95% intervals. Benjamini-Hochberg at q = 0.05; the corrected p-value is shown against each condition. Intervals crossing the zero line are nulls.

Flagged, not claimed. Volume and patient selection in elective surgery is the obvious mechanism — stronger hospitals do more joint replacement, possibly on more marginal candidates. This data cannot test that. Without multiplicity control the effect would have been buried among five nulls; with it, the effect is real but the explanation is a hypothesis.

Analysis 04Geography explains seven times more variance than hospital finances

The largest effect in the dataset, and the one with no explanation. States with at least 15 hospitals, ten highest and ten lowest.

New Jersey and Utah are five times apart, and their intervals do not overlap

Share of hospitals performing worse than expected, with 95% Wilson intervals — several states rest on fewer than 25 hospitals, so the intervals matter. The gap between the extremes survives that uncertainty comfortably.

This needs an explanation the data cannot give. State fixed effects lift model R² from 0.019 to 0.142 — geography outweighs every hospital characteristic measured here combined. It does not track margins: Florida hospitals are financially healthy (12.1% median margin) and still perform poorly. Medicare Advantage penetration, post-acute supply and regional admission practice are all plausible; none is testable here. No map is shown, because a choropleth would imply a spatial mechanism this analysis has not established.

Analysis 05A revenue-neutral cap would move $1.3B off the weakest hospitals

Three designs, each collecting the identical $5.70B modelled pool from the same hospitals. Only the distribution changes. This is the question a policy team would ask.

Capping the penalty at 25% of net income shifts 23% of the pool, at no cost to Medicare

Share of the total penalty pool borne by each financial band under each design. Bars are ordered weakest to strongest within each group; lightness distinguishes the three designs so the comparison survives greyscale.

MethodWhat this analysis can and cannot support