Medicare penalises hospitals for excess readmissions at a flat percentage of
payments. Joining CMS penalty performance for 2,373 US hospitals to their audited
financial statements shows the penalty is unrelated to who performs badly — and imposes a
14–18× difference in burden on who can afford it.
The expected finding did not hold. Readmission performance is uncorrelated with
financial health — r = , 95% CI , spanning zero.
It survives five sets of controls and 21 robustness specifications. Reported as found.
The burden is what diverges. The penalty costs roughly the same dollars everywhere
but consumes more of a struggling hospital's annual income than a
healthy one's. Because performance is independent of finances, that gap is structural:
a uniform rate meeting radically non-uniform balance sheets.
Analysis 01Financially weak hospitals do not readmit more
Three views of the same null, in increasing order of how hard they are to
dismiss: the raw relationship, the relationship after controls, and whether the design
could have detected an effect at all.
The fitted line is flat, and the confidence band is narrow enough to prove it
Each point is one hospital. The line is an OLS fit; the shaded band is its
95% interval. A 10-percentage-point improvement in margin moves the readmission ratio by
— indistinguishable from zero on a measure centred at 1.0.
Adding controls does not move the estimate
The coefficient on total margin at each stage, with 95% confidence
intervals. If the raw null were an artefact of confounding, the estimate would move as
size, capacity, payer mix, service scope and geography are absorbed. It does not budge
from zero. R² rises from 0.000 to 0.142 across the ladder — almost all of it at the
state-fixed-effects stage, which is a finding in itself.
This design could have detected an effect six times smaller than nothing
At 80% power and n = 2,373, the smallest correlation this study
could reliably detect is 0.058. The observed effect and its confidence interval sit
entirely inside the undetectable zone — which is what separates evidence of absence
from absence of evidence.
Analysis 02The same penalty costs 2% of income here, 40% there
Hospitals in five equal bands by total margin, modelled penalty 0.5% of net
patient revenue. Switch to dollar cost and note how flat it is: the gap is entirely about
capacity to absorb, not the size of the bill.
Bars show the median; the thin line spans the interquartile range. The
weakest band has no bar because every hospital in it is already loss-making — there is no
income for the penalty to be a share of.
How large the gap is depends on choices a reasonable analyst could make differently
The burden multiple under each specification tested. It holds between
14× and 18× across sample restrictions, and the penalty rate is irrelevant by
construction. But banding hospitals on operating rather than total margin gives
3.7× — a different question about a different kind of weakness. The headline is
therefore a range, not a point estimate.
Analysis 03One condition in six shows a real financial link
Testing all six conditions separately means a 26% chance of a false positive
at conventional thresholds, so p-values are corrected for the family. Five conditions are
null. Hip and knee replacement — the only elective procedure in the set — survives, in the
opposite direction to the equity critique: financially stronger hospitals readmit more.
Only elective joint replacement separates from zero after FDR correction
Correlation between total margin and each condition's excess readmission
ratio, with 95% intervals. Benjamini-Hochberg at q = 0.05; the corrected p-value
is shown against each condition. Intervals crossing the zero line are nulls.
Flagged, not claimed. Volume and patient selection in elective surgery is the
obvious mechanism — stronger hospitals do more joint replacement, possibly on more
marginal candidates. This data cannot test that. Without multiplicity control the effect
would have been buried among five nulls; with it, the effect is real but the explanation
is a hypothesis.
Analysis 04Geography explains seven times more variance than hospital finances
The largest effect in the dataset, and the one with no explanation. States
with at least 15 hospitals, ten highest and ten lowest.
New Jersey and Utah are five times apart, and their intervals do not overlap
Share of hospitals performing worse than expected, with 95% Wilson intervals
— several states rest on fewer than 25 hospitals, so the intervals matter. The gap between
the extremes survives that uncertainty comfortably.
This needs an explanation the data cannot give. State fixed effects lift model
R² from 0.019 to 0.142 — geography outweighs every hospital characteristic measured
here combined. It does not track margins: Florida hospitals are financially healthy
(12.1% median margin) and still perform poorly. Medicare Advantage penetration,
post-acute supply and regional admission practice are all plausible; none is testable here.
No map is shown, because a choropleth would imply a spatial mechanism this analysis
has not established.
Analysis 05A revenue-neutral cap would move $1.3B off the weakest hospitals
Three designs, each collecting the identical $5.70B modelled pool from the
same hospitals. Only the distribution changes. This is the question a policy team would ask.
Capping the penalty at 25% of net income shifts 23% of the pool, at no cost to Medicare
Share of the total penalty pool borne by each financial band under each
design. Bars are ordered weakest to strongest within each group; lightness distinguishes
the three designs so the comparison survives greyscale.
MethodWhat this analysis can and cannot support
Join. CMS FY2026 HRRP (18,330 hospital-condition rows) to CMS HCRIS audited cost
reports FY2023 on CMS Certification Number. 2,373 of 2,392 eligible matched (99.2%).
Validation gate. HCRIS columns are raw worksheet coordinates; misreading a line
number produces plausible but wrong figures. Every derived field is checked against
published national benchmarks and the pipeline halts on failure.
The null survives controls. Five-stage ladder — size, occupancy, payer mix,
service scope, state fixed effects — leaves the coefficient at −0.00001
(p = 0.88), HC1 robust errors throughout.
Robustness. 21 specifications across five families. The null holds throughout.
The burden multiple ranges 14–18× by sample and 3.7× under operating-margin
banding, so it is reported as a range.
A discarded variable. HCRIS Medicaid patient days were the intended social-risk
proxy. They gave a median share of 4.7% against a published ~19%. Dropped rather than
rescaled — adjusting it would have manufactured agreement and any equity finding would
then be an artefact of that adjustment.
Modelled, not observed. A uniform 0.5% penalty is applied. Actual payment
adjustment factors are in the CMS HRRP Supplemental Data File. The claim is "an
identical penalty lands unequally", not "these hospitals received these penalties".
No system affiliation. A thin standalone margin inside a large system understates
real capacity. The most likely source of overstatement in the burden figures, and not
observable here.
No causal claim, and no patient-level inference. A flat correlation rules out a
simple story, not offsetting mechanisms that cancel. All analysis is hospital-level.