Learn
Use Case

Healthcare Spending Doesn't Scale Linearly: A Decile Analysis of Cost and Risk

4 min read·July 6, 2026· MEPS dataset →
Chart of weighted healthcare expenditure and risk scores by normalized total expenditure decile, MEPS 2019-2023

Everyone in health economics knows healthcare spending is skewed — a small share of people account for most of the cost. But "skewed" is a vague word. I wanted an actual number: how much more does the top decile of spenders cost than the decile right below it, and does that gap show up in independent risk-adjustment scores or just in the raw dollars?

I asked VerbaGPT one question against MEPS — the Medical Expenditure Panel Survey, the U.S. government's household health spending survey — pooled across 2019–2023, person-level, ~126,000 rows:

"Create deciles for normalized total expenditure. Then put together a table that shows the weighted normalized total expenditure for the deciles, the actual weighted average total expenditure, and also the average normalized concurrent and prospective risk scores. Give me a bar or line chart or some other visual that relays this info well."

What "Risk Score" Means Here

MEPS carries ACA HHS-HCC risk-adjustment scores at the person level — the same family of models insurers use to price risk pools and CMS uses to adjust Medicare Advantage payments. Two flavors matter here:

Both are normalized so the population average is 1.0. A score of 2.5 means "predicted to cost 2.5x the average person" — independent of what was actually spent.

The Method

MEPS is a household survey, not a census — every record carries a person-weight so the sample reflects the U.S. population. Getting deciles right meant sorting by normalized expenditure, then cutting at boundaries of cumulative survey weight (not raw row count), and weighting every average by that same weight. Dollar figures also needed inflating to 2026 dollars via the medical CPI, since spending was pooled across five separate survey years.

None of that was written by hand — it came out of the one-sentence question above, resolved against the schema description VerbaGPT already had for the table (column meanings, weighting conventions, which columns needed the CPI multiplier).

The Table

DecileWeighted Norm. Total Exp.Weighted Actual Total Exp. (2026 $)Weighted Norm. Concurrent RiskWeighted Norm. Prospective Risk
10.000$00.3440.413
20.006$490.3920.438
30.037$2820.4340.484
40.083$6310.5180.561
50.153$1,1640.6420.683
60.265$2,0160.7820.828
70.448$3,4111.0021.027
80.786$5,9931.2701.262
91.589$12,1081.7631.700
106.634$50,5642.7072.456

Key Findings

Why the Top-Decile Gap Matters

This is the exact tension that risk-adjustment payment models (ACA marketplace transfers, Medicare Advantage) are built to handle: predicted risk and realized spend should move together, or the entity holding the risk gets paid for a population it isn't actually covering. A dataset where risk and cost track linearly for 90% of the population but diverge at the top decile is precisely where risk-adjustment models are hardest to get right — and where insurers' margins are most exposed to a handful of outlier patients.

Data source: Medical Expenditure Panel Survey (MEPS), Household Component, Full Year Consolidated files, pooled 2019–2023 (AHRQ). Risk scores are ACA HHS-HCC concurrent and prospective model scores, normalized to a population mean of 1.0. Expenditures adjusted to 2026 dollars via the medical care CPI.

Ask your own question about this dataset → MEPS is available as a live public demo datasource on VerbaGPT — no signup required for a few questions.


Related