Skip to content

Every dataset, one interface.

Each source is collected, normalised, versioned, and served through the same API, search, and SQL surface. Adding the next one does not change how you query the last.

Live
1
In the pipeline
5
Records served
63,518
Coverage
49 states and DC

Available now

Healthcare charges

What hospitals bill Medicare for outpatient procedures, and what Medicare allows. The gap between those two numbers is the most useful thing in the file.

LiveUpdated annually49 states and DCPublic domain
Centers for Medicare & Medicaid Services · data year 2024
Records
63,518
Facilities
2,993
Procedure groups
68
Widest spread
199×
Charge distribution · Level 2 Urology and Related Servicesmedian $3,137

Query it

curl
curl "https://api.vemon.io/v1/healthcare/prices?procedure=5372" \
  -H "Authorization: Bearer vm_live_…"

Medicare outpatient submitted charges. Not commercial negotiated rates.

The standard

What every dataset here has to meet

These are not aspirations for the catalog. They are the conditions a source has to satisfy before it appears on this page at all.

Normalised to one schema

Field names, casing, types, and identifiers reconciled so the second dataset does not need a second integration.

Provenance attached

Every response names its publisher, data year, and the caveat that applies. A number without its source is a liability.

Gaps disclosed, not hidden

Suppressed rows, absent jurisdictions, and known limitations are published alongside the data rather than discovered later.

Reachable four ways

API, search, SQL, and AI from the day it lands. No dataset gets a bespoke interface.

In the pipeline

What comes next, and what it costs

Listed because they are planned, not because they are available. None of these are queryable today, and the notes are honest about which are hard.

Payer negotiated rates

Planned

The rates commercial insurers actually agreed to — the number that answers what care will cost.

What it takes

Tens of terabytes per payer per month. Deeply nested JSON, provider references in separate files, index files pointing at tens of thousands more. The hard one.

Health insurers and hospitals

FDA drugs & recalls

Planned

Track approvals, label changes, and recalls without polling several unrelated endpoints.

What it takes

Several separate feeds with different update cadences and no shared identifier. Reconciliation is the work.

US Food and Drug Administration

Federal contracts

Planned

Who the government pays, for what, and how the obligation changed over time.

What it takes

Large but well structured. Awards are amended rather than replaced, so history matters more than the current row.

US Department of the Treasury

SEC filings

Planned

Company filings as data rather than as documents, with the XBRL facts extracted.

What it takes

Decades of filings in mixed formats. The structured XBRL era is clean; everything before it is not.

US Securities and Exchange Commission

US Census

Planned

The denominator for almost every other dataset — per-capita anything needs it.

What it takes

Well documented and stable. The difficulty is geography: boundaries change between vintages.

US Census Bureau

Selection

How a dataset earns a place

Two tests — and a request from someone with a real use case beats both.

It has to be public

Published by an agency or required by rule. Vemon does not license proprietary data and does not scrape what a publisher has chosen not to release.

It has to be hard enough to matter

A dataset that is already a clean, documented CSV does not need us. The value is in sources that arrive as nested JSON on an irregular schedule with undocumented suppression.

A concrete use case jumps the queue

“We need X to do Y” travels further than “X would be nice”. Tell us what you would build.

Questions

Asked most often

How many datasets are available today?
One. Medicare outpatient hospital charges — 63,518 records covering 2,993 facilities across 49 states and DC. Everything else on this page is listed as planned, and none of it is queryable yet.
How do you decide what to add next?
Two tests. The data has to be genuinely public, and it has to be hard enough to use that serving it is worth something. A dataset that is already a clean CSV download does not need us.
Can I request a dataset?
Yes, and it is the most useful thing you can send us. A request with a concrete use case behind it moves further up the queue than one without.
Does adding datasets cost more?
No. Every plan includes every live dataset. You pay for request volume, history depth, and monitoring — not for access.
What licence applies to the data?
The underlying data is published by US federal agencies and is in the public domain. Vemon claims no ownership of it; the normalisation and documentation are our own work.

Need a dataset that is not here?

6 sources are named on this page. If the one you need is not, that is worth knowing — a request with a use case behind it is the most useful thing you can send.

Catalog last generated