The bill, by industry
Four sectors, four vocabularies, one failure. In each case somebody is paid to reconstruct meaning that was known when the data was created and was not written down in a form the next system could use.
Health administration
The largest single category of waste in US health care is administrative complexity, at $265.6 billion a year. That figure is from Shrank, Rogstad and Parekh in JAMA, October 2019, in 2019 dollars. The number is not the interesting part. Every other waste domain in that paper comes with an evidence-based intervention and a projected saving. For administrative complexity the authors write that no studies were identified that focused on interventions targeting it, and the paper's headline savings range explicitly excludes it.
Worth knowing
That paper measures billing, coding and reporting burden broadly. It does not measure the cost of data that cannot describe itself, and we do not claim to recover $265.6 billion. The defensible statement is narrower: this is the largest waste domain in US health care and the only one the literature has no answer for, and a component of it is semantic.
Government-mandated version steps
When a regulator moves an exchange standard forward one version, somebody pays for it. The US government priced exactly this. In the HIPAA 5010 final rule (74 FR 3296, January 2009), with the regulatory impact analysis performed by Gartner under contract to CMS, the total industry cost of the mandated migration was put at $7.1 billion to $14.1 billion over 2009 to 2019.
Two findings inside that rulemaking matter more than the total. First, a single version step was costed at 25 to 50 percent of the original implementation, revised upward after commenters said the government's first estimate was too low and that the real figure was 50 to 75 percent. Second, and this is the one to sit with: testing was 60 percent of the cost for providers and 65 percent for health plans. Hardware and software together were about 20 percent.
Testing is what it costs to re-establish that two parties still mean the same thing. A federal rulemaking priced that at roughly two thirds of the bill. HIPAA 5010 final rule, Table 3
Careful
That analysis projected net benefits: $15.9 billion to $40.9 billion against those costs. Anyone who checks will find the benefit column, so it belongs here rather than in a footnote. These are also ex ante projections in 2009 dollars, not measured outcomes. Nobody ever published what 5010 actually cost.
Retail and supply chain
A supplier can load the right goods on the right truck and have it arrive on time, and still be fined. Walmart's on-time in-full program sets a 98 percent threshold and penalizes at 3 percent of the cost of goods sold. A share of those penalties are not logistics failures at all. They are description failures: the data about the shipment did not match what the receiving system expected it to mean.
Look at what a supplier actually buys to sell into a large retailer. A GS1 identifier is a statement of what a product is. A data pool subscription is, by its own name, a global data synchronization network, paid monthly, whose function is keeping two parties agreeing about that product. Per-retailer onboarding is the cost of building a translation between your meaning and theirs. A chargeback is the fine for getting the meaning wrong. Very little of that spend moves data. Nearly all of it maintains agreement about what the data means, per retailer, forever.
Data and ontology programs
Organizations that decide to fix this properly meet the bill in its most visible form. Michael Atkin, writing for Cutter Consortium in April 2023 after four decades advising financial institutions, puts the long-term cost of a true enterprise knowledge graph at $10 million to $20 million, with a team of five to fifteen people. Writing in Graph Praxis in February 2026, Alexander Shereshevsky puts five-year total cost of ownership at $500,000 to $1.5 million and up per domain ontology, in a lineage that runs back to the ONTOCOM cost model of 2006.
The shape of this was drawn seventeen years ago. PwC's Technology Forecast in spring 2009 published a figure showing the cost of using data sources against the number of sources, with conventional integration climbing steeply and the ontology-driven approach starting higher and staying flat. Neither axis carries a number. The curve has been public, and blank, since 2009.
Why it recurs
The common mechanism is simple enough to state in a sentence. Meaning is established once by people who understand the domain, and then discarded at the boundary, so it has to be re-derived by the next system, the next partner and the next version. Every re-derivation is billable. Testing, mapping, onboarding, reconciliation, chargebacks and the analyst who knows what the column really holds are all the same purchase.
It also gets worse rather than better. Each version step is mapped onto the workarounds the last one left behind. In 2022 HHS reported industry assessments that one pharmacy standard transition would take between two and four times the effort of the previous one.
Putting LLMs on top does not change the arithmetic, it repeats it. A pipeline that infers meaning at run time pays the full cost on the first run and the same full cost on the four thousandth, because nothing it worked out was written down in a form the next run could use. That comparison is measured and published in our substrate economics paper.
The curve that changes it
There is one mechanism that moves this from a recurring cost to an amortizing one: bind the meaning and the constraints into the record when it is authored, so the next system does not re-derive them. A component that has been modeled, constrained, given a permanent identifier and published is not modeled again. It is assembled.
The discount is not our estimate. Atkin measured it in the same paper as the $10 to $20 million: once an organization moves to an extensible platform, incremental use cases run at 30 percent of the original cost, and three times faster.
Plan for 30% of the original cost, but three times faster. Michael Atkin, Cutter Consortium, April 2023
Here is the part that is ours, and it is the whole argument. Atkin's curve is reuse inside one organization. Having paid to build your own ontology and your own pipelines, your second internal use case is cheaper. That is real, and it is also why the first one is so hard to fund: the discount arrives after the invoice.
SDC's components are published to an open catalog under an open license. The reuse is across organizations, which means the discount can apply to the first engagement in a sector rather than only to the second one inside a single client. When a supplier models a product catalog, the next supplier in that vertical starts from published components. When a health system models a care record, so does the next one. The catalog is the asset, and it is open on purpose, because a catalog only compounds if the people who need it can use it without asking us.
What we are not claiming
- We move cost from recurring to up front. Binding meaning at authoring time is design-time work by someone who knows the domain. That is a real cost and it lands before the benefit does.
- We do not remove your existing standards. A retail supplier still needs GS1 identifiers and still publishes to the data pool. HIPAA-covered transactions are still X12. SDC sits underneath those, it does not replace them.
- None of the figures above is a savings model. They describe the size and the shape of a problem from published sources. What any organization recovers depends on its own numbers, which is what a pilot is for.
- Deterministic validation does not prevent physical failures. A late truck is a late truck. It prevents the shipment being described wrongly, not shipped wrongly.
What to put in the RFP
The useful version of everything above is one line an executive can hand to a procurement team, because it is checkable before signature rather than argued after it:
Hand us the constraint set that makes the data operable, in a form we can validate without you, and let us test it before signature. The test that separates a substrate from a subscription
A vendor who can answer that has given you something that keeps working if the relationship ends. A vendor who cannot has told you that the meaning of your data lives in their platform. Both answers are worth knowing before the money moves, and it costs nothing to ask.
Three questions in the same family, for the same reason: can we validate a record without calling your service, does the model survive your next release, and can a party who has never met us verify the result on their own hardware.