About

StatOrigin is a free, open database of industry statistics. Every number on the site traces back to the organisation that first published it, with the report, the page, the verbatim sentence, and a working link.

Why this exists

Finding a reliable industry statistic is surprisingly hard. Search for almost any market figure and you will find dozens of pages quoting the same number, none of them saying where it came from. Follow the trail and it usually ends at a content farm, or at a paywalled report nobody can check.

That is bad for everyone. Writers cite numbers they cannot verify, researchers spend days on provenance, and machine-generated answers inherit the errors. The fix is not a better search engine; it is a database where provenance is part of the data rather than an afterthought.

The rules we hold ourselves to

  1. Every number has a checkable origin. If we cannot show where a figure came from, it does not get published.
  2. The numbers are never paywalled. Everything visible on a page is free and crawlable. A hidden number cannot be cited, and being cited is the entire point.
  3. We are not an article site. No listicles, no "10 interesting facts". Structured data, charts, and a source table.
  4. We do not republish report content. Only data points and their provenance, and only the figures a publisher released publicly.
  5. Disagreement is recorded, not smoothed. When sources conflict, both values stay visible.

How it is built

The site is a static-first application running entirely on Cloudflare's edge network, with a SQLite database at its core. Pages are pre-rendered where they can be and served from cache where they cannot, so the whole site stays fast without a server to babysit.

Extraction from source documents is assisted by machine reading, but nothing is published automatically: candidates are scored, checked mechanically against the sentence they came from, and anything uncertain is held for a person. The methodology page documents the process in full, including the confidence bands and what each verification badge means.

Every published record states whether it was checked by a human or by the documented automated rule, and we do not blur the two.

How it is funded

Reading and citing the data is free and always will be. If the project needs revenue later, it will come from things that do not reduce how citable the data is: bulk exports, higher API quotas, and watermark-free embeds.

What we will not do is put a number behind a paywall, gate a chart, or require registration to read a figure. That would defeat the purpose of the site.

Reusing the data

Published under CC BY 4.0: attribution required, commercial use permitted. Each record page generates a ready-made citation in APA, MLA, Chicago and BibTeX.

Attribution should credit the original publisher, and link back to the permanent record so the trail stays intact. There is an open REST API and an OpenAPI 3.1 specification if you want to build on it.

Contact

Corrections Report an error, answered within 48 hours
Data requests Ask for an industry, indicator or region to be added
API Documentation · OpenAPI
Machine readers llms.txt · sitemap

Current state

480 published data points from 47 source organisations. Snapshot generated 15 September 2026.

Coverage grows industry by industry. A hub page ships only when it clears the quality bar of at least 50 verified data points from five or more primary sources; we would rather have a few genuinely complete industries than a hundred thin pages.