Methodology

Everything here is designed around one rule: a number is only as good as the ability to check it. This page explains exactly how a figure gets from a published report onto this site, and where the limits of that process are.

1. Sourcing

We only take figures from the organisation that first published them. Where a number appears in a news article, a blog post, or another aggregator, we go to the underlying report and take it from there. If we cannot reach the original, the figure is either excluded or recorded as secondhand and labelled as such everywhere it appears.

Sources are tiered, and the tier affects how much review a figure receives:

  1. Tier 1: international and macro. Multilateral bodies and statistical agencies: World Bank, IMF, OECD, UN, IEA.
  2. Tier 2: national and official. National statistical offices, government departments, regulators, listed-company filings.
  3. Tier 3: research. Independent institutes, academic groups and market analysts, using the figures in their public summaries.
  4. Tier 4: community submitted. Always held for review and labelled secondhand until traced to a primary publisher.

2. Extraction

Each figure is recorded as a single row: one indicator, one region, one period, one value, one source. Alongside the value we store the unit, the scale, whether it is an actual, an estimate or a forecast, and (critically) the verbatim sentence the number was taken from, plus the page or table it appears in.

The quote is not decoration. It is what lets you check the extraction without re-reading the report, and it is shown on every record page.

3. Verification

Candidates are scored for confidence on the clarity of the extraction: whether the value, unit and period are unambiguous in the source, whether the link resolves, and how authoritative the publisher is. That score decides how much human attention a record gets:

Confidence Handling
≥ 0.90 Eligible for publication. A sample is still reviewed by hand, and every record must clear the automated check below.
0.70–0.89 Held for individual human review before it can be published.
< 0.70 Rejected. Not queued, not published, reason recorded.

Before publication, every record must also pass a mechanical check: the extracted value must actually appear in the quoted sentence. This catches the failure mode that matters most (a number that is not in the source at all), and it is why some records carry the label “Verified · automated check” rather than simply “Verified”.

What the badges mean. “Verified” means the record passed the quote check and either a human or the documented automated rule reviewed it. “Awaiting review” means it is in the database but deliberately not published. We do not publish unreviewed figures to make a page look fuller.

4. When sources disagree

Different organisations routinely publish different values for the same thing: different definitions, different coverage, different revision dates. We do not average them, and we do not silently pick one. Both records are kept, and each page shows the disagreement with a link to the alternative figure.

A worked example from this dataset: for China's 2025 new-energy-vehicle output, the National Bureau of Statistics reports 16.524 million units and the China Association of Automobile Manufacturers reports 16.626 million. Both are recorded, both are sourced, and the difference is visible rather than resolved.

5. Corrections

Every page carries a report link. We commit to responding within 48 hours, and corrections are logged publicly, including what was wrong and when it changed. Data is never deleted: a superseded figure keeps its URL and is marked as superseded, so citations do not break.

See the corrections log and policy →

6. What we do not do

  • We do not republish report content. We store data points and their provenance. Report snapshots are kept internally for verification only and are never redistributed.
  • We do not quote paywalled material. Where a figure sits behind a paywall, we use only the number the publisher released publicly in its summary.
  • We do not write articles. This is a structured database. There is no narrative padding on any page.
  • We do not hide numbers behind a paywall. Everything visible on a page is free and crawlable. If we ever charge, it will be for bulk export and higher API quotas, never for seeing the figure.

7. Licensing and reuse

Data on this site is published under CC BY 4.0. You may reuse it, including commercially, provided you attribute it. The citation tools on each record generate the attribution for you.

Attribution should name the original publisher (they did the work) and link to the permanent record on statorigin.org so the chain stays checkable.

8. Current dataset state

Published figures in this snapshot, by how they were checked:

Check applied Records
Automated check (quote contains value) 168
Editorial review 312
Total published 480

Snapshot generated 15 September 2026 · 47 source organisations · last data change 15 September 2026.