Every Statistic Is Earned, Not Assumed

A complete explanation of the systems, algorithms, and editorial standards behind every number in the Lighthouse database.

Last reviewed March 2026 Version 2.4 Independently audited Q3 2026 (scheduled)
4.6M+
Verified statistics
A - D
Evidence-graded
Full
Citation chains
Continuous
Updated continuously
50+
Industries covered

How Statistics Enter the Database

Statistics are collected from web sources, extracted by LLM or deterministic parsers, and then pass through evidence grading before publication. The pipeline is automated; grades are assigned by a deterministic formula, not editorial judgment.

01
Step One
Collect

Automated crawlers collect marketing statistics from web sources including industry reports, survey publishers, government data, and common crawl indices. Sources are tracked by acquisition channel. Not every source is authoritative -- source trust is scored in the grading step.

Web-scale crawl
02
Step Two
Extract

LLM extraction identifies numeric claims, their context, units, and entities from source content. Deterministic extractors handle structured sources (EDGAR, FRED, BLS). Every extraction records the model version, prompt, and raw response for reproducibility.

LLM + deterministic
03
Step Three
Validate

Deterministic validation checks value plausibility: currency magnitude sanity (detecting lost billion/million suffixes), percentage bounds, count positivity, and unit-metric consistency. Stats failing hard validation rules are blocked from publication.

Deterministic rules
04
Step Four
Grade

A multi-dimensional scoring formula evaluates evidence quality across five core dimensions: source trust, citation chain certainty, independence, methodology quality, and freshness. Sample size, contradiction detection, and topic relevance modify the score. Output: an A-D evidence grade.

5-dimension formula
05
Step Five
Corroborate

Where multiple sources report the same statistic, corroboration weight accumulates. Cross-source agreement increases confidence; contradiction reduces it. This layer is being expanded -- some stats currently lack multi-source corroboration, which is reflected in their grade.

Multi-source matching
06
Step Six
Publish

Only statistics passing all quality gates are published. Each stat carries a citation safety classification (Safe to Cite, Use with Hedging, Do Not Cite) and a provenance chain linking back to its source artifact and extraction event.

Gated publication

The A - D Evidence Grade System

Every published statistic is assigned one of four evidence grades based on a deterministic rubric -- never editorial opinion. Grades reflect the quality, verifiability, and corroboration of the underlying research.

A
High Confidence

Methodology-documented primary or institutional data with a clear provenance trail, adequate sample, and recent vintage. Highest-confidence tier; review the source context before formal citation.

B
Verified

Methodology-documented industry or commercial observation with corroboration and an adequate sample. Reliable for directional use with the stated source context.

C
Provisional

Single credible source or cross-source consensus without a primary study. Auto-generated hedging language provided. Suitable for directional use with appropriate caveats.

D
Low Confidence

Unverified origin, insufficient corroboration. Flagged as Do Not Cite. Included in the database for completeness and research context only.


Confidence Score Decomposition

Every grade is computed across exactly 6 weighted dimensions. The weights are defined by formula and applied consistently across all statistics. LLMs are never the source of a numeric score.

Source Credibility
Methodology Rigor
Sample Adequacy
Data Recency
Corroboration Breadth
Independence Scoring

Bar widths are illustrative of relative weighting. Exact dimension weights are published in the technical appendix.


Four Citation Safety Classifications

Every published statistic receives a citation safety classification that tells you exactly how and when it is safe to reference the data -- with no ambiguity.

Safe to Cite
Grade A or B, no risk flags

Cite with full confidence. Complete provenance chain is available. Suitable for published reports, enterprise presentations, and journalism.

Use with Context
Grade C, high confidence score

Reliable for directional use. Note the source limitation when citing. Recommended for internal research and supporting evidence.

Use with Hedging
Grade C, lower confidence score

Auto-generated disclaimer language is provided with every statistic. Use as supporting context, not as a primary claim.

Do Not Cite
Grade D or flagged

Insufficient provenance to support citation. Available for internal research purposes only. Not suitable for any public-facing use.


Provenance Chain

Every published statistic carries a provenance chain linking it to its source. The chain is append-only -- records can be superseded by new versions but never deleted.

1
Layer 1
Source Artifact

The original source content -- content-hashed and stored. URL, retrieval timestamp, and raw content are recorded at ingestion. Source artifacts may be raw HTML, PDFs, or structured data.

2
Layer 2
Extraction Event

Records which extraction method was used (LLM or deterministic), the model version, prompt version, raw model response, and structured output. Both LLM and deterministic extractions are fully traceable.

3
Layer 3
Version History

Every state change creates an immutable version record in stat_versions. No data is overwritten. Previous versions remain accessible. Demotions and re-approvals are tracked.

4
Layer 4
Quality Trace

The grading trace records which scoring steps passed or failed, the specific values that triggered any failure, and the evidence grade assigned. This is stored in grading_trace_steps.

5
Layer 5
Dependency Graph

Upstream sources this statistic corroborates, and downstream derivatives that depend on it. When a source stat is demoted, dependent derivatives are invalidated through the demotion cascade.

6
Layer 6
Lineage Score

A completeness score reflecting how fully the provenance chain is populated. This is used to prioritise quality assurance review and identify stats needing additional corroboration.

10M+
Stats in the database
268K+
Quality-approved statistics
5
Verification dimensions
18+
Months to build this infrastructure

How Data Freshness Works

Confidence scores are not static. Lighthouse applies a time-aware staleness model that automatically downgrades confidence as data ages -- with topic-specific thresholds for breaking news versus evergreen benchmarks.

90
90-Day Threshold -- Confidence Downgrade

Statistics that have not been refreshed by a new primary source within 90 days receive an automatic confidence penalty. The grade may drop by one level if no corroborating update is found. A staleness flag is added to the citation display.

180
180-Day Threshold -- Publication Block

Statistics older than 180 days without a refreshed source are blocked from new citations and marked as potentially stale. They remain in the database with full provenance history but are excluded from top-of-page results.

Topic-Aware Evergreen Exceptions

Structural benchmarks -- such as average human reading speed or long-run global literacy rates -- are classified as evergreen and exempt from standard staleness thresholds. Evergreen classification is assigned by formula, not editorial discretion.


Known Limitations

We believe in being honest about what the system does well and where it falls short.

  • Corroboration is incomplete. Not every statistic has multi-source corroboration yet. Stats without corroboration are graded accordingly and may carry lower confidence scores.
  • Extraction errors occur. LLM extraction can misinterpret units, magnitudes, or context. Deterministic validation catches many of these (e.g., currency magnitude checks), but some slip through. We continuously improve the validation rules.
  • Language coverage is primarily English. Non-English extraction quality is lower. We are expanding multilingual support but most verified stats are currently from English-language sources.
  • Grades reflect evidence quality, not truth. An A-grade stat has strong evidence provenance and source trust, but the underlying research may still be flawed. Always review the source methodology for critical decisions.
  • Recertification is rolling. The approved pool is re-checked on a rolling basis. Stats may be demoted if new evidence contradicts them or if validation rules catch issues that were previously missed.

See the methodology in action

Browse evidence-graded statistics -- each one with a provenance chain, citation safety classification, and confidence breakdown available on demand.