Statistics are collected from web sources, extracted by LLM or deterministic parsers, and then pass through evidence grading before publication. The pipeline is automated; grades are assigned by a deterministic formula, not editorial judgment.
Automated crawlers collect marketing statistics from web sources including industry reports, survey publishers, government data, and common crawl indices. Sources are tracked by acquisition channel. Not every source is authoritative -- source trust is scored in the grading step.
Web-scale crawlLLM extraction identifies numeric claims, their context, units, and entities from source content. Deterministic extractors handle structured sources (EDGAR, FRED, BLS). Every extraction records the model version, prompt, and raw response for reproducibility.
LLM + deterministicDeterministic validation checks value plausibility: currency magnitude sanity (detecting lost billion/million suffixes), percentage bounds, count positivity, and unit-metric consistency. Stats failing hard validation rules are blocked from publication.
Deterministic rulesA multi-dimensional scoring formula evaluates evidence quality across five core dimensions: source trust, citation chain certainty, independence, methodology quality, and freshness. Sample size, contradiction detection, and topic relevance modify the score. Output: an A-D evidence grade.
5-dimension formulaWhere multiple sources report the same statistic, corroboration weight accumulates. Cross-source agreement increases confidence; contradiction reduces it. This layer is being expanded -- some stats currently lack multi-source corroboration, which is reflected in their grade.
Multi-source matchingOnly statistics passing all quality gates are published. Each stat carries a citation safety classification (Safe to Cite, Use with Hedging, Do Not Cite) and a provenance chain linking back to its source artifact and extraction event.
Gated publicationEvery published statistic is assigned one of four evidence grades based on a deterministic rubric -- never editorial opinion. Grades reflect the quality, verifiability, and corroboration of the underlying research.
Methodology-documented primary or institutional data with a clear provenance trail, adequate sample, and recent vintage. Highest-confidence tier; review the source context before formal citation.
Methodology-documented industry or commercial observation with corroboration and an adequate sample. Reliable for directional use with the stated source context.
Single credible source or cross-source consensus without a primary study. Auto-generated hedging language provided. Suitable for directional use with appropriate caveats.
Unverified origin, insufficient corroboration. Flagged as Do Not Cite. Included in the database for completeness and research context only.
Every grade is computed across exactly 6 weighted dimensions. The weights are defined by formula and applied consistently across all statistics. LLMs are never the source of a numeric score.
Bar widths are illustrative of relative weighting. Exact dimension weights are published in the technical appendix.
Every published statistic receives a citation safety classification that tells you exactly how and when it is safe to reference the data -- with no ambiguity.
Cite with full confidence. Complete provenance chain is available. Suitable for published reports, enterprise presentations, and journalism.
Reliable for directional use. Note the source limitation when citing. Recommended for internal research and supporting evidence.
Auto-generated disclaimer language is provided with every statistic. Use as supporting context, not as a primary claim.
Insufficient provenance to support citation. Available for internal research purposes only. Not suitable for any public-facing use.
Every published statistic carries a provenance chain linking it to its source. The chain is append-only -- records can be superseded by new versions but never deleted.
The original source content -- content-hashed and stored. URL, retrieval timestamp, and raw content are recorded at ingestion. Source artifacts may be raw HTML, PDFs, or structured data.
Records which extraction method was used (LLM or deterministic), the model version, prompt version, raw model response, and structured output. Both LLM and deterministic extractions are fully traceable.
Every state change creates an immutable version record in stat_versions. No data is overwritten. Previous versions remain accessible. Demotions and re-approvals are tracked.
The grading trace records which scoring steps passed or failed, the specific values that triggered any failure, and the evidence grade assigned. This is stored in grading_trace_steps.
Upstream sources this statistic corroborates, and downstream derivatives that depend on it. When a source stat is demoted, dependent derivatives are invalidated through the demotion cascade.
A completeness score reflecting how fully the provenance chain is populated. This is used to prioritise quality assurance review and identify stats needing additional corroboration.
Confidence scores are not static. Lighthouse applies a time-aware staleness model that automatically downgrades confidence as data ages -- with topic-specific thresholds for breaking news versus evergreen benchmarks.
Statistics that have not been refreshed by a new primary source within 90 days receive an automatic confidence penalty. The grade may drop by one level if no corroborating update is found. A staleness flag is added to the citation display.
Statistics older than 180 days without a refreshed source are blocked from new citations and marked as potentially stale. They remain in the database with full provenance history but are excluded from top-of-page results.
Structural benchmarks -- such as average human reading speed or long-run global literacy rates -- are classified as evergreen and exempt from standard staleness thresholds. Evergreen classification is assigned by formula, not editorial discretion.
We believe in being honest about what the system does well and where it falls short.
Browse evidence-graded statistics -- each one with a provenance chain, citation safety classification, and confidence breakdown available on demand.