Independent R&D · Prototype in development

From complex reports to trusted data.

GovStat explores a verifiable pipeline for turning fragmented public statistical reports into structured, traceable, reusable data — without losing their meaning or source evidence.

GovTech · Data Engineering · Canonical-first ETL · SDMX-ready · Evidence-grounded AI

Technical pillars

Designed for the difficult parts of public data.

Real-world statistical publications contain split tables, layered headings, changing schemas and footnotes. GovStat treats these as data engineering problems, not just text extraction tasks.

01 / DOCUMENT STRUCTURE

Table reconstruction first

Reassemble cross-page table fragments using logical IDs and directional relationships before extracting and validating their contents.

02 / SEMANTIC DATA LAYER

Canonical-first modeling

Keep a stable intermediate representation between source documents and downstream formats, with mappings designed for metadata evolution.

03 / TRUST & VERIFICATION

Evidence-grounded answers

Link claims to source tables, pages and fields. Use deterministic calculation and human review rather than asking a model to guess numerical results.

Engineering philosophy

AI where it helps.
Rules where they matter.

GovStat separates probabilistic document interpretation from deterministic data validation. Each stage retains its own boundaries, making failures easier to detect, review and correct.

Research principle: a model-generated answer is not the source of truth. The underlying dataset, metadata, validation checks and source evidence must remain inspectable.

Architecture priorities

  • Preserve logical table structure before text recognition.
  • Use stable identities for fragments, tables and source references.
  • Keep transformations and metadata mappings auditable.
  • Validate numerical operations with deterministic code.
  • Expose missing evidence rather than fabricating an answer.
  • Design for evolving statistical yearbook formats and schemas.
Development direction

Research scope & roadmap

A focused, incremental approach to a hard document-to-data problem. The following describes development goals, not released product capabilities.

Phase 1 · Research / prototypeStatistical yearbook ingestion

Visual table reconstruction, extraction checks and end-to-end provenance for a limited public-data scenario.

Phase 2 · ValidationCanonical model + metadata

Compare mapping approaches, handle schema revisions and evaluate accuracy and processing effort.

Future directionInteroperable data services

SDMX-compatible outputs, evidence-based search and modular agent-assisted workflows.