Tesseract Foundational Research: Structured Reporting
How machine-readable are UK company accounts, really?
"Structured data beats PDF for AI-driven reporting" is now received wisdom in regulatory circles. It is also rarely measured. This is a small, open benchmark that asks the empirical version of the question: across the accounts that companies actually file, what fraction expose each core figure as a structured fact a machine can read without guessing?
The direction of travel
Digital, structured reporting is the settled direction of UK financial regulation. Companies House already receives company accounts as Inline XBRL, where every figure is machine-tagged to a UK-GAAP or IFRS taxonomy concept, and its transition to software-only filing under the Economic Crime and Corporate Transparency Act will make full tagging universal. The Financial Reporting Council and others increasingly ask what structured data makes newly possible for analysis, assurance and AI. That question has an empirical premise that usually goes unchecked: how much of a real filing is genuinely structured today?
The method: measure exposure, per concept
We took one day of the open Companies House accounts bulk archive, parsed every Inline XBRL filing, and recorded which taxonomy concepts each one exposes as a tagged fact. The benchmark quantity is concept exposure: the share of filings in which a given accounting concept appears as a structured fact an automated reader could consume with no OCR and no layout heuristics. That is the quantity the "structured beats PDF" claim actually rests on, not that data could be structured, but that in the filed population it already is, at a measurable rate that differs sharply by concept.
Share of the 8,856 parsed filings exposing each core concept as a machine-readable fact.
What the numbers say
The core of the balance sheet is genuinely structured: equity is exposed as a machine-readable fact in 97% of filings, net assets in 87%, creditors and total-assets-less-current-liabilities in the mid-70s. The tail falls away fast, and unevenly: debtors, a concept many analysts would assume is always present, appears as a tagged fact in only 28%. That unevenness is the finding. A system that assumes every concept is uniformly available will silently fail on the ones that are not; a system built on the measured exposure profile will not.
Where this goes
The same parser scales to the full daily and monthly archives for a longitudinal exposure profile, and extends to value-level extraction and cross-taxonomy comparison. It complements our work on machine-readable public data, from the quality of continuous environmental monitoring to open ontologies: in every case the useful question is not whether data could be structured in principle, but how far it is structured in the records that already exist.
Independent, self-initiated open research. Contains public sector information from Companies House licensed under the Open Government Licence v3.0. Not endorsed by Companies House or the FRC.
Explore the benchmark
Concept-exposure table, per-filing data, and the reproducible parser.
