For data scientists

Most published economic datasets are terminal artifacts — you can cite them, but you cannot rebuild them, so you have no way to tell a defensible estimate from a confident one.

You will judge this the way you judge any dataset: not by the headline figure but by whether you could rebuild it. Sources, transformations, assumptions, the places where a judgement call was made and recorded rather than buried. Most published economic data fails that test immediately — it arrives as a finished table, and the pipeline that produced it is not part of the artifact.

FAND treats the assembly as the product. The identity is simple enough to check by hand — what a place owns, what it owes, and the trust infrastructure in between — and the interesting work is entirely in how each term gets estimated from upstream agency series. That work is where the reproduction question actually lives, so it is documented rather than summarised: which agency, which series, which vintage, and what was done when a series did not cover the years needed.

The pedigree scheme matters more than it might sound. Every figure carries a NUSAP signature, so a number is never just a number — it comes with an assessment of how far it can be pushed. National-scale figures are the strongest. Subnational figures are provisional, labelled on their face, and the watermark travels with any export. That labelling is not a disclaimer bolted on at the end; it is what lets you decide which parts of the ledger are safe for your purpose and which are not.

On reproduction specifically, the honest position: the replication path is documented on the methodology page, and a public code repository is not yet the way to get it. If you want to reproduce a result, the fastest route today is to say which one, and we will point you at the sources and the steps directly. Promising a repository URL before it exists would be exactly the failure mode this page is about.

Start with a per-node download from the Where-Tree so you have real values in front of you rather than a description of them. Read the sources directory and NUSAP together — inputs and pedigree are the two halves of the same question. Then read the core-and-satellites essay, which is the architecture argument underneath all of it: what belongs in the core accounts and what belongs beside them is a decision, and it is the decision most worth arguing with.

Start here

Related pages

Talk to us about the pipeline

If you intend to reproduce any part of this, tell us which part and we will tell you plainly what is documented, what is not yet, and where the estimates are weakest. We would rather hand you the real state of it than a package that fails on your machine.

Talk to us about the pipeline →