Method: The government atlas — capturing and keeping the public record
The Government Atlas is only as trustworthy as the documents behind it. This is how we hold every report, bylaw, and budget the atlas cites — not just a link that can rot or a claim that can drift, but a kept copy anyone can independently re-check, plus a standing check that we kept the right copy in the first place.
What it produces
A durable archive behind every document the Government Atlas cites — federal, provincial, city, and a widening set of comparator municipalities. For each one, a record of whether we hold it, whether we can read it, and whether we've confirmed it's actually the complete document and not a summary, an index page, or the wrong year standing in for it.
How it works, step by step
Every document goes through the same three-layer custody chain before it's trusted as "held":
- Proof it existed. When we capture a public document, we take an archived copy in a standard web-archive format at the moment of capture — not after the fact. That archived copy is the evidence the document looked the way we say it looked, on the day we say we saw it.
- The original file itself. The archived copy is unpacked into the actual file — the PDF, the web page — stored under a naming scheme keyed to the file's own content, so duplicates are caught automatically and nothing gets silently overwritten.
- A readable text copy. Every held file is also converted to plain text, so it can be searched and checked without re-opening the original format each time. If a document turns out to be an image scan with no real text underneath, that gets flagged rather than quietly counted as "readable."
Then comes the check that matters most: the right-document check. Holding a file isn't enough if it's the wrong file — a summary or an index page standing in for the full report, or last year's budget filed under this year's name. A first pass uses mechanical signals — duplicate files, mismatched file types, suspiciously short documents, a title that doesn't match its stated year, first pages that read like "executive summary" or "table of contents" instead of the real thing — to sort documents into clean and suspect piles. The suspect pile is then reviewed, and at least one in ten of those review decisions is independently re-checked by a second reviewer before anything gets flagged in the public registry. Nothing is flagged silently: a suspect document gets a visible marker — possible summary standing in for the full document or possible wrong document — right in the row that cites it.
Where a document was captured but the underlying file is missing locally, that's tracked as a standing to-do list, not folded quietly into a "done" count. And where a report cites the law that empowers an agency to act, we hold that law under the same three-layer chain as the agency's own documents — because a claim about what a body is allowed to do is only as good as the statute behind it.
How it's checked
Every layer writes its own record rather than asserting a number in prose: a custody record tracks whether each document was successfully captured and unpacked, and a fidelity record tracks whether each one converts cleanly to readable text. A document only counts as "held" for citation purposes once both records say so. On top of that, the right-document check runs its own independent sample: whenever a first pass sorts a batch of documents into clean and suspect, at least 10% of those calls get re-checked by someone who didn't make the original call, before any flag reaches the public registry. Every flag written to the registry goes through the same controlled process — never a manual edit — so there's a single trail for every "possible wrong document" marker anyone sees.
We also keep a dated, honest list of the failure modes we've actually hit — not a hypothetical list. A retroactive audit of thirty cited claims from an earlier automated pass found zero fabricated citations, which is the finding that pushed us to make "archive at the moment of capture" mandatory rather than optional. A gap in how we extracted text from some documents was found and patched. And the exact failure the right-document check exists to catch — a summary quietly standing in for the full report — is treated as a real, expected class of error the pipeline is built to find, not an edge case we hope never happens.
What it honestly costs
Capturing and unpacking documents, and converting them to text, is mechanical work a computer does — no per-document judgment call required. Sorting the suspect pile from the right-document check is bulk classification work, but the required independent re-check of at least one in ten of those calls takes real reviewer time and cannot be skipped. Any judgment call that changes the registry's own structure — a new government level, a new document category — is logged as a dated, versioned change, not a silent edit.
As one data point: in a single working session, we recovered and matched over 99% of a batch of roughly 2,750 previously-uncertain files, sorted the suspect pile into eight review batches, and captured 160 pieces of underlying legislation the same day. That's one precedent, from one session — not a promise about how fast the next batch will go.
Day to day attention required is low for the mechanical steps. It rises sharply the moment a document gets flagged "possible wrong document," or when the registry's own categories need to widen — those are the moments a person has to look and decide, by design.
Known limits
The "at least one in ten" independent re-check is a working convention, not yet written as a formal rule with a defined next step if the re-checker disagrees with the original sort. When several copies of what looks like the same document turn up, we don't yet have a written rule for which copy becomes the one of record. The path for documents that turn out to be low-quality scans — needing real optical character recognition and, for the worst cases, a human read-through — has a known destination but no written trigger for when that extra step kicks in versus waits. And when a document is confirmed captured but the underlying file can't be found locally, that's tracked on a running list rather than resolved as a one-time sweep — it's ongoing work, not a finished project.
What stays internal, and why
The actual archive — the stored files, their text conversions, and the internal indexes that track them — is working storage, not a published dataset: it isn't part of the public site and isn't linked from it. This page describes the three-layer rule that governs it, not the storage scheme itself. Likewise, in-progress reconciliation work — sorting out duplicate or uncertain entries before they're resolved — is internal working material, not something we publish mid-process. If you're building this in your own municipality, what you need is the three-layer law itself (proof it existed, the file, the readable text, plus the right-document check) — not our internal file-naming or storage details, which are ours to manage and not part of the method you'd copy.
Every entry in the Government Atlas that cites a document links back to this discipline. If a document is flagged as a possible mismatch, you'll see the flag right on the registry row that cites it — we'd rather show you an open question than a false confidence.
Want to run this in your own municipality? Start here: Run this in your municipality. See how we check other kinds of claims across the whole site: How we verify.
— The Unknown Soldier · uniteTOlove · Toronto