FROZEN EDITION — July 2026. This page is part of the v1 Toronto homelessness research library, a complete, sealed project frozen at its July 2026 figures (92% of its 770 claims individually verified against primary sources). It is a product of its July-2026 campaign context, published as a citable historical artifact. Where our newer research disagrees, a dated margin note points at the living page — the frozen text is never silently rewritten. Why there are two editions.

REPLICATION & EXPANSION PROTOCOL v2.0

How to Build (or Rebuild, for Another Municipality) This Kind of Accountability Library

Regenerative Toronto | June 30, 2026

Document ID: MAIN-PROC-006

Status: Process documentation. Not outreach-facing. Supersedes the "§20 LLM Replication Protocol" in the archived origin document — that version is preserved unchanged in _archive/ORIGIN_Toronto_Homelessness_Reform_v1.0.md for reference; this version incorporates everything learned building the actual Toronto library (80+ documents as of July 2026 and still growing), including mistakes worth not repeating.

Use case: This document is what a human or an LLM should read before either (a) extending this Toronto library further, or (b) replicating the approach for a different municipality or topic. It does not itself execute the 444-municipality expansion — that remains a separate, explicitly-held decision.


PART 1 — WHAT THIS PROTOCOL ACTUALLY PRODUCED (the proof of concept)

Before prescribing method, an honest account of what the method produced, so a replicator can calibrate expectations:

The realistic time cost: This took the equivalent of many full research sessions across weeks, not hours. A replicator should budget accordingly — the value is in the rigor, and rigor is not fast.


PART 2 — THE DOCUMENT ARCHITECTURE (taxonomy that scaled)

A flat pile of research notes does not scale past a handful of documents. This taxonomy, arrived at empirically, did:

Prefix Purpose Audience
HMK-0XX Core evidence briefs — one topic each (population data, policy registry, root causes, ROI, etc.) Internal research base; cited by everything else
DERIV-0XX Campaign-ready derivative assets (fact sheets, candidate trackers, social copy, door scripts) Public-facing, distribution-ready
POLICY-001 The single recommendations document everything else justifies The most important document — keep it singular, not split
DATA-001 The claim ledger (CSV) Verification backbone — every other document should be falsifiable against this
REF_002_Sources Annotated bibliography with full source summaries, not just links Journalist/researcher reference
REVIEW-00X Quality audits, including external-model (Opus) review integration Internal — designated as such (see below) once the project matured
(no folder — see below) Process documents, review history, strategic memos Not for a new collaborator's first read
_archive/, archive/ Superseded versions of any document, kept for provenance Local working-directory scratch only; never uploaded

A note for anyone replicating this on a platform without folder-upload support (Claude Projects, as of this project): don't design around a folder split at all. This project originally used a real _internal/ subfolder to separate process/review documents from public-facing research — a working, sound idea for a local filesystem — but it silently broke the moment the library was uploaded as Project knowledge, since that platform flattens every upload into one directory regardless of local structure. The fix applied here (July 3, 2026) was to track the same distinction by an explicit, maintained list in MAIN_REF_001_Index.md instead of by path. Design for the flat case from the start if you're building this on a platform with the same constraint — a list is more maintenance overhead than a folder, but it's the one that actually survives upload.

Lesson: Don't split documents prematurely, but do split once a document exceeds roughly 400-500 lines or covers more than one clearly distinct question. HMK-001 became an unwieldy 1,450+-line "everything" document because early sections were never spun out — a replicator should split earlier than this project did.


PART 3 — THE RESEARCH ROUND STRUCTURE (token-efficient, cross-validated)

3.1 Search Strategy by Round

Run focused 1-3 query rounds rather than broad open-ended research. This project's effective pattern, generalized:

Round 1 — Scale and budget (2-3 queries): total system budget, per-unit/per-night cost, population count Round 2 — Performance/accountability (2-3 queries): relevant Auditor General or equivalent watchdog reports, performance/outcome data Round 3 — Comparative models (2-3 queries): what works elsewhere, with at least one international and one domestic comparator Round 4 — Root causes (2-3 queries): the upstream systems (income support, corrections, health, housing policy) that feed the problem Round 5 — Workforce/operational (1-2 queries): who delivers the service and under what conditions Round 6 — Legal/rights framework (1-2 queries): relevant constitutional, human-rights, or tribunal context

3.2 Multi-Model Cross-Validation

The single highest-value practice in this project: the same research question, run through many different models in parallel, with disagreements treated as signal, not noise.

3.3 Jina/Reader Integration for Dense Sources

For government PDFs, paywalled content, or dense budget documents: https://r.jina.ai/[TARGET_URL] Use for: municipal budget background files, Auditor General reports, advocacy-organization PDFs. This consistently extracted clean text where direct fetches returned poorly-structured output.

3.4 Source Hierarchy (unchanged from the original protocol — it held up)

  1. Primary government documents (budget notes, Auditor General reports, legislation, tribunal data, official council voting records)
  2. Peer-reviewed academic literature (cite the DOI, always)
  3. Credentialed advocacy/policy organizations (treat findings as real but attribute the organization explicitly)
  4. Established journalism (CBC, major mastheads — corroborate financial figures against primary sources before using as the sole source)
  5. Think-tank sources across the ideological spectrum (use with disclosed context; numbers are often accurate even when framing is partisan)
  6. Anything else (forums, petitions, social media) — never cite directly; use only to identify a lead worth chasing to a primary source

PART 4 — THE VERIFICATION DISCIPLINE (the part that's easy to skip and shouldn't be)

4.1 The Claim Ledger Is Not Optional

Every factual claim that will appear in a public-facing document gets a row: claim text, value, source, URL, verification status (VERIFIED / VERIFY-pending / UNVERIFIED-REMOVED), and a note. This is tedious. It is also the only thing that makes the library defensible against a hostile fact-check, and the only thing that lets you catch regressions (see Part 6).

4.2 The Three-Tier Citation Discipline

Every dollar figure, percentage, or headline statistic that will appear in public copy should be sorted into: - Locked — audited, peer-reviewed, or directly confirmed primary-document figures. Use anywhere, no hedge. - Conditional — real, traceable, but single-model or methodology-dependent. Must carry a "modelled estimate" caveat whenever cited. - Do not cite — formally investigated and found unverifiable, or superseded by a better figure.

This project's REF_005_ROI_Lock_Memo.md is the template — build the equivalent for any replication, not as an afterthought but from the start.

4.3 Red Lines (carried forward unchanged — these held up under external review)

4.4 The External Review Pass (Opus, or any stronger/different model)

Once a derivative document is built, before public deployment, route it through a different model for adversarial review — specifically looking for: (a) true facts presented with vulnerable framing that a hostile reader could exploit even though the underlying fact is correct, (b) claims stated more confidently than their source supports, (c) numbers that are technically accurate but invite a misleading inference when read quickly. This project's Opus review caught exactly this category of problem — not factual errors, but framing exposure — across multiple "true but vulnerable" statistics. This step is not optional for anything going to media or political contacts.


PART 5 — BRANDING AND OPERATIONAL SECURITY


PART 6 — THE SELF-AUDIT CADENCE (what this project learned the hard way)

Run a periodic full-library regression check — even (especially) on figures you believe are already corrected. This project found a stale, previously-corrected figure had silently reappeared in 5 documents, almost certainly from copying content out of an older saved draft during a later merge. The fix:

  1. Periodically grep the entire library for any figure that has ever been corrected, not just check it once
  2. Maintain a running corrections log (this project's TRACK_003_Corrections_Log.md) as the single source of truth for "has this been fixed everywhere," not just "was this fixed somewhere once"
  3. Treat a found regression as informative, not embarrassing — it's evidence the verification system is working, provided it gets caught before publication

PART 7 — IF REPLICATING FOR A NEW MUNICIPALITY

(Process notes only — does not constitute authorization to execute a multi-municipality expansion, which remains a separate decision.)


MAIN_PROC_006_Replication_Protocol.md | Version 2.0 | June 30, 2026 | Regenerative Toronto Supersedes the §20 protocol in _archive/ORIGIN_Toronto_Homelessness_Reform_v1.0.md (preserved unchanged for reference)

Frozen v1 edition (July 2026) · published 2026-08-17 · corrections to the living library are welcome — tell us where we’re wrong.