REPLICATION & EXPANSION PROTOCOL v2.0
How to Build (or Rebuild, for Another Municipality) This Kind of Accountability Library
Regenerative Toronto | June 30, 2026
Document ID: MAIN-PROC-006
Status: Process documentation. Not outreach-facing. Supersedes the "§20 LLM Replication Protocol" in the archived origin document — that version is preserved unchanged in _archive/ORIGIN_Toronto_Homelessness_Reform_v1.0.md for reference; this version incorporates everything learned building the actual Toronto library (80+ documents as of July 2026 and still growing), including mistakes worth not repeating.
Use case: This document is what a human or an LLM should read before either (a) extending this Toronto library further, or (b) replicating the approach for a different municipality or topic. It does not itself execute the 444-municipality expansion — that remains a separate, explicitly-held decision.
PART 1 — WHAT THIS PROTOCOL ACTUALLY PRODUCED (the proof of concept)
Before prescribing method, an honest account of what the method produced, so a replicator can calibrate expectations:
- 58 collaborator-facing documents, 332+ individually sourced claims, ~98% independently verified against primary sources (audited budgets, Auditor General reports, peer-reviewed journals, council voting records) — figures current as of June 30, 2026; check
MAIN_REF_001_Index.mdfor the live count - A working claim ledger (one CSV row per factual claim, with source, URL, verification status) that makes every public-facing statement traceable
- A documented self-correction record — claims found wrong were marked REMOVED rather than quietly deleted, and the reasoning is preserved
- Cross-model research synthesis — the same questions were run through 15-20+ different LLMs in parallel rounds, with disagreements between models surfaced and resolved against primary sources rather than averaged
- A real regression caught and fixed during this very protocol-writing session (a corrected figure had silently reverted to a stale value in 5 documents during an earlier merge) — proof that even a disciplined process needs periodic self-audits, not just one clean pass
The realistic time cost: This took the equivalent of many full research sessions across weeks, not hours. A replicator should budget accordingly — the value is in the rigor, and rigor is not fast.
PART 2 — THE DOCUMENT ARCHITECTURE (taxonomy that scaled)
A flat pile of research notes does not scale past a handful of documents. This taxonomy, arrived at empirically, did:
| Prefix | Purpose | Audience |
|---|---|---|
HMK-0XX |
Core evidence briefs — one topic each (population data, policy registry, root causes, ROI, etc.) | Internal research base; cited by everything else |
DERIV-0XX |
Campaign-ready derivative assets (fact sheets, candidate trackers, social copy, door scripts) | Public-facing, distribution-ready |
POLICY-001 |
The single recommendations document everything else justifies | The most important document — keep it singular, not split |
DATA-001 |
The claim ledger (CSV) | Verification backbone — every other document should be falsifiable against this |
REF_002_Sources |
Annotated bibliography with full source summaries, not just links | Journalist/researcher reference |
REVIEW-00X |
Quality audits, including external-model (Opus) review integration | Internal — designated as such (see below) once the project matured |
| (no folder — see below) | Process documents, review history, strategic memos | Not for a new collaborator's first read |
_archive/, archive/ |
Superseded versions of any document, kept for provenance | Local working-directory scratch only; never uploaded |
A note for anyone replicating this on a platform without folder-upload support (Claude Projects, as of this project): don't design around a folder split at all. This project originally used a real _internal/ subfolder to separate process/review documents from public-facing research — a working, sound idea for a local filesystem — but it silently broke the moment the library was uploaded as Project knowledge, since that platform flattens every upload into one directory regardless of local structure. The fix applied here (July 3, 2026) was to track the same distinction by an explicit, maintained list in MAIN_REF_001_Index.md instead of by path. Design for the flat case from the start if you're building this on a platform with the same constraint — a list is more maintenance overhead than a folder, but it's the one that actually survives upload.
Lesson: Don't split documents prematurely, but do split once a document exceeds roughly 400-500 lines or covers more than one clearly distinct question. HMK-001 became an unwieldy 1,450+-line "everything" document because early sections were never spun out — a replicator should split earlier than this project did.
PART 3 — THE RESEARCH ROUND STRUCTURE (token-efficient, cross-validated)
3.1 Search Strategy by Round
Run focused 1-3 query rounds rather than broad open-ended research. This project's effective pattern, generalized:
Round 1 — Scale and budget (2-3 queries): total system budget, per-unit/per-night cost, population count Round 2 — Performance/accountability (2-3 queries): relevant Auditor General or equivalent watchdog reports, performance/outcome data Round 3 — Comparative models (2-3 queries): what works elsewhere, with at least one international and one domestic comparator Round 4 — Root causes (2-3 queries): the upstream systems (income support, corrections, health, housing policy) that feed the problem Round 5 — Workforce/operational (1-2 queries): who delivers the service and under what conditions Round 6 — Legal/rights framework (1-2 queries): relevant constitutional, human-rights, or tribunal context
3.2 Multi-Model Cross-Validation
The single highest-value practice in this project: the same research question, run through many different models in parallel, with disagreements treated as signal, not noise.
- When models agree on a figure, confidence rises sharply — but still verify at least once against a primary source before calling it "Locked."
- When models disagree, that's the most important thing to investigate, not average away. In this project, exactly one model (out of ~20) fabricated a council vote, a budget item, and a council-approval date that did not exist — convincingly enough that it would have entered the library undetected without cross-model comparison.
- Track which models proved reliable vs. unreliable across rounds — this project found certain models consistently produced single-source, unverifiable claims, while others consistently grounded claims in checkable primary documents. Weight accordingly in future rounds, but never skip verification even for "reliable" models.
3.3 Jina/Reader Integration for Dense Sources
For government PDFs, paywalled content, or dense budget documents:
https://r.jina.ai/[TARGET_URL]
Use for: municipal budget background files, Auditor General reports, advocacy-organization PDFs. This consistently extracted clean text where direct fetches returned poorly-structured output.
3.4 Source Hierarchy (unchanged from the original protocol — it held up)
- Primary government documents (budget notes, Auditor General reports, legislation, tribunal data, official council voting records)
- Peer-reviewed academic literature (cite the DOI, always)
- Credentialed advocacy/policy organizations (treat findings as real but attribute the organization explicitly)
- Established journalism (CBC, major mastheads — corroborate financial figures against primary sources before using as the sole source)
- Think-tank sources across the ideological spectrum (use with disclosed context; numbers are often accurate even when framing is partisan)
- Anything else (forums, petitions, social media) — never cite directly; use only to identify a lead worth chasing to a primary source
PART 4 — THE VERIFICATION DISCIPLINE (the part that's easy to skip and shouldn't be)
4.1 The Claim Ledger Is Not Optional
Every factual claim that will appear in a public-facing document gets a row: claim text, value, source, URL, verification status (VERIFIED / VERIFY-pending / UNVERIFIED-REMOVED), and a note. This is tedious. It is also the only thing that makes the library defensible against a hostile fact-check, and the only thing that lets you catch regressions (see Part 6).
4.2 The Three-Tier Citation Discipline
Every dollar figure, percentage, or headline statistic that will appear in public copy should be sorted into: - Locked — audited, peer-reviewed, or directly confirmed primary-document figures. Use anywhere, no hedge. - Conditional — real, traceable, but single-model or methodology-dependent. Must carry a "modelled estimate" caveat whenever cited. - Do not cite — formally investigated and found unverifiable, or superseded by a better figure.
This project's REF_005_ROI_Lock_Memo.md is the template — build the equivalent for any replication, not as an afterthought but from the start.
4.3 Red Lines (carried forward unchanged — these held up under external review)
- Every factual claim requires a source citation in the ledger
- Criticism of any government must be grounded in documented policy action, not characterization of intent
- Distinguish documented actions ("X policy was frozen since date Y") from inferences about motive
- No single unsourced statistic in any derivative campaign document
- Dollar figures from advocacy organizations get cross-checked against primary budget documents before use
- When a claim sounds too clean or too damning to be true, that is the signal to verify hardest, not to deploy fastest
4.4 The External Review Pass (Opus, or any stronger/different model)
Once a derivative document is built, before public deployment, route it through a different model for adversarial review — specifically looking for: (a) true facts presented with vulnerable framing that a hostile reader could exploit even though the underlying fact is correct, (b) claims stated more confidently than their source supports, (c) numbers that are technically accurate but invite a misleading inference when read quickly. This project's Opus review caught exactly this category of problem — not factual errors, but framing exposure — across multiple "true but vulnerable" statistics. This step is not optional for anything going to media or political contacts.
PART 5 — BRANDING AND OPERATIONAL SECURITY
- Decide the public-facing identity (campaign name, author persona, domain) before writing derivative assets, not after — retrofitting a rebrand across dozens of documents (as this project had to do once, when it was smaller than it is now) is mechanical but error-prone, and stragglers hide in headers, social-post copy, and door scripts longer than expected.
- Run a full-library grep for old branding after any rebrand decision, including in volunteer-facing scripts specifically — those are the highest-consequence place for a leftover old name to surface, since a volunteer will say it out loud at a door.
- Never reference the real identity of any individual organizer in public-facing documents — internal/conversational references are fine; document content is not.
- Maintain the internal-vs-public-root split from the start of a project, not retrofitted later — whether that's a real folder (if your platform supports uploading one) or an explicit tracked list (if it doesn't, per the note above), it's much easier to default new process documents into "internal" from the start than to sort 10 of them out after the fact.
PART 6 — THE SELF-AUDIT CADENCE (what this project learned the hard way)
Run a periodic full-library regression check — even (especially) on figures you believe are already corrected. This project found a stale, previously-corrected figure had silently reappeared in 5 documents, almost certainly from copying content out of an older saved draft during a later merge. The fix:
- Periodically grep the entire library for any figure that has ever been corrected, not just check it once
- Maintain a running corrections log (this project's
TRACK_003_Corrections_Log.md) as the single source of truth for "has this been fixed everywhere," not just "was this fixed somewhere once" - Treat a found regression as informative, not embarrassing — it's evidence the verification system is working, provided it gets caught before publication
PART 7 — IF REPLICATING FOR A NEW MUNICIPALITY
(Process notes only — does not constitute authorization to execute a multi-municipality expansion, which remains a separate decision.)
- Start with the same six research rounds (Part 3.1), substituting the new municipality's specific budget/audit/legislative bodies
- Identify the local equivalent of: the municipal auditor, the shelter/homelessness system administrator, the provincial/state oversight body, and any local academic research partner (in Toronto's case, CAMH played this role) — these four sources carry the most weight
- Before treating a small municipality as "no work needed": confirm actual population and homelessness scale first. A municipality with fewer than approximately 5 documented homeless individuals may genuinely warrant being grouped into a single shared regional briefing rather than a full standalone library — but this should be confirmed against an actual count for that specific place, not assumed from population size alone, since under-counting is itself a documented methodological problem (see Fleet Task A in
MAIN_TASK_001_LLM_Fleet_Tasks.md). - Reuse this entire Part 1-6 structure unchanged — the document taxonomy, verification discipline, and self-audit cadence are not Toronto-specific.
MAIN_PROC_006_Replication_Protocol.md | Version 2.0 | June 30, 2026 | Regenerative Toronto
Supersedes the §20 protocol in _archive/ORIGIN_Toronto_Homelessness_Reform_v1.0.md (preserved unchanged for reference)