The Regenerative Toronto Assembly — Blueprint
Blueprint page source — derived render of the frozen Assembly Blueprint
Derived 2026-08-12 (old-estate fires dispatch, Leg 1) by pandoc (gfm) from
research/assembly/assembly_kb_import/Regenerative_Toronto_Assembly_Blueprint.docx, AFTER the
D-008 scrub recorded in that tree's _IMPORT_LOG.md (2026-08-12 entry: cover line + docx
creator metadata → "The Unknown Soldier"; zero residuals verified inside the zip). This file is
the committed page source for /assembly/blueprint.html so the deploy leg needs no pandoc
dependency. Content is the frozen June-2026 working draft, published as-is (cite-don't-adopt;
the second-person voice addresses the project's founder). Everything above the first --- is
stripped at render by tools/assembly_site_gen.py.
THE REGENERATIVE TORONTO ASSEMBLY
A simulated Citizens’ Assembly built from representative AI agents and Pol.is-style deliberation
A reusable civic-deliberation instrument — Toronto as flagship pilot, designed to scale beyond it
Design & Implementation Blueprint
Prepared for The Unknown Soldier
Working draft — June 2026
Contents
1. Executive summary
This blueprint describes how to build and run a simulated Citizens’ Assembly for Toronto: roughly 300 AI agents, each constructed to be statistically representative of a real Torontonian, deliberating one question — the regenerative, flourishing future we want for the city and our grandchildren, and how to get there. The agents argue, listen, and vote like the members of a real civic lottery would. Their votes are mapped with the same mathematics Pol.is uses to surface where a divided public actually agrees. When they need facts, they can summon a panel of frontier “expert” models for balanced briefings. The system outputs a synthesized vision, a set of bridging priorities that hold across demographic groups, and candidate pathways to get there.
What this is: a rehearsal instrument and hypothesis generator. It lets you design and pressure-test the agenda, the framing, the expert materials and the process for a fraction of the cost and time of a real assembly, and it produces testable predictions about where Torontonians might find common ground. It is engineered to run first for $0 on your own hardware, then for \~$20 on cheap hosted models, then — only if warranted — for \~$10k on frontier models.
What this is not: a substitute for real Torontonians. Simulated deliberation approximates group averages far better than it captures the lived diversity within any group, and it carries documented biases. Its outputs are inputs to a real, human assembly — never a replacement for one, and never a manufactured “mandate.” Every published result must be labelled as simulated. Section 9 sets out the validity limits in detail, and they are not optional reading.
Strategic positioning. Per your steer, this is framed as a movement instrument rather than a campaign vehicle: an open, forkable method for any city to convene its own regenerative conversation, with Toronto as the proof-of-concept. It is built to help you lead the conversation and scale it well beyond Toronto, regardless of who ultimately stands for office in October.
2. Strategic framing & purpose
Your thesis is that humanity’s hardest problems are not, at root, technology or resource problems but coordination, governance, incentive and system-design problems — and that the missing piece is adoption and integration of known solutions in one place. A Citizens’ Assembly is one of the strongest known instruments for exactly that: it takes a representative slice of a population, gives them balanced information and good facilitation, and lets them reason their way to recommendations that carry democratic legitimacy. The OECD has documented nearly 300 such processes and found they consistently produce considered, common-ground outcomes that ordinary polling and politics miss [16].
Running one is slow and expensive, which is precisely why a simulation is valuable. A simulation lets you iterate the hard design choices — what question to ask, what evidence to put in front of people, how to frame contested issues without leading the room — cheaply and repeatedly before a single real participant is recruited. It serves three honest purposes:
-
Rehearsal and de-risking. Test agendas, framings and expert briefings across many runs; find where the process breaks or biases the outcome before spending real money and real people’s time.
-
Hypothesis generation. Produce concrete, testable predictions about where Toronto’s groups might converge — e.g. “across every cluster, renters and owners both endorse X” — which a real assembly or survey can then confirm or refute.
-
A replicable public artifact. An open-source kit other cities can fork. This is how the conversation scales past Toronto, which is the actual goal.
The October 2026 context. The municipal election is the catalytic occasion, not the object. The assembly question is framed around the regenerative future; its outputs can include a citizen-authored set of priorities that any candidate could be invited to endorse, and the model can be used to explore the city’s contested relationship with the Province under Doug Ford — where a coordinated, representative citizenry could press for change. This blueprint keeps campaign mechanics light by design; that strategy belongs in its own thread.
3. What we are borrowing — grounded in real systems
Nothing here is speculative architecture. Every component maps to a working system or a peer-reviewed result. The table summarizes; the notes that follow give the mechanism we actually borrow.
| Precedent | What it is | What we borrow |
| Pol.is + Jigsaw Sensemaking [1–4] | Open-source deliberation platform; opinion-clustering on a vote matrix; LLM summarization layer | The clustering math and bridging metric; grounded LLM summaries that cite source statements |
| DeepMind Habermas Machine [7] | AI mediator that drafts group statements maximizing endorsement (Science, 2024) | The generate → predict-endorsement → rank → critique loop for drafting consensus statements |
| Stanford generative agents [9,10] | Interview-grounded agents replicating 1,052 real Americans; memory/reflection architecture | Agent construction from a life-narrative + an “expert reflection” module; memory stream |
| Silicon sampling [11] | Conditioning an LLM on real demographic backstories to simulate survey responses | The method — with its fidelity tested per model, never assumed |
| OECD / sortition [16–19] | Best practice for real assemblies: stratified civic lottery, balanced learning phase | Stratified selection on census strata + an attitudinal stratum; informants-vs-advocates evidence |
Pol.is and the bridging metric [1–4]
Pol.is, from the Computational Democracy Project, runs entirely on a vote matrix — participants vote agree / disagree / pass on short statements and submit their own. The engine reduces that matrix with PCA to a 2-D opinion map, then clusters it with k-means (sweeping k = 2–5, picking the best by silhouette score). The metric we most want is Group-Informed Consensus: for each statement it estimates each group’s probability of agreement (with a Bayesian +1/+2 smoothing prior) and multiplies those probabilities together. Because it is a product, one dissenting group pulls the score toward zero — by design, this rewards statements that \~70–80% of every group endorses and structurally defeats tyranny of the majority [2]. Google Jigsaw’s Sensemaking tools add a Gemini layer that summarizes agreement and disagreement while citing the underlying statements, so summaries stay grounded [4]. The whole stack is open source and reproducible in Python (Section 6).
Note on “Pol.is 2.0”: that term does not map to a single shipped product. It is an informal label for CompDem’s in-progress LLM-augmented rebuild. The real, citable evidence of AI-augmented Pol.is is the Jigsaw Sensemaking partnership and Talk to the City [4,5] — we cite those rather than a version number.
The Habermas Machine [7]
DeepMind’s mediator drafts a single group statement that maximizes predicted endorsement. Mechanically it is a simulated ranked-choice election over candidate statements: a generative model proposes candidates from each member’s written opinion; a reward model predicts how each individual would rank them; a Condorcet (Schulze) method picks the winner; then a critique round lets members object and the draft is regenerated. Across 5,734 UK participants, people preferred the AI-mediated statement to a human mediator’s 56% to 44%, group agreement rose, and minority positions were preserved, not steamrolled. Crucially the authors flag that it does no fact-checking — garbage in, garbage out — and that their virtual assembly modelled only the deliberation phase, without expert testimony. We borrow the loop and explicitly fix the two gaps with the expert panel (Section 7).
Interview-grounded agents [9,10]
Stanford and Google built agents of 1,052 real Americans by injecting a two-hour interview transcript into each agent’s prompt plus an “expert reflection” module (the model reviews the transcript as an economist, a psychologist and a sociologist and stores \~20 high-level observations). These agents reproduced their real humans’ survey answers at about 85% of the humans’ own test-retest consistency, and — importantly — grounding in a narrative reduced demographic bias versus thin persona prompts. The earlier “Smallville” paper supplies the canonical memory architecture: a memory stream scored by recency, importance and relevance, periodic reflection into higher-level beliefs, and planning. We use both: rich narratives over thin personas, and a memory/reflection loop so agents evolve their views across rounds rather than re-answering from scratch.
4. System architecture overview
The system is a seven-stage pipeline. Stages 1–2 build the room; 3–6 run the meeting; 7 writes it up. Each stage is a discrete module you can run, inspect and swap independently — which is what makes the free engineering run (Phase 0) tractable.
| # | Stage | Function | Core technique |
| 1 | Sampling | Allocate 300 agent “seats” to match Toronto | Stratified lottery + raking to census marginals |
| 2 | Instantiation | Turn each seat into a believable person | Life-narrative + expert-reflection + memory store |
| 3 | Learning | Give agents balanced evidence on request | Frontier “expert” panel: informants vs advocates |
| 4 | Deliberation | Multi-round statement generation & voting | Agent debate; Pol.is comment routing |
| 5 | Mapping | Find the clusters and the bridges | Vote matrix → PCA → k-means → group-informed consensus |
| 6 | Mediation | Draft statements the whole room can accept | Habermas loop: generate → predict → Schulze → critique |
| 7 | Synthesis | Vision, priorities, pathways, report | Grounded LLM summarization (Jigsaw-style) |
Data flow: Census + strata → [1] agent roster → [2] agent objects (profile + memory) → [4] each round, agents read shared context and the routed statement subset, emit a statement and votes → votes accumulate in a participant×statement matrix → [5] clustering identifies groups and bridging statements → [3] unmet information needs trigger expert calls that feed back into the next round → [6] consensus drafting → [7] report. Everything — every vote, every expert call, every prompt — is logged for transparency and replay.
5. Component A — the 300 representative agents
5.1 Selection: a stratified civic lottery, executed by raking
Real assemblies use a two-stage civic lottery: thousands of random invitations, then a stratified weighted draw from the respondents so the final panel matches the population on agreed strata [17,18]. Because we generate agents rather than recruit them, we collapse this to one step: we define target quotas from Toronto’s 2021 census marginals and use iterative proportional fitting (raking) to allocate 300 agents so that every marginal is matched simultaneously. The table below shows illustrative single-dimension targets and the resulting agent counts; the real allocation rakes across all dimensions jointly, because StatCan publishes the marginals freely but the full joint cross-tabulation (e.g. income × race × ward) requires a custom tabulation [20,21].
| Dimension | Target (City of Toronto, 2021) | Agents / 300 |
| Geography (community council areas) | Toronto–East York 30% · North York 25% · Scarborough 23% · Etobicoke–York 22% | 90 / 75 / 69 / 66 |
| Gender | \~52% women · \~48% men · \<1% non-binary / trans | 156 / 142 / \~2 hard-seated |
| Age (adults 18+) | 18–24 \~10% · 25–44 \~38% · 45–64 \~31% · 65+ \~21% | 30 / 114 / 93 / 63 |
| Racialized population | 55.7% racialized · 44.3% not | 167 / 133 |
| — largest groups | South Asian 14% · Chinese 10.7% · Black 9.6% · Filipino 5.8% | 42 / 32 / 29 / 17 |
| Immigrant status | \~47% foreign-born | 141 |
| Tenure | \~48% renter · \~52% owner | 144 / 156 |
| Low income (LIM-AT, 2020) | \~13% (structurally \~18–20%; 2020 was COVID-deflated) | \~45 hard-seated |
| Education | \~41% bachelor’s degree or higher | 123 |
| Indigenous | \~1.7% (First Nations, Métis, Inuit) | \~5 hard-seated, all three identities present |
| Attitudinal stratum | e.g. trust-in-government & openness-to-change, balanced terciles | \~100 / 100 / 100 |
Figure values are illustrative single-dimension marginals; the production roster rakes all dimensions jointly. Income figures for 2020 are distorted by pandemic transfers — flag this when low-income representation matters. Sources: Statistics Canada 2021 Census Profile for Toronto (CSD 3520005) [20] and City of Toronto census backgrounders [21].
The attitudinal stratum matters. Climate Assembly UK deliberately stratified not only on age, gender, ethnicity, education and geography but on attitude to climate change, so the room was not packed only with people who already cared [19]. We adopt the same move: stratify on one or two attitudinal axes (for example trust in government and openness to systemic change) so the simulated assembly is not quietly stacked with enthusiasts for your own vision. This is one of the strongest guards against a flattering, self-confirming result.
5.2 Instantiation: from a row in a table to a person
A demographic row is not a person. Following the Stanford result that narrative grounding both raises fidelity and lowers demographic bias [9], each agent is built from a synthesized life-narrative consistent with its profile — occupation, neighbourhood, household, migration story, daily pressures, what it worries about and hopes for — plus an expert-reflection pass that distills the narrative into stable beliefs and values stored in the agent’s memory. During deliberation the agent retrieves from a memory stream scored by recency, importance and relevance [10], so positions taken in round 2 inform round 4.
Honest gap: Stanford’s agents were grounded in real two-hour interviews of the actual people they simulate. Ours are grounded in synthetic narratives generated from demographics. That is a meaningful fidelity downgrade and the single largest validity risk in the whole design (Section 9). In the $10k phase it is partly closed by grounding narratives in real qualitative material — public Toronto consultation transcripts, ethnographic and oral-history archives — rather than inventing them whole.
5.3 Varying capacities and limitations
Real rooms are not uniform. Some members are articulate and numerate; others are quiet, uncertain, distrustful, or short on time. We give each agent profile-linked parameters: verbosity, baseline knowledge, numeracy, attention, civic engagement, persuadability, and trust in institutions. Some agents abstain often; some over-participate; some misunderstand a briefing and need it re-explained. This heterogeneity is not cosmetic — it is what keeps the simulation from collapsing into an unrealistically articulate, agreeable consensus, and it is what makes the expert-request mechanism necessary: an agent that flags “I don’t follow this” or “I want evidence on that” is what routes a question to the panel.
6. Component B — the deliberation engine & Pol.is integration
Each deliberation round has the agents (a) read a short shared context and a routed subset of existing statements, (b) vote agree / disagree / pass on them, and (c) optionally author a new statement or an information request. Pol.is’s comment-routing logic decides which statements each agent sees next — weighting by information gain, consensus potential and novelty — so no agent has to read all of them and the system scales [1]. Votes accumulate into the participant×statement matrix that drives the mapping.
The clustering, in practice. After each round the vote matrix is reduced with PCA to two dimensions and clustered with k-means (k = 2–5, best by silhouette). For each candidate statement we compute group-informed consensus — the product of per-group agreement probabilities — and representativeness for each group [1,2]. This yields two outputs every round: the opinion map (how many camps, and what divides them) and the bridge list (statements that hold across all camps). The bridge list is the heart of the deliverable: it is where a divided Toronto actually agrees.
6.1 Self-hosting the math
Cheap / fast path — red-dwarf. The Python library red-dwarf reproduces Pol.is’s pipeline exactly using stock scikit-learn (PCA → KMeans → silhouette, plus the representativeness statistics) and can even load real Pol.is conversations for calibration [26]. For Phases 0 and 1 this is all you need — no server, just a function over the vote matrix.
High-fidelity / real path — full Pol.is. The full Pol.is platform is open source (AGPL-3.0) and self-hostable via Docker [3]. Stand it up when you want the real participation and reporting UI — in particular when you move to a hybrid run that mixes simulated agents with real Torontonians, or when you want the Jigsaw Sensemaking summaries on top [4].
6.2 The deliberation arc
The round structure mirrors a real assembly’s learning → deliberation → decision arc [16,19]: (1) framing and seed statements; (2) a learning phase where agents request and receive balanced expert briefings; (3) debate and statement generation; (4) voting and clustering; (5) Habermas-style consensus drafting on the contested items; (6) a critique-and-amend round; (7) a final vote. Three to five rounds is enough to see convergence in Phase 0–1; the $10k run can afford more rounds and re-runs.
7. Component C — the expert-opinion mechanism
This is the component that fixes the Habermas Machine’s acknowledged weakness — no fact-checking, no evidence phase [7] — and that mirrors what makes real assemblies legitimate: a balanced, contestable body of evidence. Its design follows Climate Assembly UK’s split between informants and advocates, overseen by a balance check [19].
-
Informants. Neutral briefings in response to an agent’s information request (“what would a neighbourhood-cell network cost?”, “what does the evidence say about fourplexes and rents?”). The panel returns a grounded, citation-bearing summary.
-
Advocates. On contested questions, the panel is asked to present the strongest case for each side, explicitly, so agents hear genuine disagreement rather than a single smoothed answer.
-
Balance check. A meta-step (and, in the $10k phase, a human reviewer) audits whether the evidence set is balanced before it reaches the floor — the institutional guard that real assemblies place around who gives evidence.
Implementation. The panel is a small ensemble of frontier models acting as expert witnesses — e.g. Claude Opus, a GPT-5-class model, a Gemini Pro-class model, and DeepSeek. Fanning each request out to several models and reconciling their answers (flagging where they disagree) reduces single-model bias and is itself a useful signal. Two guardrails are mandatory: grounding (every claim cites a source) and contestability (any agent can challenge an expert claim, which triggers a re-query). Every expert call is logged, so the evidence base is fully auditable after the fact.
Cost note. Expert calls hit frontier models and are the main cost driver at the top end (Section 12). Budget them deliberately: cache the shared briefing context, batch where latency allows, and reserve the most expensive models for genuinely contested questions.
8. The deliberation agenda — the regenerative future
The central question: what is the regenerative, flourishing future we want for Toronto and our grandchildren, and how do we get there? Around it sit a set of balanced sub-themes, each seeded with both supporting and opposing statements so the room is not led:
-
Housing & land — supply, affordability, public and cooperative housing, density and tenure
-
Food, green space & permaculture — urban agriculture, the ravine system, regenerative land use
-
Mobility & the 15-minute city — transit, active transport, car dependence
-
Governance & decentralization — neighbourhood-level democracy, participatory budgeting, the “10,000 puroks” idea
-
Economy & post-scarcity — local economies, cooperatives, who accumulates wealth and power
-
Climate resilience & energy — adaptation, retrofits, emergency preparedness
-
Belonging & multiculturalism — Toronto’s diversity as civic capacity; cohesion and trust
-
The city–province relationship — municipal autonomy and democratic pushback under provincial control
-
AI, data & transparency — the city as a leader in open, accountable technology adoption
8.1 A flagship stimulus: the 10,000 puroks idea
One concrete proposal makes an excellent deliberation stimulus because it is vivid, evidence-backed, and genuinely two-sided. A purok is a Filipino sub-neighbourhood cell of roughly 20–50 households with a named volunteer leader, a roster of every resident, recurring peacetime functions, and an emergency activation protocol [22]. The municipality of San Francisco in the Camotes Islands, Cebu, built its disaster response on this unit and won the UN’s 2011 Sasakawa Award for it [23]. The proof came in 2013: when Typhoon Haiyan struck, the purok network pre-emptively evacuated the \~1,000 residents of the offshore islet of Tulang Diyot — whose roughly 500 houses were then destroyed — with zero casualties [24].
Asking the assembly: could Toronto build a network of \~10,000 neighbourhood cells, and should it? The balanced evidence pack includes the strong analogues — Indonesia’s near-universal RT/RW associations and Brazil’s Porto Alegre participatory budgeting, which lifted sanitation coverage in poor districts dramatically [25] — and the cautionary one: Cuba’s block-level CDRs, which deliver real mutual aid but also function as instruments of surveillance and political control. A dense neighbourhood-cell network is dual-use; the assembly’s job is to weigh mutual aid against social control, and the simulation should surface exactly where Toronto’s groups land on that trade-off rather than assume the answer.
9. Validity, bias & ethics — reading the results honestly
This section is load-bearing. A simulated assembly that is presented as if it were real, or whose biases go unstated, would be worse than useless — it would be discrediting. The literature on simulating people with LLMs is clear about the failure modes:
-
Variance collapse / false precision. Even when synthetic group means track reality, synthetic responses are far less varied than real ones — in one study the synthetic standard deviation was roughly half the real one — which manufactures false statistical confidence and exaggerated polarization, and can flip the sign of relationships between variables [12].
-
Caricature and flattening. LLM personas reproduce out-group stereotypes and are systematically less internally diverse than real people — worst for minority and intersectional groups [13]. This directly threatens the parts of Toronto we most need to hear.
-
Reliably mis-modelled groups & default skew. Models poorly represent some groups (e.g. the elderly, certain religious communities) and aligned models carry a measurable WEIRD/liberal skew; prompt-steering only partly corrects it [14].
-
Alignment can hurt fidelity. The same RLHF that makes a model pleasant can make it worse at reproducing a subpopulation’s actual distribution of views [11].
-
The inclusion illusion. Even accurate simulation does not deliver the representation that human deliberation exists to provide [15]. Simulated Torontonians are not Torontonians.
What follows for use. The simulation is good at central tendencies and at generating hypotheses about bridging statements; it is unreliable about within-group diversity, the intensity of views, and minority nuance. So: treat every output as a hypothesis to be tested with real people, not a finding. Run on multiple base models and report where they disagree. Where any real Toronto survey data exists, calibrate against it and report the gap. Never publish a synthetic number as though it were a measured one. And label everything as simulated, every time.
Ethics. Three commitments keep this honest: full transparency that participants are AI; an open method others can inspect and fork; and a hard line that the simulation informs and mobilizes but never substitutes for a real democratic mandate. The endpoint of this work is a real, human assembly — the simulation is how we design and de-risk it, not how we replace it.
10. Three-phase implementation plan
10.1 Phase 0 — free feasibility run (engineering proof)
Goal: prove the whole pipeline runs end-to-end on hardware you already own. Quality is secondary; “the machine runs” is the bar.
-
Compute: your RTX 4070 Ti SUPER (16 GB), serving a quantized 8–14B model (Qwen2.5-14B or Llama-3.1-8B at 4-bit) under vLLM, which uses continuous batching to run \~16–32 agents concurrently rather than one at a time [27,28]. 300 agents run as \~10–20 batched waves — an overnight job, not a week.
-
Clustering: red-dwarf in-process [26]. Orchestration: a few hundred lines of custom Python asyncio (see Section 11). Expert panel: Claude Max headroom plus free OpenRouter models.
-
Method: start at 30–50 agents to exercise every stage and debug cheaply, then scale to 300. Cost: $0 marginal. Deliverable: a complete run and a short report showing each stage produced sane output.
10.2 Phase 1 — \~$20 quality run (hosted, cheap)
Goal: a credible first vision report, plus the ability to iterate and run sensitivity checks across base models.
-
Models: cheap non-thinking models via OpenRouter — Qwen3-235B, Mistral Small 3.2, Llama 3.3 70B, or DeepSeek V3.1 — for the agents; a mid-tier model for the expert panel [29]. Avoid reasoning models here: their hidden reasoning tokens are billed as output and make 300-agent costs unpredictable.
-
Economics: a full 300-agent × 5-round run costs well under $1 on these models (Section 12). $20 therefore buys roughly 20–45 complete runs — enough to vary framings, swap base models, and run ablations. Deliverable: a v1 regenerative-Toronto report with a cross-model sensitivity analysis.
10.3 Phase 2 — \~$10k frontier run (only if warranted)
Goal: a defensible, publishable study — and the design spec for a real, human (or hybrid) assembly.
-
Models: frontier agents (e.g. Claude Sonnet 4.6 with prompt caching and the batch API) and a frontier expert panel (Claude Opus 4.8, a GPT-5-class model, a Gemini Pro-class model) [30].
-
Where the money goes: caching and batching make the 300-agent loop itself surprisingly cheap (\~tens of dollars). The budget is consumed by richer interview-grounded narratives, larger contexts and more rounds, multi-model ensembles, frontier expert calls, calibration against real Toronto data, and human oversight/facilitation. Deliverable: a publishable method + the blueprint for the real thing.
| Phase 0 — free | Phase 1 — \~$20 | Phase 2 — \~$10k | |
| Proves | Pipeline runs end-to-end | Credible results + iteration | Defensible, publishable study |
| Agents on | Local 8–14B (vLLM) | Cheap hosted (OpenRouter) | Frontier (Sonnet-class + cache/batch) |
| Experts on | Claude Max + free tier | Mid-tier hosted | Frontier ensemble (Opus / GPT-5 / Gemini Pro) |
| Agent count | 30 → 300 | 300 | 300+ × ensembles |
| Rounds | 3–5 | 3–5 × many runs | More rounds, more re-runs, calibration |
| Marginal cost | $0 | \<$1 per run; \~20–45 runs | \~$10,000 all-in |
| Timebox | A few days | 1–2 weeks | 4–8 weeks |
11. Recommended tech stack
| Layer | Recommendation | Why |
| Local serving | vLLM | Only mature option with continuous batching + PagedAttention; \~10–19× Ollama under load on one GPU [28] |
| Local models | Qwen2.5/Qwen3-14B or Llama-3.1-8B (4-bit) | 14B at 4-bit is the realistic ceiling for 16 GB; \~37 tok/s on your card [32] |
| Hosted aggregation | OpenRouter | One OpenAI-compatible API over 300+ models; free tier for testing, cheap tier for Phase 1 [29] |
| Frontier (Phase 2) | Anthropic / OpenAI / Google direct | Use prompt caching (\~90% off repeated context) + batch (50% off) to keep the agent loop cheap [30] |
| Orchestration | Custom Python asyncio first; LangGraph if you need an explicit phased graph | 300 near-identical agents need a semaphore-bounded async loop, not a heavy framework [31] |
| Clustering | red-dwarf (Python) → full Pol.is (Docker) later | Reproduces Pol.is math exactly with scikit-learn; upgrade to the platform for real/hybrid runs [3,26] |
| Data | StatCan 2021 census + City of Toronto profiles | Marginals for raking the 300-agent roster [20,21] |
| Storage | SQLite (Phase 0–1) → Postgres | Holds the vote matrix, agent memory and the full expert-call log |
On “OpenClaw” / “MyClaw.” These do not map to a 300-agent orchestrator and are not recommended for this. “OpenClaw” most prominently names an unrelated game-engine project; a same-named AI-assistant gateway exists but is a single message-router, not a multi-agent framework, and its provenance is hard to verify. “MyClaw” and “ClawCloud” are hosting/deployment SaaS, not orchestration. Use a custom asyncio loop (cheapest, most transparent, exact token control) and reach for LangGraph only when you want explicit, checkpointed deliberation phases. Avoid AutoGen-style conversational frameworks at this scale — their per-turn overhead multiplies token cost across 300 agents.
12. Cost model
Working assumption per full run: 300 agents × \~5 rounds ≈ 4M input tokens and 1.5M output tokens, before the expert panel. Two levers dominate the top end: prompt caching (cached reads bill at \~10% of input price — decisive because every agent reads the same shared context each round) and the batch API (50% off, async).
| Model (per 1M tokens, in / out) | Cost of one 300×5 run | Use |
| Qwen3-235B ($0.07 / $0.10) | \~$0.43 | Phase 1 — cheapest credible |
| Mistral Small 3.2 ($0.075 / $0.20) | \~$0.60 | Phase 1 |
| Llama 3.3 70B ($0.10 / $0.32) | \~$0.88 | Phase 1 — strong open model |
| DeepSeek V3.1 ($0.15 / $0.75) | \~$1.73 | Phase 1 — upper cheap tier |
| Claude Sonnet 4.6 ($3 / $15), cached + batched | \~$20–40 (core loop) | Phase 2 agents |
| Frontier expert panel (Opus 4.8 / GPT-5-class / Gemini Pro) | Main Phase 2 driver | Phase 2 evidence |
Illustrative $10k allocation: frontier deliberation core across a 3-model ensemble and many scenario variants (\~$1.5k); frontier expert panel, thousands of grounded briefings (\~$2k); rich interview-grounded narrative generation for 300 agents (\~$1k); calibration, validation and ablations against real survey data (\~$1.5k); human facilitation, comms and cloud compute for full Pol.is hosting (\~$2.5k); contingency and frontier-price-volatility buffer (\~$1.5k). The headline: the crowd is cheap; large uncached contexts and frontier expert calls are what spend the budget.
Pricing is volatile and was current to mid-2026 at the time of writing; re-verify the Anthropic, OpenAI, Google and OpenRouter pricing pages before locking any cost model [29,30].
13. Risks & mitigations
| Risk | Mitigation |
| Validity — variance collapse, caricature, skew (Section 9) | Multi-model ensembles; calibrate to real data; report uncertainty; label simulated; treat outputs as hypotheses |
| VRAM ceiling on 16 GB | Cap local models at \~14B (4-bit); quantize KV cache for context; use vLLM batching; offload heavy runs to hosted models |
| Reputational — “fake mandate” accusation | Radical transparency: open method, open logs, simulated labelling, and a stated commitment to a real assembly as the endpoint |
| Leading the room | Balanced seed statements (pro + con); attitudinal stratification; informants-vs-advocates evidence; independent balance check |
| Political — provincial (Ford) friction | Frame as citizen-led and non-partisan; let the assembly’s bridging statements, not you, carry the message |
| Cost overrun at frontier tier | Cache + batch aggressively; reserve frontier models for contested questions; hard budget caps per run |
| Ethics — inclusion illusion | Simulation informs, never replaces; real human deliberation remains the goal |
14. Roadmap & how it feeds the movement
A pragmatic sequence that keeps the scale-beyond-Toronto goal in front:
-
Days 0–30 — build the free pilot. Stand up Phase 0 end-to-end at small scale, then 300 agents. Publish the method openly and invite others to fork it. The artifact itself is the first act of leading the conversation.
-
Days 30–60 — iterate cheaply. Run Phase 1 across framings and base models; produce a first regenerative-Toronto report with sensitivity analysis; recruit collaborators (the Computational Democracy Project, AI Objectives Institute, the Sortition Foundation, sympathetic academics) to harden the method.
-
Days 60–90 — decide on the $10k run. Only if Phase 1 shows something worth the spend. Use it to produce a defensible study and the design spec for a real assembly.
-
Beyond — the real thing, and the fork. Convene a real (or hybrid simulated-plus-human) assembly; release the open-source kit so other cities run their own; let Toronto’s result be the demonstration that invites the world.
How outputs feed the work. The bridge list becomes a citizen-authored set of priorities any October candidate can be invited to endorse; the opinion maps show where the city is genuinely divided versus where it is quietly united; the purok deliberation yields a concrete, evidence-tested flagship proposal; and the whole method becomes a reusable instrument for the larger conversation you want to lead — in Toronto first, and then well past it.
Appendix A — Glossary
| Term | Meaning |
| Sortition / civic lottery | Selecting an assembly by random draw rather than election, then stratifying to match the population. |
| Raking (IPF) | Iterative proportional fitting — adjusting a sample so it matches several known marginal distributions at once. |
| Group-informed consensus | Pol.is metric: the product of each opinion group’s probability of agreeing with a statement; rewards cross-group agreement. |
| PCA / k-means / silhouette | Dimensionality reduction, clustering, and the score used to pick the number of clusters — the Pol.is opinion-mapping stack. |
| Habermas loop | Generate candidate statements → predict each person’s endorsement → rank (Schulze) → critique → regenerate. |
| Expert reflection | Distilling an agent’s narrative into stable beliefs via simulated domain experts (Stanford agent design). |
| Continuous batching / PagedAttention | vLLM techniques that let one GPU serve many agents concurrently by sharing and paging the KV cache. |
| Silicon sampling | Conditioning an LLM on real demographic backstories to generate matched synthetic survey respondents. |
Appendix B — Sources
[1] Pol.is clustering algorithms (Computational Democracy Project) — https://compdemocracy.org/algorithms/
[2] Group-Informed Consensus (CompDem) — https://compdemocracy.org/group-informed-consensus/
[3] Pol.is source code, AGPL-3.0 (GitHub) — https://github.com/compdemocracy/polis
[4] Jigsaw Sensemaking tools (Gemini + Polis) — https://github.com/Jigsaw-Code/sensemaking-tools
[5] Talk to the City, AI Objectives Institute — https://ai.objectives.institute/talk-to-the-city
[6] vTaiwan case study (CompDem) — https://compdemocracy.org/case-studies/2014-vtaiwan/
[7] Tessler, Bakker et al., 'AI can help humans find common ground' (Habermas Machine), Science 2024 — https://www.science.org/doi/10.1126/science.adq2852
[8] Habermas Machine code (DeepMind) — https://github.com/google-deepmind/habermas_machine
[9] Park et al., 'Generative Agent Simulations of 1,000 People', 2024 — https://arxiv.org/abs/2411.10109
[10] Park et al., 'Generative Agents: Interactive Simulacra of Human Behavior', 2023 — https://arxiv.org/abs/2304.03442
[11] Argyle et al., 'Out of One, Many' (silicon sampling), Political Analysis 2023 — https://arxiv.org/abs/2209.06899
[12] Bisbee et al., 'Synthetic Replacements for Human Survey Data? The Perils of LLMs', Political Analysis 2024 — https://www.cambridge.org/core/journals/political-analysis/article/B92267DC26195C7F36E63EA04A47D2FE
[13] Wang, Morgenstern, Dickerson, 'LLMs ... misportray and flatten identity groups', Nature Mach. Intell. 2025 — https://arxiv.org/abs/2402.01908
[14] Santurkar et al., 'Whose Opinions Do Language Models Reflect?', ICML 2023 — https://arxiv.org/abs/2303.17548
[15] Agnew et al., 'The Illusion of Artificial Inclusion', CHI 2024 — https://arxiv.org/abs/2401.08572
[16] OECD, 'Catching the Deliberative Wave', 2020 — https://www.oecd.org/en/publications/innovative-citizen-participation-and-new-democratic-institutions_339306da-en/full-report.html
[17] Sortition Foundation stratification algorithm (GitHub) — https://github.com/sortitionfoundation/sortition-algorithms
[18] Flanigan et al., 'Fair algorithms for selecting citizens' assemblies', Nature 2021 — https://www.nature.com/articles/s41586-021-03788-6
[19] Climate Assembly UK (stratification incl. attitude) — https://www.climateassembly.uk/about/citizens-assemblies/
[20] Statistics Canada, 2021 Census Profile — Toronto, City (CSD 3520005) — https://www12.statcan.gc.ca/census-recensement/2021/dp-pd/prof/details/page.cfm?Lang=E&DGUIDlist=2021A00053520005
[21] City of Toronto, 2021 Census backgrounders & ward/CCA profiles — https://www.toronto.ca/city-government/data-research-maps/neighbourhoods-communities/ward-profiles/
[22] Purok (definition), Wikipedia — https://en.wikipedia.org/wiki/Purok
[23] San Francisco, Camotes wins 2011 UN Sasakawa Award (Official Gazette PH) — https://www.officialgazette.gov.ph/2011/05/18/camotes-island-municipality-awarded-2011-un-sasakawa-award-for-disaster-risk-reduction/
[24] UNDRR, 'Evacuation saves whole island from Typhoon Haiyan' — https://www.undrr.org/news/evacuation-saves-whole-island-typhoon-haiyan
[25] WRI, Porto Alegre participatory budgeting — https://www.wri.org/research/porto-alegre-participatory-budgeting-and-challenge-sustaining-transformative-change
[26] red-dwarf: Pol.is math reproduced in Python (GitHub) — https://github.com/polis-community/red-dwarf
[27] vLLM documentation — https://docs.vllm.ai/
[28] Red Hat, 'Ollama vs vLLM' benchmarking — https://developers.redhat.com/articles/2025/08/08/ollama-vs-vllm-deep-dive-performance-benchmarking
[29] OpenRouter API limits & pricing — https://openrouter.ai/docs/api/reference/limits
[30] Anthropic / Claude pricing — https://platform.claude.com/docs/en/about-claude/pricing
[31] LangChain, AI agent frameworks overview — https://www.langchain.com/resources/ai-agent-frameworks
[32] LocalScore benchmark — RTX 4070 Ti SUPER — https://www.localscore.ai/accelerator/136