Method: The platform substance baseline
Once a registry exists of who is running and what they've published, a second question follows: has this candidate actually offered voters something to read — a real, specific plan — or just a page of slogans? This method answers exactly that one question, and only that one. It never asks whether a candidate's direction is agreeable. The only finding this instrument can produce is that published substance is thin or absent; it cannot and does not produce a finding that a position itself is wrong.
What it produces
A scored baseline of how much substantive, specific, actionable written policy each candidate has published, rolled up into aggregate bands — never a public ranking of named individuals against each other. From that baseline, a small, separately-gated shortlist of candidates whose published platforms are confirmed, through multiple independent checks, to be genuinely exceptional gets named publicly, in praise only. Everyone else's individual standing stays unpublished by design; the public surface is the method, the aggregate picture, and that praise-only shortlist.
How it works, step by step
Capture. A candidate's website and any linked platform or policy pages are saved into a dated file, with the source link and the date it was captured recorded on its face. If a page can't be captured — for instance a site that only loads through interactive scripts a plain fetch can't run — that is recorded as its own honest category, never quietly treated as a zero.
Fit checks before scoring. Before any score is assigned, three checks run. Does the captured material actually belong to this candidate, in this municipality? Is what's left, after stripping login screens, boilerplate, and site-template debris, actually substantial enough to score — a bare minimum length applies below which nothing is scored at all? And is the content current, or is it a leftover page from a past campaign or a different office, in which case it is held and never scored as this candidate's current offer?
Classification before scoring. Every candidate is placed into exactly one category first: a scoreable current platform; a brochure-style page too thin to score in depth; a page that is really a record of past accomplishments rather than a forward plan (reported only in aggregate, never averaged in as if it were a scored platform); stale or out-of-date content; too technically unreachable to assess; or no published platform found at all.
A concrete-commitment check. A candidate can only be fully scored if a reader can point to at least one specific forward commitment in their own words — a pledged action tied to a number, a mechanism, or a timeline. If no such commitment can be found and quoted, the candidate is not scored as a full platform, however polished the page looks.
Dimension scoring. Mechanical counts — word count, number of distinct policy planks, whether costs or timelines or supporting evidence are mentioned, how many issue areas are covered — are computed automatically and are purely descriptive, never a judgment call on their own. On top of that, trained readers score each platform against a fixed set of written anchors across eight dimensions, on a 0–5 scale. Every score of 3 or higher must be backed by a quote from the candidate's own text; every score of 1 or lower must state plainly what was missing. Language that is all hedges and maybes — "could," "might," "would consider" — caps how high a score can go, regardless of how many topics it touches, and only a candidate's forward-looking offer is scored, never their past record.
Rolling up to a headline. Individual dimension scores combine into breadth (how many issue areas are covered), depth, quality, and an overall composite score, using fixed published thresholds to sort candidates into bands. This rollup step is fully automatic and deterministic once the dimension scores exist — no further judgment is applied at that stage, and only rollups covering at least three people are ever reported, so no aggregate can ever identify one or two individuals by elimination.
Naming the exceptional, not the rest. Only candidates in the top scoring band are even considered for public, named praise, and even then only after a second, independent reader re-reads their entire platform from scratch to confirm it really is a current, forward program, and a specific quoted commitment is checked to actually appear, word for word, in the source material. Individual scores themselves are never published — only the confirmed shortlist and the aggregate bands are.
How it's checked
This is the part we'd expect a skeptical outsider to press hardest on, so here is exactly how it's checked, in plain terms.
- Independent double-scoring. A separate, blind sample of platforms is scored a second time by a different reader who never sees the first reader's scores, specifically to measure how consistent the judgment calls really are before any aggregate is allowed to finalize. The full method for that check is published separately: see Method: Blind double-scoring.
- Every real disagreement gets a named, human resolution. Any case where two readers reached a different category call, or scored a platform meaningfully differently, is individually re-read by a senior reviewer against the original material before that score counts toward anything public — never silently averaged away.
- Quotes are checked to actually exist. Before any commitment quote appears on the public exceptional shortlist, it is checked to appear verbatim in the source material — the last of several independent gates a platform must clear before it's named at all.
- The aggregation step can't identify individuals. The rule that no aggregate is published for a group of two or fewer people is built directly into the automated rollup step, not left to a person to remember each time.
- The instrument's own history is on the record. When earlier scoring runs made mistakes — crediting a record of past work as if it were a forward plan, or letting vague, hedged language score as if it were specific — those failures are written up and the scoring rules were changed specifically to catch them earlier next time, not covered up.
The single most important thing to understand about this instrument, and the thing that makes it fair to challenge us on: no named candidate is ever ranked or criticized by it. Any adverse finding — thin, absent, stale, or unclear substance — is published in aggregate only. A candidate's individual score is never released, in either direction, except through the praise-only exceptional shortlist and its own extra gates.
What it honestly costs
The mechanical counts run with no real judgment involved at all — word counts and plank counts are computed automatically. First-pass sorting of ambiguous cases can run on lighter tools too. But the dimension scoring itself, the resolution of any disagreement, and any decision to change the instrument's own rules all require a trained human reader applying real judgment against written anchors — that part does not shortcut. In one real run covering an entire province's worth of candidates, roughly a dozen and a half primary readers scored the full field, with five more readers running the independent blind check described above over a large stratified sample. From dozens of candidates in the top scoring band, an independent multi-gate confirmation process narrowed the list down to fewer than two dozen who were named publicly in praise. That is one real session's numbers, not a promised throughput — treat it as an order of magnitude, not a rate.
Known limits
Published on purpose, because the honesty is the point.
- The rule for exactly what counts as a big-enough change to the scoring instrument to require a version bump — versus a small clarifying edit — is stated as a principle but not yet written down as a hard, specific threshold.
- The written scoring rules themselves are carefully version-tracked, but we can't yet confirm that the exact wording sent to every scoring reader is snapshotted and preserved alongside each version the same rigorous way.
- When a large region is scored by different groups of readers at different times, we check consistency within each group's own work, but we have not yet drawn a sample to check whether one group scored systematically higher or lower than another. A fix for this — a modest cross-group calibration sample — is designed and planned but not yet run as of this writing.
What stays internal, and why
What's public: the method itself, the mechanical counts, the aggregate bands with their group sizes stated, and the confirmed, praise-only exceptional shortlist. What stays internal, named openly here rather than hidden: every individual candidate's dimension scores, any flag noting a score was made with lower confidence, and any per-candidate category finding that could read as adverse (thin, stale, or a record-only page) — none of that is ever published tied to a name. This is not an oversight; it is the rule that makes the aggregate-only law real rather than nominal. A separate, even more sensitive layer — reading a platform against this project's own policy priorities — is never part of any public render of this method at all.
If you want to build a version of this for your own municipality's candidates, start with Run this in your municipality for what to build first, and see How we verify for how every public claim on this site carries a receipt.
— The Unknown Soldier · uniteTOlove · Toronto