Method: Blind double-scoring — how we check our own judgment

Any process that asks a person to read something and assign a score is, at some level, asking for a judgment call — and judgment calls vary from one reader to the next. Rather than assume our scoring is consistent, we test it. This is the method we use to check our own work: a second, independent reader scores a sample blind, we publish how often they agreed with the first reader, and we name, on the record, how every real disagreement was resolved. It exists so that anyone reading a headline number from our platform scoring knows how much of that number is solid instrument and how much is one reader's judgment call.

What it produces

A published, dated agreement report for a scored batch of platforms: how often two independent readers landed on the exact same score, how often they landed within one point of each other, and whether one reader tended to score systematically higher or lower than the other. Alongside that, a named list of every case where the two readers genuinely disagreed, with a plain statement of how each one was resolved. This report is a required read before any scored baseline's aggregate numbers are allowed to be finalized and published.

How it works, step by step

The sample is drawn first, before anyone scores anything. A fixed, randomized selection of candidate-platform pairs is chosen ahead of time, spread across different regions and office types so the check doesn't accidentally land only on the easy cases or only on the hard ones. The selection method and the exact count are written down and dated on the sample file itself, before scoring starts — so the sample can never be quietly redrawn after someone has already seen how the numbers came out.

Two readers, blind to each other. For every platform in the sample, one reader scores it as part of the normal, full-scale scoring run, and a second, separate reader scores the exact same platform independently, using the identical written instructions, with no visibility into what the first reader scored. This isn't a lighter spot-check — the second reader runs the complete scoring process, start to finish, exactly as they would for any other candidate.

Agreement is measured properly, not glossed over. For every pair where both readers produced a score, we compute how often they matched exactly, how often they were within one point of each other, and the average size of the gap between them, for every single dimension scored — not collapsed into one flattering headline number. Whether the two readers agreed on the basic category call (is this a real platform at all, or a brochure, or a record of past work) is checked first and separately, because if they disagreed on the category, comparing their detailed scores on top of that would be meaningless.

We check for a systematic pattern, not just noise. We also look at whether one reader consistently scored higher or lower than the other across the whole sample, rather than the disagreements just being random scatter in both directions. We report this even when it shows nothing unusual, because "no systematic pattern found" is itself a useful, publishable finding.

Every real disagreement is resolved by a named person, on the record. Any pair where the two readers' overall scores diverged by a meaningful amount, or disagreed on the category entirely, is individually re-read by a senior reviewer against the original source material, and the outcome — which reader's score stood, or whether the two were reconciled to a new value — is written down with the reasoning. The score that counts toward any published aggregate always comes from the first, primary reader's pass, decided as a fixed rule in advance, specifically so no one can later pick whichever reader's number looks better.

Everything ships together. The agreement numbers, the list of resolved disagreements, and an honest statement of what this check does and doesn't cover are always published as one report — never a headline agreement percentage on its own, stripped of its caveats.

How it's checked

This method is itself a checking mechanism, so here is what keeps the checker honest.

The most direct way to challenge our platform scoring is to ask for this report and read the disagreement list yourself — every resolved case is there, with the reasoning, not just the final tally.

What it honestly costs

The second, blind scoring pass costs the same as the primary scoring pass it mirrors — it is not a cheap add-on, it's a second full read by a trained reader for every platform in the sample. Computing the agreement statistics and resolving disagreements both require a person applying real judgment, not something that runs unattended. In one real run covering a province-wide scoring pass, roughly a fifth of the full candidate field was drawn into the blind-check sample and scored independently by a small dedicated team of reviewers, and about one in ten of the comparable pairs needed individual, named resolution by a senior reviewer in the same working session as the main scoring run. Those are one real run's figures, not a promised rate — treat them as an order of magnitude.

Known limits

We would rather publish the gaps than let someone assume this process is more mature than it is.

What stays internal, and why

The agreement table itself — how often readers matched, by dimension, with no names attached — is exactly the kind of thing we publish, because it's what lets someone judge how much to trust a scored baseline's headline numbers. What stays internal: the specific reasoning tied to a named candidate in the disagreement-resolution list. That material carries the same sensitivity as the underlying per-candidate platform scores it's checking, and those individual scores are never published, in either direction — only in aggregate, and only through the separately-gated praise-only shortlist described in Method: The platform substance baseline. Keeping that boundary here is what keeps the aggregate-only rule real rather than a rule with a side door.

If you're running this kind of scoring process yourself and want a way to test whether your own readers are applying it consistently, this method generalizes to any judgment-scored batch of material — start with Run this in your municipality, and see How we verify for how every claim on this site is expected to carry a receipt.

— The Unknown Soldier · uniteTOlove · Toronto