Method: Blind double-scoring — how we check our own judgment
Any process that asks a person to read something and assign a score is, at some level, asking for a judgment call — and judgment calls vary from one reader to the next. Rather than assume our scoring is consistent, we test it. This is the method we use to check our own work: a second, independent reader scores a sample blind, we publish how often they agreed with the first reader, and we name, on the record, how every real disagreement was resolved. It exists so that anyone reading a headline number from our platform scoring knows how much of that number is solid instrument and how much is one reader's judgment call.
What it produces
A published, dated agreement report for a scored batch of platforms: how often two independent readers landed on the exact same score, how often they landed within one point of each other, and whether one reader tended to score systematically higher or lower than the other. Alongside that, a named list of every case where the two readers genuinely disagreed, with a plain statement of how each one was resolved. This report is a required read before any scored baseline's aggregate numbers are allowed to be finalized and published.
How it works, step by step
The sample is drawn first, before anyone scores anything. A fixed, randomized selection of candidate-platform pairs is chosen ahead of time, spread across different regions and office types so the check doesn't accidentally land only on the easy cases or only on the hard ones. The selection method and the exact count are written down and dated on the sample file itself, before scoring starts — so the sample can never be quietly redrawn after someone has already seen how the numbers came out.
Two readers, blind to each other. For every platform in the sample, one reader scores it as part of the normal, full-scale scoring run, and a second, separate reader scores the exact same platform independently, using the identical written instructions, with no visibility into what the first reader scored. This isn't a lighter spot-check — the second reader runs the complete scoring process, start to finish, exactly as they would for any other candidate.
Agreement is measured properly, not glossed over. For every pair where both readers produced a score, we compute how often they matched exactly, how often they were within one point of each other, and the average size of the gap between them, for every single dimension scored — not collapsed into one flattering headline number. Whether the two readers agreed on the basic category call (is this a real platform at all, or a brochure, or a record of past work) is checked first and separately, because if they disagreed on the category, comparing their detailed scores on top of that would be meaningless.
We check for a systematic pattern, not just noise. We also look at whether one reader consistently scored higher or lower than the other across the whole sample, rather than the disagreements just being random scatter in both directions. We report this even when it shows nothing unusual, because "no systematic pattern found" is itself a useful, publishable finding.
Every real disagreement is resolved by a named person, on the record. Any pair where the two readers' overall scores diverged by a meaningful amount, or disagreed on the category entirely, is individually re-read by a senior reviewer against the original source material, and the outcome — which reader's score stood, or whether the two were reconciled to a new value — is written down with the reasoning. The score that counts toward any published aggregate always comes from the first, primary reader's pass, decided as a fixed rule in advance, specifically so no one can later pick whichever reader's number looks better.
Everything ships together. The agreement numbers, the list of resolved disagreements, and an honest statement of what this check does and doesn't cover are always published as one report — never a headline agreement percentage on its own, stripped of its caveats.
How it's checked
This method is itself a checking mechanism, so here is what keeps the checker honest.
- The sample can't be gamed after the fact. It's drawn, sized, and dated before any scoring happens, on a file that records exactly how it was drawn.
- Category disagreement is caught before it can hide inside a detail score. We check whether both readers even agreed on what kind of platform this was before we ever compare their point-by-point scores, so a fundamental disagreement can't get buried in an average.
- Every meaningful disagreement gets a named human resolution, not a silent average. Nothing is smoothed over by splitting the difference; a senior reviewer re-reads the actual material and states, on the record, why one score stood or how it was reconciled.
- The full table is published, not a summary. All dimensions, both the exact-match and within-one-point rates, are published — a favorable overall number can't be used to paper over one dimension where readers disagreed a lot.
- The report can't ship without its own limitations attached. A rule we hold ourselves to: the honest gaps in this method (see below) are published in the very same document as the headline agreement numbers, not filed somewhere easy to miss.
The most direct way to challenge our platform scoring is to ask for this report and read the disagreement list yourself — every resolved case is there, with the reasoning, not just the final tally.
What it honestly costs
The second, blind scoring pass costs the same as the primary scoring pass it mirrors — it is not a cheap add-on, it's a second full read by a trained reader for every platform in the sample. Computing the agreement statistics and resolving disagreements both require a person applying real judgment, not something that runs unattended. In one real run covering a province-wide scoring pass, roughly a fifth of the full candidate field was drawn into the blind-check sample and scored independently by a small dedicated team of reviewers, and about one in ten of the comparable pairs needed individual, named resolution by a senior reviewer in the same working session as the main scoring run. Those are one real run's figures, not a promised rate — treat them as an order of magnitude.
Known limits
We would rather publish the gaps than let someone assume this process is more mature than it is.
- We haven't yet written down a formal justification for why the sample size and the way it's spread across regions and office types is the right design — it follows a stated general rule, but that rule itself isn't derived from a specific confidence target in writing yet.
- When a large region is scored in separate batches, at separate times, by different groups of readers, we currently verify consistency within each batch but haven't yet drawn a sample to check consistency across batches — so a direct comparison between two different batches' average scores could carry an unmeasured calibration gap. A modest cross-batch calibration sample is designed as the fix and is planned, not yet run as of this writing.
- The agreement statistics themselves are currently computed by hand for each report rather than through one fixed, reusable calculation tool, so we can't yet guarantee the exact same method is applied identically every single time without a person re-deriving it.
- This method is young. It has been run only a couple of times so far, both on the same underlying scoring process, in a short window. Nothing has gone wrong with the check's own mechanics yet — but the honest reading of that is that it hasn't had enough time to be tested against failure, not that it's proven durable.
What stays internal, and why
The agreement table itself — how often readers matched, by dimension, with no names attached — is exactly the kind of thing we publish, because it's what lets someone judge how much to trust a scored baseline's headline numbers. What stays internal: the specific reasoning tied to a named candidate in the disagreement-resolution list. That material carries the same sensitivity as the underlying per-candidate platform scores it's checking, and those individual scores are never published, in either direction — only in aggregate, and only through the separately-gated praise-only shortlist described in Method: The platform substance baseline. Keeping that boundary here is what keeps the aggregate-only rule real rather than a rule with a side door.
If you're running this kind of scoring process yourself and want a way to test whether your own readers are applying it consistently, this method generalizes to any judgment-scored batch of material — start with Run this in your municipality, and see How we verify for how every claim on this site is expected to carry a receipt.
— The Unknown Soldier · uniteTOlove · Toronto