Vouch · the evaluation

A second reader’s job is to read independently.

Vouch routes a referral letter only when two independent readers name the same specialty. If they differ it says so and leaves the choice with the clinician. This page is what that rule costs and what it buys, measured on letters no reader had seen.


How it was measured

153 invented referral letters, 5 samples from every reader on every letter, so 765 calls per reader. The specialty each letter was written to was recorded before any reader saw it. Results below are the 78 held-out letters only, because the abstention threshold was fitted on the other split and quoting that would be quoting the letters the rule was tuned on.

25 of the 78 were written to read two ways, so both readings are defensible. Lenient counts either as right. Strict counts only the seeded one. Both are shown, because quoting one without the other would let the scoring rule do the arguing. The asks are identical under both: the rule changes what counts as correct, never who disagreed with whom.

Each reader alone

ReaderRightWrongStrict
Hosted 1681058 / 20
Open-weights671161 / 17
Hosted 2641456 / 22

The open-weights reader is a peer of the hosted ones on its own. That is what makes the pairing question interesting rather than a fallback.

Each routing rule

Forced errorsAsksout of 78 held-out letters
022Hosted 1 + Open-weights613Hosted 2 + Open-weights716Hosted 2 + Hosted 1717all three agree622majority vote82
Read the bottom row against the ones above it. Majority voting has the shortest ask bar and the longest error bar, because a rule that refuses to ask has to resolve every disagreement silently. The exact counts, and how many each rule got right, are in the table below.
RuleCorrectForced errorsAsksStrict
Hosted 1 + Open-weights5961353 / 12 / 13
Hosted 2 + Open-weights5571651 / 11 / 16
Hosted 2 + Hosted 15471749 / 12 / 17
all three must agree5062246 / 10 / 22
majority of three688261 / 15 / 2

Read the last row first. Majority voting scores highest and is the trap: it almost never asks, so every disagreement between readers is resolved silently and would land as a forced error on a real referral. A forced error is a letter in this test assigned to a specialty other than its label, with nobody told. An ask is a clinician being asked a question they were always the right person to answer.

Vouch never votes between readers. It does vote inside one: each reader is sampled five times and the modal answer is taken, and how many of the five agreed is the consistency score. That vote makes an ask more likely, because low agreement raises the nonconformity and trips the threshold. A vote between readers does the opposite: it resolves a disagreement and nobody is told.


How far this goes

  • It measures the routing rule, not clinical accuracy. The specialty each letter was written to was set by the model that wrote it, which makes it a clean test of whether two readers agree and a poor one for what a GP would do. Clinician labels for a subset come next, and their answer replaces this one.
  • Counts, not percentages. At 78 letters a single letter moves a rate by more than a point, so the counts are given and the arithmetic is left to you.
  • Two readers serve. Three were tested. The third is an open-weights model, brought in as an independent opinion to check whether the disagreements were real or an artefact of one model family. It has never answered a clinical request and is not in the serving path. Which models these are is not published; the providers are named in the data processing agreement.
  • Outcomes are a pilot’s job. Whether a patient waits less is measured in a practice over months, not in an evaluation, and nothing here claims it.

Every figure above is computed from the committed run records when this page is built, not typed in. See the boundary it all sits behind on the referral screen: ask the live service to choose a hospital and it answers 409.