Vouch · the evaluation
A second reader’s job is to read independently.
Vouch routes a referral letter only when two independent readers name the same specialty. If they differ it says so and leaves the choice with the clinician. This page is what that rule costs and what it buys, measured on letters no reader had seen.
How it was measured
153 invented referral letters, 5 samples from every reader on every letter, so 765 calls per reader. The specialty each letter was written to was recorded before any reader saw it. Results below are the 78 held-out letters only, because the abstention threshold was fitted on the other split and quoting that would be quoting the letters the rule was tuned on.
25 of the 78 were written to read two ways, so both readings are defensible. Lenient counts either as right. Strict counts only the seeded one. Both are shown, because quoting one without the other would let the scoring rule do the arguing. The asks are identical under both: the rule changes what counts as correct, never who disagreed with whom.
Each reader alone
| Reader | Right | Wrong | Strict |
|---|---|---|---|
| Hosted 1 | 68 | 10 | 58 / 20 |
| Open-weights | 67 | 11 | 61 / 17 |
| Hosted 2 | 64 | 14 | 56 / 22 |
The open-weights reader is a peer of the hosted ones on its own. That is what makes the pairing question interesting rather than a fallback.
Each routing rule
| Rule | Correct | Forced errors | Asks | Strict |
|---|---|---|---|---|
| Hosted 1 + Open-weights | 59 | 6 | 13 | 53 / 12 / 13 |
| Hosted 2 + Open-weights | 55 | 7 | 16 | 51 / 11 / 16 |
| Hosted 2 + Hosted 1 | 54 | 7 | 17 | 49 / 12 / 17 |
| all three must agree | 50 | 6 | 22 | 46 / 10 / 22 |
| majority of three | 68 | 8 | 2 | 61 / 15 / 2 |
Read the last row first. Majority voting scores highest and is the trap: it almost never asks, so every disagreement between readers is resolved silently and would land as a forced error on a real referral. A forced error is a letter in this test assigned to a specialty other than its label, with nobody told. An ask is a clinician being asked a question they were always the right person to answer.
Vouch never votes between readers. It does vote inside one: each reader is sampled five times and the modal answer is taken, and how many of the five agreed is the consistency score. That vote makes an ask more likely, because low agreement raises the nonconformity and trips the threshold. A vote between readers does the opposite: it resolves a disagreement and nobody is told.
How far this goes
- It measures the routing rule, not clinical accuracy. The specialty each letter was written to was set by the model that wrote it, which makes it a clean test of whether two readers agree and a poor one for what a GP would do. Clinician labels for a subset come next, and their answer replaces this one.
- Counts, not percentages. At 78 letters a single letter moves a rate by more than a point, so the counts are given and the arithmetic is left to you.
- Two readers serve. Three were tested. The third is an open-weights model, brought in as an independent opinion to check whether the disagreements were real or an artefact of one model family. It has never answered a clinical request and is not in the serving path. Which models these are is not published; the providers are named in the data processing agreement.
- Outcomes are a pilot’s job. Whether a patient waits less is measured in a practice over months, not in an evaluation, and nothing here claims it.
Every figure above is computed from the committed run records when this page is built, not typed in. See the boundary it all sits behind on the referral screen: ask the live service to choose a hospital and it answers 409.