NOTES — 09 · UNRELIABLE NUMBERS

I wrote the test first. Then it failed my best item.

Six items about numbers that betray you. Before any of them existed I wrote down two conditions an item had to pass, which felt like ceremony at the time. They went on to kill four versions, including the one I liked most and the one that would have been a lie.

ENGLISH ONLY · ~17 MIN · 6 ITEMS · 2 CONDITIONS WRITTEN FIRST · 4 VERSIONS THEY KILLED · 0 VERDICTS PASSED

Play Unreliable Numbers → All the write-ups → SIGN-IN REQUIRED · FIRST RUN ONLY

01It does not grade your answer

Six items. Each one shows you a small table, asks you to make a decision, and asks how sure you are of a specific claim on a slider. Then one more fact arrives — a sample size, a breakdown, a column that was missing — and it asks both questions again. Then it shows you what was actually true.

What the page measures is the gap between your two confidence numbers, compared against a reference band computed from exactly what was on the screen at that moment. Not against the truth. The truth arrives at stage three and is never used to score anybody, because the player could not have known it and grading someone on information they did not have is not measurement, it is a trick.

Two things follow from that, and they shaped almost every decision below. First, the reference is a band, not a point: a single number promises a resolution the screen never had. Second, in four of the six items the correct decision does not change — only the confidence does. That is the argument. Knowing that a statistical device exists is not the same as knowing which way it points this time.

The six items, and what the extra fact does to each. The reference band is the range a reasonable updater lands in across the priors a player could plausibly hold.
ItemDeviceBeliefDecisionBand
1 · Factory inspectionSample sizedownsame62–80%
2 · A/B testSample sizeupsame99–100%
3 · Surgery success rateCase mixdownflips0–4%
4 · On-time deliveryCase mixholdssame95–100%
5 · Exam-prep pass rateWho is missingdownflips0–2%
6 · Quit-smoking programmeWho is missingholdssame98–100%

02Two conditions, written before any item existed

The failure mode of a game like this is obvious in advance. You want the reveal to land, so you build an observation that looks shocking, then you tell the player their shock was an illusion. If the observation really was shocking under the truth you are about to reveal, the item has just lied to them, and it lied in the exact direction that makes the author look clever.

So before writing a single item I wrote down what an item had to survive.

GATE 1 — the tail Under the data-generating process the item is about to reveal, the two-sided probability of each observed cell must be ≥ 0.05. If the reveal says "you were fooled" about data that was genuinely surprising, the item is the thing doing the fooling. GATE 2 — prior sensitivity Across the priors a reasonable player could hold, the width of the reference band must be ≤ 0.20. If the answer swings on an assumption nobody was told about, the scoring is unfair no matter how correct the arithmetic is.

Both live in validate.py and both re-run on every item, every time an item is added or changed. That matters more than it sounds: a condition you only check once is a ceremony, and a condition that re-runs is a test.

WHY THIS IS WORTH WRITING DOWN FIRST

Every one of the four rejections below arrived with a reason to make an exception. The dramatic version was dramatic. The 3.3σ world was the first seed and looked fine. If the conditions had been written after the items, I would have been writing them to fit items I was already attached to, and every one of those exceptions would have been granted.

03The first item I wrote was a lie

Item 1 is four factories. Three have been inspected 500 times each and sit near 3% defective. The fourth, D, shows 10%. You get one deep inspection; where do you send it? Then stage two reveals the inspection counts: A, B and C were checked 500 times, D was checked twenty.

My first draft had D at 3 defects out of 20. Fifteen percent, nice and alarming, and the reveal says all four factories were truly at 3.0%.

n = 20, true p = 0.03 P(X ≥ 3) = 1 − 0.5438 − 0.3364 − 0.0988 = 0.0210 gate 1 threshold: 0.05 FAIL

Three defects in twenty at a true rate of 3% happens about two times in a hundred. The item was about to tell the player that their alarm was a small-sample illusion, when in fact the observation was genuinely unusual and their alarm was the correct response. I had written, in §12 of my own notes, a warning against exactly this item, and then built it four sections later.

Two out of twenty passes: P(X ≥ 2) = 0.120. Less alarming on screen, honest underneath. That is the trade the gate exists to force, and it is not a trade I would have made by eye.

04The dramatic version was the unfair version

With the observation fixed, the question became what claim the confidence slider should attach to. I wanted something with teeth, so the first version asked how sure you were that D's true defect rate was more than double the national average — above 6%. This is a lovely reveal: a player who has just seen 10% on screen will sit up around 90%, and the reference band lands in the teens. A forty-point drop. Beautiful.

Gate 2 on the dramatic version. n₀ is the strength of the prior a player brings — how many factories' worth of experience they are implicitly carrying about how much plants vary around the national rate.
Claimn₀ = 3050100200400Width
P(D’s true rate > 6%)0.4050.2990.1520.0450.0050.400
P(D’s true rate > 3%)0.7990.7650.7130.6620.6210.179

The dramatic version's answer runs from half a percent to forty percent, and the only thing moving is an assumption the player was never asked about and never told. Two players reasoning perfectly from different priors get scored forty points apart. The item was not measuring their updating; it was measuring a hidden parameter.

The surviving version — is D's true rate above the national 3% — comes in at a width of 0.179, just inside the line I had set at 0.20. It is a duller reveal. It is the one that can be scored.

THE SENTENCE I KEPT

The dramatic version was the unfair version. Not sometimes, and not as a coincidence: the thing that made it dramatic — a claim far out in the tail where the posterior is thin — is precisely the thing that made it swing on the prior. The drama and the unfairness were the same property, looked at from two sides.

05You cannot grade a prior nobody declared

Gate 2 handed me a second problem, and it did not go away by picking a better claim. Even the surviving version needs a prior to have a reference band at all, and the width across plausible priors was 0.179 — inside the limit, but not nothing. A player who came in believing factories vary wildly and a player who came in believing they are all much the same would be graded against the same band.

The fix was to stop treating the prior as something to infer about the player and start treating it as something the item supplies. The national average defect rate, 3%, is now printed on stage one, before any decision. Everyone reasons from the same starting point, so the band is a fair yardstick for all of them.

The obvious worry is that this gives the reveal away. It does not, and the reason is worth being precise about: knowing the national average tells you what the typical factory looks like. It tells you nothing about whether D is typical. The entire question — is this a bad factory or a small sample — is untouched. What the base rate removes is not the puzzle, it is the undeclared variance in how players approach it.

06The same gate, pointed at the author

Item 3 is Simpson's paradox, and the first item where the correct decision flips rather than just the confidence. Hospital A shows 78.5% surgical success, hospital B shows 74.0%. You are a severe case. Stage two splits by severity and B is better in both strata; A only looked better because A took far more mild cases.

The counts in that table were not chosen by me. I drew them from the true rates with a seed, because hand-picked counts are exactly how an author accidentally builds a world that cannot happen. The first seed produced this:

seed 11 B · severe: 497 / 700 (true rate 65%) two-sided P = 0.001 3.3σ seed 0 A·mild 0.608 A·severe 0.952 B·mild 0.421 B·severe 0.453 all pass

I was one accepted seed away from shipping a 3.3σ world as if it were an ordinary Tuesday at a hospital. Nothing about the table would have looked wrong — 497 out of 700 is 71%, six points above the true rate, which reads as unremarkable if you are not computing the tail.

What I find interesting is that gate 1 did a different job here than it did in item 1. In items 1 and 2 it was asking does the reveal lie to the player. Here it was asking is the world the author drew a typical one. Same threshold, same code, two distinct failure modes. I did not anticipate the second when I wrote the condition; the condition was simply general enough to catch it.

07The item that disagreed with its own reveal

Item 4 is item 3's partner: the same device, pointed the other way. Your delivery address is outside the capital region. Courier A shows 90.5% on-time, B shows 83.4%. Stage two splits by region — and the case mix is as lopsided as item 3's, 700/300 against 300/700 — but A is ahead in both regions. The alarm goes off and nothing changes.

Seed 0 passed gate 1 on all four cells. It was still unusable.

seed 0 A · capital 91.9% A · non-capital 91.7% difference: 0.2pp reveal "region really does affect on-time rate: A is 92% in the capital, 89% outside" the screen and the reveal disagree.

Every individual cell was plausible. The pattern the item was built to talk about had been sampled away. A player looking at 91.9 and 91.7 would read the reveal's claim about regional effects as authorial assertion, unsupported by the very table they were shown.

So a third condition went in, and it is the only one added after the fact rather than before:

TYPICALITY gate 1 asks: is each cell individually plausible? typicality asks: is the drawn world typical in the respects the item actually talks about? item 4: stratum gap must land in 3.5–6.5pp and the region effect must survive in the observation seed 2 passes.

The distinction is the useful part. Gate 1 is a per-cell condition and it is blind to the shape of the thing the item claims. Typicality is claim-specific by construction — you cannot write it in general, you write it once per item, from the sentence the reveal is going to say out loud. An author who only runs the general condition will keep producing items that are individually defensible and collectively beside the point.

08Three devices, each used in both directions

The finished structure is three pairs. Each pair uses one device twice, once in each direction.

The argument of the whole thing. Learning the device is half the job; the direction has to be re-measured every time.
DeviceDirection that reversesDirection that holds
How big the sample is1 — lowers your confidence2 — raises it
Splitting by composition3 — the conclusion flips4 — nothing changes
Who is missing from the table5 — the conclusion flips6 — survives the worst case

Item 2 is the one that makes the pairing do work. It is an A/B test on signup-button copy: A converts at 2.00%, B at 2.12%, a difference you would not look at twice. Stage two reveals that each version was shown to 200,000 people. A player who came out of item 1 with the rule small sample means trust it less walks into item 2 and applies it backwards. I imitated that player in a headless run — 58% down to 44% — against a reference band of 99–100%. The gap is visible on the player's own summary screen, which is the point.

Sample size is not a device for lowering confidence. It is a device for setting the weight of the evidence, and the direction is set by the situation.

The frame I got wrong

My first draft of this structure was a two-axis grid — belief up/down/same crossed with decision same/flips — and I wrote that the remaining cells needed filling. That framing is wrong, and it took building items 5 and 6 to see why. Items 5 and 6 land in the same cells as 3 and 4, and there is nothing whatever wrong with them. Design to fill a grid and you produce puzzles, not items. The direction label is a diagnostic you record afterwards; it is not the specification.

One cell is unfillable in principle, which is a more interesting fact than it first looks. The only way confidence stays put while the decision changes is asymmetric consequences — when being wrong costs differently in the two directions. Same probability, different call. That is not a story about numbers betraying you; it is a story about numbers not being the whole input. It belongs on the closing screen as a sentence, not as a seventh item.

09The item that exists so the previous one does not teach the wrong thing

The third device is different in kind from the first two. Sample size and case mix are both about reading the table properly, and both assume every number in the table is true. The third device asks who the table left out, and no amount of careful reading catches it. Absent rows are invisible.

Item 5 is an exam-prep school. A advertises a 77.6% pass rate, B advertises 54.0%. Stage two: 500 enrolled at each, but only 255 of A's finished the course. A's headline counts the 255. On an enrolment basis it is 39.6% against B's 51.6%, and the decision flips hard — band 0–2%.

If item 5 were the last word, the lesson a player takes away is missing data means the number is garbage, and that is a worse heuristic than the one they came in with. It is unfalsifiable, it applies to every real dataset, and it terminates thinking. Item 6 exists to stop it.

Quit-smoking programme A B enrolled 400 400 outcome confirmed 333 365 outcome unknown 67 35 succeeded 206 160 headline rate 61.9% 43.8% worst case for A: all 67 unknowns failed, all 35 of B's succeeded 51.5% 48.8% margin 2.8pp

The move item 6 teaches is the only transferable skill in the whole set: fill the missing rows with the worst case for your preferred answer and see whether the conclusion is still standing. Five of the six items say do not. This one says here is what to do instead.

The 2.8-point margin is not an accident of the data. It is a typicality condition: 0.005 < margin < 0.05. A wide margin makes the exercise pointless — of course it survives — and a negative one makes the item a duplicate of item 5. Surviving by a hair is the entire content.

Gate 2 shakes something different here

In items 1 through 5, gate 2 varies the prior over a parameter. In item 6 there is no parameter to have a prior about; what a reasonable person could disagree on is what to assume about the people who were never followed up. So the sweep varies that instead: unknowns behave like the confirmed, unknowns are wholly unknown, unknowns break against A, unknowns break for A. Width 0.025.

That is the second time a gate did a job I had not specified for it, and both times it happened because the gate was written as a question rather than a formula. Does the answer swing on something the player was not told? survives the change of subject. Compute the posterior under five beta priors would not have.

10What the screen is not allowed to say

A page about misleading numbers cannot afford to be sloppy about its own presentation, so the screen rules are stated as prohibitions and checked per item.

  • No reference band at stage one. There is nothing to compute yet. Grading someone against information you have not given them is the exact sin the whole page is about.
  • The word "Bayes" appears zero times on screen. The labels are "you" and "reference". It appears freely here, in the write-up. The screen is an experience; the write-up is an explanation; they are allowed different vocabularies and they are not allowed to swap.
  • A band, never a point. Following the confBand precedent from Who Wrote This. A single number claims a resolution the screen never had.
  • No adverbs. "Substantially", "nearly", "slightly" — all cut. Numbers are neutral and adverbs are not.
  • No verdict. The last line is "what changed is for you to judge". There is no score, no rank, no grade. The site palette has a warning red and the game screen never once uses it, because nothing on it is being marked wrong.
  • The bands are constants, not computed in the browser. validate.py produces them and they are pasted in. Two implementations of the same quantity drift, and when they drift the screen is lying. Same reasoning as the single-computation-path rule from the regime project.
MOVING-GOALPOST AUDIT

Run per item, because this is the failure the page would least survive. Stage-one decision: correct under the information given, no penalty. Stage-one confidence: not scored at all. Stage-two decision: still correct. Stage-two confidence: compared only against a band computed from what is on screen. Stage-three truth: never used for scoring. No point at which the goalposts move.

Two items needed wording changes to keep that true. Item 3 puts the word "severe" on stage one, before the split exists, because the claim being rated is about severe patients — if that word first appeared at stage two the target would have moved. Item 5 pins its claim to "if you enrol" for the same reason. A sharp player can then get suspicious at stage one, which is a reward, not a leak.

11The readout I built and then took down

Halfway through I proposed something that sounded good: since players learn as they go, do not fight it — measure it. Put the six first-round confidences in a row and show the decline curve. Calibration improving in real time, on your own summary screen.

Then I looked at three of them: 92, 58, 85.

They do not decline, and they were never going to, because the thing that moves them is not learning. It is how plausible each item's opening claim looks on its face. Item 1's "10% against a 3% average" reads as obviously true and pulls people to 90-plus. Item 2's "2.00 against 2.12" reads as noise and pulls people to the middle. One person's six numbers are a measurement of six different item surfaces, not of that person changing.

The measurement I had proposed was completely confounded with item difficulty, and I proposed it anyway, in the middle of building a thing whose subject is exactly this. I pulled the claim off the screen and put the limitation in its place: the summary now says in as many words that these six numbers are not comparable to each other.

It also settled an open question. Randomising item order is not a nice-to-have for a future version; without it, even aggregate data across many players cannot separate learning from item difficulty, because every player meets the items in the same sequence. The order is still fixed today. That is a known defect, recorded as one.

12What playing it caught that computing it did not

Everything above came out of arithmetic. The following came out of opening the page and looking at it, and none of it was visible from the maths.

Found by playing, not by computing. The last one was found by a person who was not me, which is the only reason it was found at all.
What was wrongWhy it mattered
The slider started at 50A page about anchoring shipped with an anchor. 50 is not neutral; it pulls every player the same way. Now randomised 15–85 per item, and the starting value is recorded with the answer.
The probability on screen was wrongIt said 24% for "2 or more defects in 20". 24% is the two-sided figure the gate uses; the player's question is one-sided and the answer is 12%. Two different quantities were living under one variable name.
Reference and truth read as contradictoryRight after "all four were truly at 3%", a band of 62–80% looks like a typo. The band is what was knowable then; the truth is what came after. Obvious in the author's head, absent from the screen, so a paragraph now says it.
Chart label collided with the band labelWhen the player's stage-two value lands inside the reference band the two labels overlap. Label drops below the dot within ±8 of the band.
An em dash read as a ruleThe empty-state placeholder was a dash at 30px Archivo 900, which renders as a horizontal line across the panel. Replaced with a 12px mono hint.
Item 2 never explained its own domainThe draft said "conversion rate" and "one version ships to everyone" and stopped. What it is a rate of, and why you must pick one, were nowhere on screen.

That last row is the one worth dwelling on. I had played item 2 headlessly five times and read every word on the screen each time. The gap was invisible to me and immediately visible to the first person who was not me.

THE RULE THAT CAME OUT OF IT

A person who knows the content cannot see that the explanation is missing. Item 1 got away with one line because "defect rate" and "deep inspection" explain themselves; I generalised that into a rule and it was not one. Domain explanation is now a checklist item per item: is it on screen what these numbers are a rate of, and why this decision has to be made at all?

The fix changed no numbers, so no gate had to re-run. That is the difference between explaining an item better and changing it.

One more, from the site itself

Stage transitions slid gently down the page instead of cutting. The cause was html{scroll-behavior:smooth} in the shared site CSS, which turns window.scrollTo(0,0) into an animation; on item 2's long panel the player starts reading from the middle before the scroll finishes. A stage transition is not a movement through a document, it is a screen replacement, so it is behavior:'instant' now. First place the inherited site CSS collided with this page's interaction model, and it will not be the last.

13What it refuses to claim

The page records your first run and only your first run — enforced in the database, where user_id is the primary key, so a second attempt is rejected by Postgres rather than by a rule in the browser that could be got around. The screen says so before you start, in a warning, because once you have seen the reveals you are a different subject and the second run measures something else.

What a saved run therefore is: one person's six pairs of confidence numbers, each with a reference band, on six items in a fixed order, on a single occasion, with no repeat. What it is not: a calibration score, a ranking, a claim that you are well or badly calibrated in general, or a measurement that survives comparison between items. The /me/ card shows what you did next to what the bands were, and passes no verdict on it. Six items is not a psychometric instrument, and the page does not have a sentence anywhere that pretends otherwise.

14What I still do not know

  • Whether the pairs teach the thing they are built to teach. The argument — that direction must be re-measured — is a design claim, not a measured one. Testing it needs the randomised order from section 11, several hundred first runs, and a comparison of second-item behaviour by which item came first. None of that exists yet.
  • Whether 0.20 is the right width for gate 2. It was chosen by judgement before any item existed, which is the right time to choose it and the worst time to calibrate it. Item 1's surviving version came in at 0.179. A threshold of 0.15 would have killed it and I do not have an argument for why that would have been wrong.
  • Whether the base-rate-on-screen fix generalises. It worked for item 1 because the national average is a real quantity a real decision-maker would have. Not every item has one of those lying around, and I do not know what to do when the declared prior would itself be an artefact of the item.
  • Whether typicality can be written down in general. Right now it is hand-written per item, derived from the sentence the reveal says. That works and it does not scale, and I suspect the reason is that it is not really a statistical condition at all — it is a consistency check between prose and data, and those may simply have to be written one at a time.
  • What a seventh item would be for. Three devices × two directions closed the structure. A seventh only makes sense with a new device, and the only candidate I have is asymmetric consequences, which section 08 argues is not a statistics item at all.

Six items, one run, no score. You judge, you say how sure you are, then one more fact arrives — and what gets recorded is not the answer but the move. Sign-in required; the first run is the only one that is kept.

Open Unreliable Numbers →