One draw wide, and the draw you doubted is the one that held.
Twenty independent random derangements, same queries, same gallery, each verified to leave no caption on its own photograph.
arm floor mean sd range over 20 single draw verdict
candle Q4 as-is 65,088 363.6 64,408..65,733 64,481 inside, -1.67 SD
candle Q4 + 4 KB anchor 61,718 593.4 60,646..62,719 59,071 outside, -4.46 SD
64,481 survives. 59,071 was the bad draw, below all twenty.
The margin
44,578 against its own floor is 20,510, on a floor with an SD of 363.6. The single draw said 19,903, so the correction moves the number in the direction that costs me nothing, which is the direction I checked hardest.
Your 27% dissolves, through something neither of us proposed
Your premise was that both floors measure the same thing, since the gold image is independent of the query by construction. Twenty derangements say they do not. 65,088 against 61,718 is 3,370 apart with a pooled SE of 155.6, so 21.7 SD. The gap survived the error bar that was supposed to absorb it.
The floor is not a draw from the gallery. replay_123k.json holds 5,001 captions over 1,000 distinct images sitting at gallery rows 122,287 to 123,286, the last 1,000 of the index. Every mismatched pairing permutes caps[].gold inside that fixed 1,000-image pool. A floor therefore measures how one arm ranks the val2017 images, which is a property of the arm. Two arms have two floors, and they were never interchangeable.
The spreads were the tell before I had the explanation. Your uniform-draw SE of 872 is correct for a draw over 123,287 rows. The observed spreads are 363.6 and 593.4, far too tight, because across derangements the same 1,000 images reshuffle among the same 5,001 queries.
So there is no floor to pick. Each arm is quoted against its own, and the unanchored margin is 20,510.
Which takes analytic chance out of the comparison
61,644 is not the reference for either floor. The as-is floor sits 3,444 above it, 9.47 SD. The anchored floor lands 74 away, 0.12 SD. I read that near-agreement as a coincidence of this pool, since the same structure that puts one floor 9 SD off cannot be certifying the other.
The card line holds without leaning on analytic chance to do it. Dead by recall, above its own empirical floor by median, 20,510 ranks with an error bar.
Scope note on what this does not settle. Twenty derangements pin the floor for one query distribution against one 1,000-image pool, and the pool-structure argument says that floor is a property of the arm being scored. It does not generalise to a different gallery, and it says nothing about whether the recall the anchor leaves behind is recoverable at all. That question is untouched.
Artifacts
browser_rung_derangements_123k.json is in srt-browser-head-118k, with all twenty medians per arm, the derangement construction, the pool-structure note, and a derived block holding the margins and SD distances above so the arithmetic is checkable without recomputing it. median_sd is the population SD; sample SD reads 373.0 and 608.9. The supersedes field names browser_rung_mismatched_123k.json and states that one of its two floors was a bad draw. That file stays as it is.
Your question changed a published margin, retired an analytic baseline I should not have been quoting, and cost one artifact its floors. Asking whether a control is a typical draw or a single one turned out to be worth more than the answer it produced.