You are running a **blind** scoring pass. The point of this exercise…

← back to session · subagent transcript, not indexed

2026-08-28

User:

You are running a blind scoring pass. The point of this exercise is that two previous scorers, using the same named instruments on the same text, disagreed by a factor of five — one scored a sample at 7% and the other at 40%. Before anything can be concluded from those runs, we need one hand scoring everything under one standard.

The samples

Three prose samples, in redacted: private: scene-A.md, scene-B.md, scene-C.md`. They are different tellings of the same events — a guild's roster being corrected at a table, with the same cast.

Their provenance is withheld from you deliberately, and you must not try to work it out. Do not search the repository for matching text, do not open samples/round6-isolated/, samples/round7-wants/, or chapters/019-a-question-of-rosters.md, and do not speculate in your report about which is which. If you notice something that looks like a tell, ignore it and say in your report that you did. One sample is substantially longer than the others; that is an artefact of sampling, not a clue, and your instruments are rates rather than counts.

You may read notes/dialogue-voices.md (the instruments and the cast's voice markers), notes/canon/characters/ and notes/voices.md. Read CLAUDE.md first and obey it.

The instruments, and why you must pin them down first

From notes/dialogue-voices.md:

Before you score anything, write down your operationalisation in enough detail that a third party could reproduce your numbers: what exactly counts as a "turn" or a "line" (a paragraph of speech? a sentence? an uninterrupted speech act?); what you mask and what you leave; what threshold of confidence counts as "places"; who you swap each speaker with and how you choose; and what "would reject as impossible" means as against "would find surprising". The disagreement between the previous two scorers is almost certainly hiding in these definitions, so make yours explicit and then hold to them identically across all three samples.

Score the three samples in a randomised order of your own choosing, not A-B-C, and do not look back at an earlier sample's score while doing a later one.

Deliverable

Write redacted/samples/round8-blind/scores.md:

Then commit.

Do not edit anything outside `samples/round8-blind/.

Your final message: the three-way table, the spread against your error bar, and your verdict on whether the instruments can support cross-run comparison.

I'll start by reading the required context.

18 tool calls

🔧 Read CLAUDE.md

✅ 109 lines, 6054 chars

🔧 Read notes/dialogue-voices.md

✅ 1098 lines, 71979 chars

🔧 Bash: cd redacted: private && wc -w *.md && ls redacted/notes/canon/characters/

✅ 13 lines, 175 chars — 29 BRIEF.md / 1576 scene-A.md / 4189 scene-B.md

🔧 Bash: cd redacted: private && cat BRIEF.md && echo ===== && cat scene-A.md

✅ 5 lines, 211 chars

🔧 Bash: cd redacted: private && cat scene-A.md

✅ 175 lines, 8133 chars

🔧 Bash: cd redacted: private && cat scene-C.md && echo "=====B=====" && cat scene-B.md

✅ 258 lines, 29161 chars

🔧 Bash: cd redacted/notes/canon/characters && for f in redacted.md edwin.md peri.md minor.md;…

✅ 43 lines, 2239 chars

🔧 Bash: cd redacted/notes/canon/characters && for f in redacted.md edwin.md peri.md minor.md;…

✅ 244 lines, 14854 chars

🔧 Bash: cd redacted/notes/canon/characters && grep -n -i -A25 "redacted\|redacted\|redacted

✅ 118 lines, 7670 chars

Now I'll write the operationalisation and scores. Let me set up the file.

🔧 Bash: mkdir -p redacted/samples/round8-blind && cat > redacted/samples/round8-blind/scores.…

✅ 1 lines, 4 chars

🔧 Bash: git add samples/round8-blind/scores.md && git commit -q -m "Round 8: one-hand blind scori…

✅ 1 lines, 89 chars — 8a20653 Round 8: one-hand blind scoring of three samples, with the instruments' error bar

Written to redacted/samples/round8-blind/scores.md and committed (8a20653). Scored in the order C, A, B, with C's swap re-scored after B as the error-bar check.

The three-way table

A B C
Turns scored 33 53 16
Prop-cover 18/33 (55%) 25/53 (47%) 10/16 (63%)
Swap: rejections on standing/fact 9/15 8/15 9/14
Swap: voice alone 4/15 (27%) 3/15 (20%) 3/14 (21%)

Spread against error bar. The swap re-score of C returned the same total (3/14) but with one call flipped and one reconsidered underneath it — so a swap score is good to about ±1–2 lines, i.e. ±7 to ±14 points. Against that, 27/21/20 is one number three times: the swap test cannot separate these samples. Prop-cover's 16-point range is about equal to my uncertainty on C, whose sixteen turns make every single call worth 6 points. The real spread is between speakers, not samples: redacted runs 36% in one sample and 100% in another, redacted is worst everywhere (his bare-figure opener, "Nineteen on the books. Fifteen.", appears in all three and is Edwin's declared passport in every one), redacted best. In A the most-heard speaker is the least-placed, which moves that total several points on staging alone.

redacted — the word "Agreed." — and is unscoreable, which is not the same as 0%.

Verdict on the instruments. The swap test cannot support cross-run comparison. Three reasons: resolution (one judgement is 7 points); it is not scale-free, because the swap partner is whoever happened to get the most turns, and lines swapped into redacted survive where the same lines swapped into redacted break — B's low score is partly just that; and the bar is a property of the scorer. The single choice that most changes the result is impossible vs surprising: I used impossible and got ~20%, and would have got ~60% on the same lines under surprising — a factor of two and a half, enough to explain the 7-vs-40 on its own. Separating standing from voice is worth another factor of four (standing rejections ran 8–9/15 against voice-alone 3–4/15). Prop-cover is better but under-reports any record-centred scene by its own 2026-08-28 ruling, and rewards long turns. Both work as per-line diagnostics — they found real miscast lines, quoted in the file — and should not be quoted as totals.

Two samples carried provenance headers. I read them before I knew what they were, held them out of scoring, and did not search the repo.