highlight · part iv · 3

Refinement ends with a decision, not with zero.

A pre-test shows which kinds of split are still standing. You decide which kinds get a policy. What remains is a residual, and how to handle it is a second decision.


what still splits · after round 5your call
Blinded, but of whom wrote a ruleround 2 · name the assessor
Sample size: randomized, or analyzed wrote a ruleround 4 · two fields, not one
Registry entry conflicts with the paper ⚠ leave itno policy would hold
Pre-1990 papers, methods in one paragraph ⚠ leave ita rule here would name the paper

Five pilot papers cannot tell you how many cells will split. They can tell you what kind. Each kind is one question: is there a policy here, or only a patch for one paper?

Stop when what still splits has no policy behind it. The residual is the team's to handle: read it, or accept it as error.

@karlrohe

A codebook is never finished, and a run still has to ship. The loop in Part III has no natural floor. Every rule you write can open a new edge, and a corpus of two thousand papers holds cases your ten pilot papers did not. So the loop cannot end on zero. It ends when what is left has no policy behind it. What is left is a residual, and the team decides what to do with it.

A pre-test shows kinds, not rates. Ten papers and three readers give you thirty readings per question. That is enough to see that registry conflicts split and dose does not. It is not enough to say how often. Three readers on a genuine coin flip all agree one time in four by luck, so the pilot understates what is ambiguous. What the pilot gets right is the shape of each split. The shape is what you are deciding on.

Each remaining kind is either a policy or dust. A policy is a rule a colleague on another project would recognize as part of the definition. Write it: the codebook takes a position and every document gets the same treatment. Dust is the rest: small, isolated fragments, no two alike. A registry that contradicts its paper. A 1985 methods section with no subsections. A table whose totals do not reconcile. A rule for any one of them names the paper. It makes the codebook a list of exceptions, and every exception can open a new edge on the next batch. So dust stays out of the codebook.

What happens to the dust is a second decision, and it belongs to the team. Read the flagged cells by hand, as a review queue. Or accept them as error and say so in the methods. Both are defensible. Neither is the codebook's job. Some dust gets flagged and some does not; the flag is how you see part of it, not what defines it.

The rate shows up in deployment, as a readout of the decision you already made. Run the corpus. Cells where the readers agree pass through. Cells where they split are mostly the dust you chose to leave, and now you know how much of it there is. That is when the second decision gets made with a number in hand. That is the two conjectures from Part 0 doing their work: agreement licenses trust on what passed, and disagreement directs your attention to what you kept. Neither says the codebook is right. The next card is about that.