Addendum (2026-08-16) — status of the five agenda items above, and nine questions this round is asked to consider — 9 decided 2026-08-16
Status of the original five. Item 1, the dual-gate structure, is partly overtaken and newly complicated: the AND combinator no longer joins two pass decisions because there are none, but on the epoch-2 bank the gates behave differently — Gate A produces a resolvable adaptation gap and Gate B does not — which is the divergence the structure was built to expose, arriving from an unexpected direction; and whether 0.70 is defensible as a floor now has a measured answer, which is no. Item 2, the per-provision minimum, is partly overtaken: the 0.50 floor no longer blocks issuance, it labels a provision line, and the substantive questions remain open with the epoch-2 spread (Art. 15 at 0.6033 taught against Art. 26(5) at 0.7983) as concrete material. Item 3, seeded serving, is materially changed and has now been exercised: opaque handles mean a seed no longer reproduces a paper on its own, and reproduction requires the attempt document, which an auditor has and a candidate does not — comment is specifically invited on whether that is an acceptable audit basis; the commitment scheme itself is unchanged and has been run through a full epoch roll. Item 4, retakes and nondeterminism, is dissolved in one direction and sharpened in two: with no pass line there is nothing to retry toward, but the honest margin is 0.16 rather than plus or minus 0.05 and its dominant term is the item draw, and the epoch surfaced a condition the round never contemplated — a token ceiling that deterministically produces no measurement rather than a noisy one. Item 5, rotation and exposure, is unchanged in mechanism, exercised in practice and newly pressured: 32 items retired and published in full with one held item disclosed, but none of the 32 retired by exposure, which is a case the policy as written does not cover. Nine questions are put for comment. One, should the reference band stay population-relative, become model-relative where a reference point exists, or be replaced by a required two-arm measurement reporting a within-model delta. Two, the adapted arm is published per model so an operator can reproduce the comparison, but it is equally a published target — publish both arms per model, the untaught arm only, or both only in aggregate. Three, two reports differing by less than 0.16 are not distinguishable measurements: print the margin on the report face, keep it in the reference file, or report an interval rather than a point score. Four, raising items served per attempt from 24 to 54 roughly halves draw noise, doubles exposure per attempt and invalidates every reference figure published so far until the panel is re-run — adopt and re-measure, or defer. Five, items replaced for being too easy were never exposed out, and publishing them publishes a list of what untaught models answer correctly; this round took the publish option for all 32 retired EU items, so the question is now whether that was right against a concrete published file rather than a hypothetical. Six, above the untaught range is not a pass and the file says so, but it will be read as one — is there a formulation that resists that reading, or should the band be numeric only. Seven, adaptation-context quality as the Gate B blocker: the round rebuilt the bank and the Gate B gap did not move because the context moves the value layer by 0.035 on this pack against 0.294 on the Korean one, yet the context is an unversioned build artifact with no quality bar, pinned by sha only so panels stay comparable — should an adaptation context be a first-class versioned quality-gated object with published metrics, and does rebuilding it invalidate the pack's reference entry the way a bank epoch does. Eight, Gate A as the primary discriminator: on the epoch-2 bank the public gate produced the only resolvable adaptation signal while the private gate stayed inside noise, and Gate A scores are floors precisely because its items are published, which is why the design never let it certify alone — should a score report lead with the Gate A delta where it is the only resolvable signal, and what does it mean for the dual-gate argument if the public gate is the more informative instrument on a repaired bank. Nine, the reasoning-token budget as a signed measurement condition: should the ceiling be a signed field in the score report conditions block alongside item counts and label masking, should the reference distribution record the ceiling each cell was measured at given that it currently does not, and is a panel in which one model was measured at a different ceiling still one panel or two. UPDATE (2026-08-16): the operator decided Q2, Q10-Q12, Q14, Q16-Q19; Q16/Q18/Q19 (report margins, gate-discriminator notes, signed maxTokens) are deployed. Comment remains open on the undecided questions until 2026-09-13.