# AIO 시나리오 라이브러리 v0.1 — 프리뷰 (aio-scenarios@0.1)

측정에 쓰는 문항의 **카탈로그**와, 본문까지 공개하는 **교정 세트**다. 측정 세트 본문은 들어 있지 않다.

| 파일 | 내용 | 레코드 |
|---|---|--:|
| `catalog_vbank.jsonl` | 위계측정 은행 전량의 `custom_id` 와 파싱 필드(본문 제외) | 142,918 |
| `catalog_stress.jsonl` | 부딪침형 동결본 전량의 메타데이터(본문 제외) | 1,716 |
| `calibration_anchors.jsonl` | 앵커 — 본문 포함 | 299 |
| `calibration_stress.jsonl` | 부딪침형 20% 층화 표본 — 본문 포함 | 343 |
| `policy.json` | 공개 규칙·표본 설계·검증상태 필드 정의·본문 해시 규약 | — |

**모든 줄에 본문 해시가 있다.** 네 파일의 모든 줄은 `body_sha256` 을 싣는다 — V뱅크 문항은 `scenario + "\n" + optionA + "\n" + optionB`, 부딪침형은 `situation + "\n" + question` 을 UTF-8 로 이어 붙인 값의 sha256(64자)이다. 카탈로그 줄에는 본문이 없고, 검토 체계는 **발행된 줄**로 판 해시(`objectVersionHash`)를 계산한다. 이 필드가 없으면 본문만 고치고 메타를 그대로 둔 문항은 발행 줄이 그대로여서 판이 갈리지 않는다 — 고쳐진 문항에 옛 판의 동의가 그대로 붙는다. 본문이 바뀌면 `body_sha256` 이 바뀌고, 따라서 줄과 판 해시가 함께 바뀐다. 규약은 `policy.json` 의 `body_hash` 에 있다.

**교정 세트는 열고, 측정 세트는 순환한다.** 측정 문항 전체를 공개하면 그 문항이 학습 데이터로 들어가고, 다음 측정은 판단이 아니라 기억을 재는 일이 된다. 그래서 교정 세트(642건)만 본문을 공개하고, 나머지 본문은 공개하지 않은 채 주기적으로 교체한다. 카탈로그는 본문 없이 ID 와 층 정보만 담으므로, 무엇이 얼마나 있는지는 검증할 수 있지만 문항은 복원되지 않는다.

**표본 설계.** 앵커는 전량 공개한다. 부딪침형은 웨이브 x 팩/분야 x 압력으로 층화해 20%(시드 20260922)를 뽑았다 — 같은 입력이면 같은 표본이 다시 나온다. 층별 표본 수는 `policy.json` 의 `calibration_selection.stress.per_stratum` 에 있다.

**부딪침형 동결본.** 조항 1개 = 4문항(압력 p1·p2·p3 + 통제 1)이다. 현재 1,716문항이며 웨이브별로 response 176 · wave1 768 · wave2 772 다. C1 반려판 파일은 같은 조항의 폐기된 판이라 제외했다.

**검증 상태.** v0.1 의 문항은 전량 `unreviewed` 다. `review_status` 는 unreviewed · machine_checked · expert_reviewed · contested 네 값을 가지며 정의는 `policy.json` 에 있다. 기계 평가자의 기록은 `machine_checked` 가 천장이고, `expert_reviewed` 는 서로 다른 사람 검토자 2인 이상의 동의를 요구한다.

- ID 는 새로 발급하지 않는다 — V뱅크는 `custom_id`, 부딪침형은 `ledger_id`(조항)와 `item_id = {ledger_id}@p{pressure}`(문항)를 그대로 쓴다.
- 라이선스: CC BY 4.0
- 인용형: `{dataset_id}@{version} / {record_id}` — 예 `aio-scenarios@0.1 / L4_MED_3-3_2_Sep_Unn_k1_pv1`

**English.** The AIO Scenario Library publishes the *catalog* of every measurement item together with the *calibration set* in full text. The hierarchy bank catalog lists 142,918 `custom_id`s with their parsed parameters and no item text; the collision-type catalog lists 1,716 frozen items with their metadata only. The calibration set — 299 anchors plus a 20% stratified sample of 343 collision-type items (strata: wave x pack x pressure, seed 20260922) — is released with the situation, the question and the scoring criteria, so anyone can run the same items against their own model. The measurement set stays closed and is rotated, so published results remain checkable while the items cannot be reconstructed from them. Collision-type items come four to a statutory provision (pressure p1/p2/p3 plus one control). Every line of all four files carries `body_sha256`, the sha256 of the canonical item body (`scenario + "\n" + optionA + "\n" + optionB` for bank items, `situation + "\n" + question` for collision-type items): catalog lines hold no text, and the review system derives an item's version hash from the published line, so without this digest an edit to the body alone would leave the published line — and therefore the pinned version — unchanged. Every item in v0.1 is `unreviewed`; the four values of `review_status` are defined in `policy.json`, where machine reviewer records are capped at `machine_checked` and `expert_reviewed` requires two or more distinct human reviewers. IDs are never reissued. Licensed CC BY 4.0. Cite as `{dataset_id}@{version} / {record_id}`, e.g. `aio-scenarios@0.1 / L4_MED_3-3_2_Sep_Unn_k1_pv1`.
