The scenario library
The catalog of every item is published in full; the item text only for the calibration set. The measurement set stays closed and rotates — once published items enter training data, the next measurement tests recall rather than judgment. A few measurement-set items approved as observed cases are published as explicit exceptions (see the release rule below).
Source edition prod_v_scenarios(k1) · anchors_300 · 부딪침형 동결본
What the library is made of
Forced-choice items across value pairs, domains, and context cells — the body of the measurement set. The catalog carries their IDs and strata only.
Reference items that pin down where a judgment turns. Released in full, item text included.
Items that raise pressure step by step to see whether the same judgment holds. One statutory provision yields four items: one control plus p1, p2, p3.
| Wave | Items |
|---|---|
| Response form (first person) | 176 |
| Wave 1 (early international) | 768 |
| Wave 2 (national provisions) | 772 |
Excluded: stress_wave1_coe-huderia_v1.1_C1반려.jsonl · stress_wave2_BASE_v1_통제반려.jsonl — withdrawn editions of provisions that are already covered.
The calibration set and the catalog
The calibration set is 299 anchors plus a stratified sample of 343 collision-type items — 642 items in all. It carries the situation, the options, and the scoring criteria, so anyone can run the same items against their own model. The two catalog files carry no item text.
| File | Records | Size | Description |
|---|---|---|---|
| catalog_vbank.jsonl | 142,918 | 36 MB | Every hierarchy-bank custom_id with its parsed parameters; no item text. Each line carries body_sha256, the sha256 of scenario + \n + optionA + \n + optionB, so that editing the body alone still changes the published line |
| catalog_stress.jsonl | 1,716 | 666 KB | Metadata for every frozen collision-type item; no item text. Each line carries body_sha256, the sha256 of situation + \n + question |
| calibration_anchors.jsonl | 299 | 471 KB | Anchor calibration items, with text and the matching body_sha256 (recomputable from the text on the same line) |
| calibration_stress.jsonl | 343 | 817 KB | A 20% stratified calibration sample of collision-type items, with text and the matching body_sha256 |
| policy.json | — | 33 KB | Release rule, sampling design, the review_status field definition including the machine-record status ceiling, and the body_sha256 specification |
| README.md | — | 4.7 KB | Dataset description |
Citation form aio-scenarios@0.1 / {record_id} · License CC-BY-4.0
A 20% stratified sample was drawn from the 1,716 collision-type items. The strata are wave x pack_id x pressure and the seed is fixed at 20260922, so the same source data reproduces the same sample. Per-stratum counts are in policy.json.
IDs are never reissued
The source data already carries IDs used as join keys. They stitch together responses, bundles, and evidence files, so issuing new numbers would break every existing citation. Publication therefore leaves the IDs untouched.
- custom_id
- One hierarchy-bank item. e.g.
L4_MED_3-3_2_Sep_Unn_k1_pv1. The domain and cell codes are the parameters that generated the item, not labels applied later, so reclassification cannot break them.custom_id — {layer}_{domain}_{scale}-{reversibility}_{time}_{valueA}_{valueB}_{batch}_{perspective} - ledger_id · item_id
- For collision-type items the provision is the ledger_id and a single item is the item_id. e.g.
KR-MED#R-001@wave2-p2ledger_id (조항 단위) · item_id = {ledger_id}@p{pressure} (문항 단위) - rule_id
- One reference rule. e.g.
KR-EDU-T-101. Reference dataset →
The calibration set opens; the measurement set rotates
The calibration set is published in full, item text included, so that anyone can measure their own model on the same items and place the result beside ours.
The measurement set is not published and is rotated periodically. Raw responses carry only record IDs and choices, so results stay checkable while the items cannot be reconstructed.
If a published calibration item is suspected of having entered training data, it is withdrawn and replaced with a new item drawn from the same stratum.
The measurement set stays closed. The 3 items below are the exceptions: their original text and stored response were released, by operator approval (2026-09-30), for use as observed cases. Later measurement and analysis flag them as published items; observations and scores already recorded are neither deleted nor recomputed. No other measurement-set item text is published.
L4_MED_1-1_1_Sda_Sep_k1_pv1See the case →L4_TEC_4-1_2_Sdt_Ses_k1_pv1See the case →L4_LAW_1-1_2_Sdt_Fac_k1_pv1See the case →
Machine-readable list: /content/cases/disclosure_exceptions.json · It will be carried in the dataset's policy.json from the next release.
How far an item has been checked
Every item carries a verification state. In this release they are all unreviewed; the field fills in per record once item auditing begins.
| Value | Meaning |
|---|---|
| unreviewed | generated, not yet checked by machine or person |
| machine_checked | passed machine audit for form, duplication and statutory match |
| expert_reviewed | two or more independent reviewers agreed |
| contested | reviewer disagreement stands; excluded from headline figures |
The fields carried by one item
| Field | What it holds |
|---|---|
| Domain | one of the 22 domain codes (e.g. MED, health) |
| Value conflict | the two value codes in tension (e.g. Sep vs Unn) |
| Three condition axes | scale, reversibility, time — one of the 96 cells |
| Pressure level | how hard a collision-type item pushes (control, p1, p2, p3) |
| Norm reference | the reference-dataset rule ID the item targets |
| Verification state | unreviewed, machine_checked, expert_reviewed, contested |
| Generation | the `gen_ver` edition, raised whenever the item text changes |