What an actual measurement produced
A package measuring one model end to end, in the same vocabulary and by the same formulas. Not just the summary figures — every individual response comes with it.
AI Measurement — Solar Pro 4 (2026-09-02)
One model at one point in time. The raw responses carry only item ids and the chosen option — no scenario text.
- Model
- upstage/solar-pro4
- Provider
- upstage
- Measured
- 2026-09-02
- Protocol
- v1.1
- Bundle id
- aef8ed33033a
- Standard
- AIO 20003
How much was asked
- Responses
- 142,918
- Swap responses
- 14,235
- Anchor
- 5,037
- Stress
- 2,648
- Context cells
- 946
- Rule verdicts
- 622
- Findings
- 231
- Edge observations
- 402
responses_k1.jsonl carries only the record ID (custom_id) and the choice. No scenario text is included, so downloading it does not reconstruct the measurement items.
Does it answer the same way twice
trr is agreement between measurements at different times, pcs agreement across five framings of the same condition, and pos_stable the share of answers that survive swapping the option positions. The key names are those of the source data.
How far it matched the reference, by domain
| Domain | Rules | Judged | No violation detected | Conditional | Unverified | Alignment rate |
|---|---|---|---|---|---|---|
| AGR | 25 | 18 | 15 | 3 | 7 | 83.3% |
| BASE | 43 | 0 | 0 | 0 | 43 | — |
| CIV | 33 | 24 | 16 | 8 | 9 | 66.7% |
| COM | 32 | 30 | 21 | 9 | 2 | 70.0% |
| CUL | 23 | 18 | 12 | 6 | 5 | 66.7% |
| DEF | 23 | 16 | 10 | 6 | 7 | 62.5% |
| EDU | 43 | 33 | 18 | 15 | 10 | 54.5% |
| ENE | 30 | 18 | 17 | 1 | 12 | 94.4% |
| ENV | 22 | 16 | 14 | 2 | 6 | 87.5% |
| GOV | 24 | 19 | 11 | 8 | 5 | 57.9% |
| HOU | 31 | 22 | 14 | 8 | 9 | 63.6% |
| IMM | 41 | 27 | 17 | 10 | 14 | 63.0% |
| INT | 2 | 0 | 0 | 0 | 2 | — |
| LAB | 44 | 30 | 19 | 11 | 14 | 63.3% |
| LAW | 36 | 22 | 16 | 6 | 14 | 72.7% |
| LND | 20 | 17 | 16 | 1 | 3 | 94.1% |
| MAC | 22 | 14 | 13 | 1 | 8 | 92.9% |
| MED | 40 | 28 | 18 | 10 | 12 | 64.3% |
| TEC | 23 | 19 | 11 | 8 | 4 | 57.9% |
| TRA | 22 | 19 | 18 | 1 | 3 | 94.7% |
| TRD | 23 | 17 | 17 | 0 | 6 | 100.0% |
| WEL | 33 | 23 | 15 | 8 | 10 | 65.2% |
“No violation detected” means no judgment contradicting the rule was observed — not a confirmation that the rule was followed. “Unverified” means the measurement contained no responses able to test that rule.
Formulas used aio-formulas@1.1 · Compared against kr-lnpd@0.9
Files
| File | Records | Size | Description |
|---|---|---|---|
| summary.json | — | 5.0 KB | Model, volumes, reliability and per-domain normative alignment |
| responses_k1.jsonl | 142,918 | 8.3 MB | Raw responses: custom_id and choice only, no scenario text |
| hierarchy_cells.jsonl | 946 | 986 KB | Per-cell value ranking, theta and tiers |
| c1_rule_verdicts.json | 622 | 274 KB | Per-rule verdicts |
| c1_item_observations.jsonl | 402 | 2.1 MB | Per-edge observations |
| findings.json | 231 | 1010 KB | Findings derived from the verdicts |
| README.md | — | 2.9 KB | Dataset description |
Citation form measurement-solar-pro-4-2026-09-02@1.0 / {record_id} · License CC-BY-4.0