What the AIO Framework is
The AIO Framework is a governance framework for AI integrity. It is the body of standards and tools that lets an organization declare the criteria its AI should judge by, then record, compare, and correct whether the AI actually judges that way.
The AIO Framework is not a single tool. It is a system: four standards — set, log, analyze, apply — resting on a shared vocabulary and a shared procedural base. Each tool is useful on its own, but only together do they answer one question: "Is this AI actually judging by the criteria it declared?"
1. Why a governance framework
AI governance has so far existed in roughly three forms.
| Form | What it does | What it leaves open |
|---|---|---|
| Ethics charters, principle statements | States intent in prose | No way to compare the statement against actual judgments |
| Safety evaluation, red teaming | Finds dangerous outputs | Never asks what criteria produced the non-dangerous ones |
| Compliance checklists | Confirms documents and procedures exist | Says nothing about what to record, or in what vocabulary |
None of the three reaches the content of the judgment. The AIO Framework fills that gap. It does not rule on whether an AI is right; it makes the distance between the declared criteria and the actual judgment measurable.
And AI is not the only object of governance. People and organizations are what decide what comes first, so the first thing the framework measures is the organization itself. The prompt injected into the AI is a derivative of that.
2. A first-principles approach — three starting points
The AIO Framework does not begin with a new list of rules. It begins again from what a judgment is.
First, every judgment is a hierarchy. Any judgment reduces to what was placed ahead of what. That ordering is observable along three axes: values (V), evidence (E), and sources (S). The minimal unit of a judgment is not a score but an order.
Second, a hierarchy can be declared, recorded, and compared. Write the hierarchy in a code that humans read and machines emit, and the declaration (setting) and the reality (log) sit in the same language. What shares a language can be compared; what can be compared can be corrected. This is where governance starts to work.
Third, whoever writes the standard does not judge compliance with it. AIO supplies methods and tools; verification belongs to independent auditors. A standard whose author also certifies against it cannot be trusted. That principle is why the final step of the loop is apply, not certify.
3. What the framework is made of — the integrity loop
Set (20001) → Log (20002) → Analyze (20003 · 20023) → Apply (20004) → (re-set)
Each of the four standards is a tool in its own right and a joint in the loop. Without the earlier step, the later one has nothing to compare against.
Set — AIO 20001
An organization fixes its value hierarchy (direction) and its red lines (limits) in a single working session, and exports the result as machine-readable code.
What is different — a conventional AI ethics charter ends at sentences like "we respect fairness." Sentences do not say what wins when they collide. Instead of abstract declarations, 20001 leaves an order obtained by forced choice and red lines that were explicitly declared — and that becomes a reference value something can be checked against.
Log — AIO 20002
Each judgment leaves one line of hierarchy code: context (C), values (V), evidence (E), sources (S), and the ordering among them.
What is different — a conventional audit log records what was done. 20002 records what was put first. It applies at the prompt layer without swapping models or opening them up, and an axis that was never engaged is declared honestly as U (unspecified) rather than left blank.
Analyze — AIO 20003 · 20023
Two branches. 20003 (experimental evaluation) measures a subject's actual disposition through controlled scenarios; 20023 (field evaluation) compares operational logs against the organization's declared setting.
What is different — conventional benchmarks measure accuracy and rank models. What is measured here is not correctness but disposition and coherence. This step also diagnoses integrity hallucination — judgments that wobble across structurally identical situations — by decomposing it into reproducibility (TRR) and perspective consistency (PCS).
Apply — AIO 20004
Discrepancies are returned to the system and to the organization. Combining the setting and the record standard into a single system prompt is stage one of that.
What is different — most governance procedures end at a report or a certificate. If the loop never closes, the next judgment is no different from the last. 20004 is the step that pushes results back into the setting — and, because of the independence principle above, it is not certification.
Shared base — AIO 00010–00014
Everything on the loop rests on the same base: 00010 terminology, 00011 vocabulary, 00012 the public RFC process, 00013 numbering, and 00014 scenario generation. Scenario generation in particular is shared infrastructure between the setting step and experimental evaluation.
4. Why the vocabulary comes first
Governance begins with language. If the same problem is described in different words, neither agreement nor audit is possible. So the framework fixes vocabulary before it fixes rules.
- C — context: 22 domains × scope · reversibility · time
- V — 19 values: Schwartz value theory (2012 refinement)
- E — 10 evidence types: Walton's argumentation schemes + the GRADE/CEBM evidence hierarchy
- S — 10 source types: Hovland–Kelley source credibility theory
This is not an arbitrary taxonomy but a vocabulary wired into existing scholarly traditions — a new language built on verified roots.
The three axes are not independent; they cascade. Values constrain evidence standards, evidence standards constrain source preferences, and source preferences ultimately decide which data is taken up. A value system that coherently tunes the criteria beneath it is legitimate cascading; distorting facts and data without declaring it is authority pollution. Only by observing the three axes separately can you say at which layer the pollution occurred.
5. What this makes possible
| For whom | What becomes possible |
|---|---|
| Organizations | Put the declared criteria into the running system and see the gaps, with evidence |
| Auditors and regulators | Submit and review judgment paths in a vendor-independent form (EU AI Act Art. 13–14 interpretability layer) |
| Researchers | Compare and reproduce value / evidence / source structures across models |
| The public | See, in published form, what a given AI puts first |
6. Why now
The wider the range of decisions AI takes part in, the more a society needs a language for the path those decisions took. If the path is not recorded, responsibility cannot be divided, disagreement cannot be adjudicated, and consensus cannot accumulate.
The AIO Framework is not trying to make one value system win. It is trying to establish, first, the conditions under which arguments about values can be fair. Whether an AI-based society moves in a healthy direction depends, in the end, on whether that society can check that it wrote down its own criteria and kept to them. This framework defines the minimal unit of that check — one line of record.
7. Where this stands today (honesty notice)
- Measurement status: base (7 domains · 10 values · 366,120 responses, validated) / Extended (22 domains · 19 values, in progress).
- Measurement used the Schwartz basic 10; the vocabulary standard is the refined 19. Measurement at the finer level is planned.
- 20002 is positioned as an EU AI Act Art. 13–14 interpretability layer. No claim is made that it satisfies Art. 12 on its own.
- The 20004 standard document is not final; it will be settled through the public RFC process (00012).
- Publishing the limits of the standard, the data, and the method is one of the framework's own principles. Not hiding them is the basis of trust.