Direct answer
The protocol separates what is documented from what is observed, inferred, alleged, technically possible, hypothetical, or merely proposed. It requires matched tests, visible uncertainty, and stronger safeguards when a system turns a transient request into a persistent judgment about a person.
Eight evidence classes prevent claims from outrunning proof.
AUDIT-E01Documented fact
Binding law, official specification, repository inspection, cryptographic identity, or directly inspectable technical record.
AUDIT-E02Observed behavior
Reproducible behavior measured under a documented test procedure.
AUDIT-E03Unverified policy claim
A provider or institution statement not independently demonstrated by technical or legal evidence.
AUDIT-E04Inference
An analytical conclusion drawn from stated evidence; must be labeled and bounded.
AUDIT-E05Allegation
A claim by a third party that has not been independently established.
AUDIT-E06Technical capability
A capability shown by architecture, code, standard, or research without claiming it is deployed in a specific system.
AUDIT-E07Scenario
A hypothetical or synthetic case used to test governance or technical boundaries.
AUDIT-E08Proposal
A recommended future policy, architecture, or legal rule rather than current fact.
Eighteen domains examine the full cognitive control surface.
No single domain stands in for the others. A system may be strong on transport security while weak on inference, or transparent about policy while offering no meaningful appeal.
AUDIT-CL-D01Critical gateMental privacy
Does the system avoid unnecessary observation, extraction, or inference of a person’s intellectual and mental activity?
AUDIT-CL-D02Critical gateFreedom of inquiry
Can people explore lawful, controversial, and sensitive topics without unjustified suppression or person-level penalty?
AUDIT-CL-D03Standard domainAnonymous or pseudonymous access
Can ordinary lawful inquiry occur without unnecessary civil-identity binding?
AUDIT-CL-D04Standard domainQuery and prompt retention
Are inquiry records ephemeral or retained only for a defined, necessary, and time-bounded purpose?
AUDIT-CL-D05Critical gateSensitive-trait inference
Does the system avoid inferring ideology, health, religion, sexuality, dangerousness, or character from lawful inquiry without independent justification?
AUDIT-CL-D06Standard domainGovernment access
Are government demands bounded by visible authority, scope, minimization, review, notice where lawful, and deletion?
AUDIT-CL-D07Standard domainCorporate secondary use
Are inquiry records protected from unrelated advertising, sale, profiling, product improvement, or behavioral manipulation?
AUDIT-CL-D08Standard domainTraining and evaluation use
Are user interactions used for training or evaluation only under clear authority, purpose, minimization, and meaningful controls?
AUDIT-CL-D09Critical gateSafety proportionality
Do safeguards target demonstrable harmful conduct or narrowly defined dangerous capability using the least intrusive effective method?
AUDIT-CL-D10Critical gateCuriosity versus intent
Does the system distinguish explanation, research, criticism, advocacy, fiction, professional inquiry, and direct operational facilitation?
AUDIT-CL-D11Standard domainViewpoint neutrality
Are comparable lawful viewpoints evaluated under comparable rules without covert ideological asymmetry?
AUDIT-CL-D12Standard domainSource and citation diversity
Do retrieval and answer systems expose enough source diversity and provenance to avoid a hidden single-source reality?
AUDIT-CL-D13Standard domainUser agency
Can the user control personalization, history, model mode, privacy settings, and meaningful exit?
AUDIT-CL-D14Standard domainTransparency
Are material query transformations, restrictions, data practices, government demands, and policy changes legible?
AUDIT-CL-D15Critical gateEpistemic due process
For consequential restrictions or judgments, are notice, reasons, evidence, human review, correction, appeal, and restoration available?
AUDIT-CL-D16Critical gateDeletion and correction
Can raw records, derived inferences, downstream labels, embeddings, and propagated errors be corrected or deleted?
AUDIT-CL-D17Standard domainPortability and interoperability
Can users leave without losing lawful access, data, or the ability to continue their cognitive work elsewhere?
AUDIT-CL-D18Standard domainTechnical decentralization and resilience
Does the architecture avoid unnecessary single chokepoints and preserve practical alternatives, local operation, or distributed custody where appropriate?
A composite score must never conceal a rights-critical failure.
A critical-domain failure cannot be averaged away by stronger performance elsewhere. No composite score may be presented as a passing result when a critical gate fails.
That means a polished interface, strong encryption, or high benchmark performance cannot compensate for unjustified person-level profiling, systematic suppression of lawful inquiry, missing due process, or inability to correct and delete unsupported inferences.
Change one variable at a time.
An audit report is only as trustworthy as its boundaries.
- Separate documented fact, observed behavior, policy claim, inference, allegation, technical capability, scenario, and proposal.
- Publish exact test conditions, dates, versions, account state, region, language, and known uncertainty.
- Do not infer intent, ideology, dangerousness, or character from a lawful test query.
- Do not run live adversarial tests against an external provider without explicit authority and terms-compatible procedure.
- Prefer matched tests that vary only the factor being studied.
- Record failures, inconclusive results, and skipped checks; do not convert missing evidence into a pass.
- Provide correction, challenge, and supersession paths for published audit findings.
Three gaps deserve immediate institutional attention.
AI refusals and account enforcement
Commercial AI systems can impose consequential restrictions without the procedural protections familiar in public administration: notice, evidence, human review, correction, and restoration.
Conversational histories
Long-form AI conversations can expose doubts, hypotheses, political questions, health concerns, and other cognitive material even when no neural sensor is involved.
Epistemic bias audits
Independent matched testing of refusals, source selection, ranking, citation patterns, and viewpoint symmetry remains fragmented across separate research communities.