The Evidence Map for AI Incident Reporting
The action layer behind the verdict: how to stand up an AI incident regime inside existing incident command without overstating what law, standards, or public incident databases actually support.
Use this before an AI system with real-world consequence goes live, and before a regulator asks what your AI incident process is. The point is not a new bureaucracy. Mature regimes in aviation, grid reliability, securities disclosure, and drug safety already show the pattern: defined thresholds, named recipients, staged reports, follow-ups, and a protected learning lane. The work is wiring AI into the incident machinery you already trust, then rehearsing it.
The findings in full
The cleanest verified base definition: development, use, or malfunction of AI directly or indirectly leads to harm to health, critical-infrastructure disruption, violations of protected rights, or harm to property, communities, or the environment. The EU AI Act adds a serious-incident ceiling with a causal gate. Enterprise practice should add two categories law often misses: a near miss where harm was narrowly prevented, and a control failure where a stop, override, authorization, monitoring, data, or evidence control did not work. The deliberate consequence: a bad model output with no harm and no failed control stays in the quality backlog.
Aviation's confidential reporting system demonstrates that voluntary, protected reporting surfaces near misses a punitive channel suppresses; its sanction waiver is conditional and it never replaces mandatory duties. Grid, securities, and drug-safety regimes supply the accountability lane: defined recipients, triggering tests, deadlines, preliminary reports, and follow-ups. The enterprise version protects good-faith internal reports while preserving discipline for recklessness, deliberate misconduct, concealment, or failure to make a required external report.
The worked example, current as of July 2026: CIP-008-6 and EOP-004-4 are the mandatory NERC baselines. An AI-related cyber event meeting the CIP-008-6 criteria goes to E-ISAC and CISA on the existing one-hour clock after determination; a reliability event follows EOP-004-4 and the entity's Operating Plan; the DOE form uses its own faster tiered clocks through its own channel, and submitting it does not automatically notify every other required recipient. One internal case drives every route. Successor standard versions exist but were not yet in force at verification.
Internal reportability starts on a reasonable possibility that AI contributed to harm, a near miss, an unauthorized action, a failed stop or override, compromised data or model integrity, loss of evidence, or a recurring pattern. External reportability starts when the underlying event meets the applicable safety, reliability, cyber, privacy, employment, environmental, securities, or AI-law threshold. The two triggers are intentionally different, and related low-severity events must be aggregated: a series can be material when no single event is.
Before remediation changes the scene, preserve the system and model version, use case, inputs and retrieved context, prompts, tool calls, permissions, human approvals and overrides, affected assets and people, expected controls, observed failure, containment, vendor involvement, and causal confidence. Then report in stages: fast initial facts, updates as the picture changes, a root-cause and corrective-action record, and a trend review that catches repeated low-level failures across systems. An event that cannot be reconstructed cannot support learning, disclosure, or assurance.
First moves before hiring anyone
Incident means realized harm with material AI contribution. Near miss means harm prevented by intervention, redundancy, or chance. Control failure means a failed safeguard that could plausibly enable harm. Keep routine quality defects in the product backlog unless they aggregate, recur, or expose a failed control.
People or environmental safety; critical-service or operational disruption; cyber, data, or model integrity; legal, privacy, employment, or rights impact; financial or investor materiality; unauthorized autonomy, failed stop, irreversibility, or cross-site propagation.
System and model version, use case, owner, vendor, inputs and context, tools and actions, permissions, approvals, override state, affected assets and stakeholders, expected and failed controls, evidence-preservation status, containment, causal confidence, and every applicable reporting clock. A field set is a sprint; a parallel AI portal is a mistake that splits ownership and clocks.
Immediate operational escalation; a named incident commander and legal triage within one hour for severe events; a preliminary evidence packet within four hours; an initial enterprise record within 24 hours. These are internal targets. The external clock stays whatever the applicable reliability, cyber, disclosure, privacy, safety, or AI law requires.
For a NERC-registered utility that means mapping E-ISAC, CISA, the reliability regulator and Regional Entity, the Reliability Coordinator, the DOE Operations Center, the state commission, law enforcement, securities disclosure, and privacy authorities against each AI use case. Other regulated industries substitute their own recipients; the matrix discipline is the same.
Run one tabletop with three scenarios: an unsafe AI recommendation rejected by an operator, an agent acting outside its authorization, and a vendor model compromise. Measure time to classify, evidence completeness, clock identification, stop authority, recipient accuracy, update quality, corrective-action closure, and recurrence.
Owner, briefing, proof
Owner
A named incident commander per event, plus standing owners for classification, evidence preservation, and each external reporting route. The AI label never moves ownership out of incident command.
Briefing
A one-page read per event: class, severity gates crossed, lane, applicable external thresholds and clocks, and the current causal confidence, updated as the picture changes.
Proof
The preserved evidence packet plus the staged record: initial facts, updates, root cause, corrective action, and the trend review that catches repeated low-level failures.
Start with one tabletop: run an unsafe AI recommendation, an out-of-authorization agent action, and a vendor model compromise through the incident command you already have, and score time to classify, evidence completeness, and recipient accuracy. If the gaps are material, widen to a definition, trigger, and routing review across the AI portfolio. Build the full regime, protected lane included, only when the sponsor wants it run.
Claim ledger
The verified base establishes the definitions, the current mandatory clocks, and the design pattern mature regimes share. Two things it does not establish. First, public incident databases measure documented reports, not the AI incident rate: the counts are voluntary, media-dependent, skewed toward high-visibility English-language cases, and lack any exposure denominator, so a rising count proves accumulation of documented cases, not a measured rise in harm. Second, the two-lane architecture rests on cross-regime analogy and a peer-reviewed design review, not on controlled evidence that any single reporting design reduces incidents. There is also an open institutional question the briefing names honestly: no one has yet built the protected cross-firm sharing layer that would let one enterprise's incident prevent another's.
- The EU Digital Omnibus publishes in the Official Journal, or the European Commission finalizes its serious-incident guidance and reporting template, changing Article 73 timing or content.
- FERC approves the successor event-reporting standard, NERC changes an implementation date, or the next cyber-incident standard version is accelerated; the routing example then needs re-verification.
- DOE replaces the current disturbance-reporting form or resolves its known wording inconsistencies.
- A standards body, regulator, or information-sharing organization publishes a common AI incident taxonomy, severity scale, or protected sharing protocol.
- A regulator or court decides how much AI contribution triggers causation, materiality, disclosure, or liability.
- A controlled or audited multi-enterprise study shows protected near-miss reporting reduces recurrence, or that reporting burden overwhelms signal.
- A serious AI-related utility or industrial incident exposes missing evidence, failed routing, or an ineffective human override in practice.