← Back to the Overview
Deep Dive · AI Incident Reporting

The Evidence Map for AI Incident Reporting

The action layer behind the verdict: how to stand up an AI incident regime inside existing incident command without overstating what law, standards, or public incident databases actually support.

Source  Storm Research Verification  20 citations checked Prepared for  Leader discussion
How to use this

Use this before an AI system with real-world consequence goes live, and before a regulator asks what your AI incident process is. The point is not a new bureaucracy. Mature regimes in aviation, grid reliability, securities disclosure, and drug safety already show the pattern: defined thresholds, named recipients, staged reports, follow-ups, and a protected learning lane. The work is wiring AI into the incident machinery you already trust, then rehearsing it.

The findings in full

9/10
Define the event by harm plus AI contribution, and keep near misses distinct.

The cleanest verified base definition: development, use, or malfunction of AI directly or indirectly leads to harm to health, critical-infrastructure disruption, violations of protected rights, or harm to property, communities, or the environment. The EU AI Act adds a serious-incident ceiling with a causal gate. Enterprise practice should add two categories law often misses: a near miss where harm was narrowly prevented, and a control failure where a stop, override, authorization, monitoring, data, or evidence control did not work. The deliberate consequence: a bad model output with no harm and no failed control stays in the quality backlog.

9/10
Build two lanes: protected learning and mandatory accountability.

Aviation's confidential reporting system demonstrates that voluntary, protected reporting surfaces near misses a punitive channel suppresses; its sanction waiver is conditional and it never replaces mandatory duties. Grid, securities, and drug-safety regimes supply the accountability lane: defined recipients, triggering tests, deadlines, preliminary reports, and follow-ups. The enterprise version protects good-faith internal reports while preserving discipline for recklessness, deliberate misconduct, concealment, or failure to make a required external report.

9/10
Inside a regulated utility, reuse incident command and add an AI evidence layer.

The worked example, current as of July 2026: CIP-008-6 and EOP-004-4 are the mandatory NERC baselines. An AI-related cyber event meeting the CIP-008-6 criteria goes to E-ISAC and CISA on the existing one-hour clock after determination; a reliability event follows EOP-004-4 and the entity's Operating Plan; the DOE form uses its own faster tiered clocks through its own channel, and submitting it does not automatically notify every other required recipient. One internal case drives every route. Successor standard versions exist but were not yet in force at verification.

8/10
Trigger on outcome, failed control, scale, and irreversibility, not one dollar threshold.

Internal reportability starts on a reasonable possibility that AI contributed to harm, a near miss, an unauthorized action, a failed stop or override, compromised data or model integrity, loss of evidence, or a recurring pattern. External reportability starts when the underlying event meets the applicable safety, reliability, cyber, privacy, employment, environmental, securities, or AI-law threshold. The two triggers are intentionally different, and related low-severity events must be aggregated: a series can be material when no single event is.

8/10
Make the preserved evidence packet part of the definition.

Before remediation changes the scene, preserve the system and model version, use case, inputs and retrieved context, prompts, tool calls, permissions, human approvals and overrides, affected assets and people, expected controls, observed failure, containment, vendor involvement, and causal confidence. Then report in stages: fast initial facts, updates as the picture changes, a root-cause and corrective-action record, and a trend review that catches repeated low-level failures across systems. An event that cannot be reconstructed cannot support learning, disclosure, or assurance.

First moves before hiring anyone

01
Put the three-part taxonomy in policy.

Incident means realized harm with material AI contribution. Near miss means harm prevented by intervention, redundancy, or chance. Control failure means a failed safeguard that could plausibly enable harm. Keep routine quality defects in the product backlog unless they aggregate, recur, or expose a failed control.

02
Adopt six severity gates, any one of which escalates.

People or environmental safety; critical-service or operational disruption; cyber, data, or model integrity; legal, privacy, employment, or rights impact; financial or investor materiality; unauthorized autonomy, failed stop, irreversibility, or cross-site propagation.

03
Add an AI field set to the incident system you already run.

System and model version, use case, owner, vendor, inputs and context, tools and actions, permissions, approvals, override state, affected assets and stakeholders, expected and failed controls, evidence-preservation status, containment, causal confidence, and every applicable reporting clock. A field set is a sprint; a parallel AI portal is a mistake that splits ownership and clocks.

04
Set internal clocks that beat the fastest external clock.

Immediate operational escalation; a named incident commander and legal triage within one hour for severe events; a preliminary evidence packet within four hours; an initial enterprise record within 24 hours. These are internal targets. The external clock stays whatever the applicable reliability, cyber, disclosure, privacy, safety, or AI law requires.

05
Build one routing matrix by use case, entity, jurisdiction, and harm type.

For a NERC-registered utility that means mapping E-ISAC, CISA, the reliability regulator and Regional Entity, the Reliability Coordinator, the DOE Operations Center, the state commission, law enforcement, securities disclosure, and privacy authorities against each AI use case. Other regulated industries substitute their own recipients; the matrix discipline is the same.

06
Exercise the regime and measure it.

Run one tabletop with three scenarios: an unsafe AI recommendation rejected by an operator, an agent acting outside its authorization, and a vendor model compromise. Measure time to classify, evidence completeness, clock identification, stop authority, recipient accuracy, update quality, corrective-action closure, and recurrence.

Owner, briefing, proof

Owner

A named incident commander per event, plus standing owners for classification, evidence preservation, and each external reporting route. The AI label never moves ownership out of incident command.

Briefing

A one-page read per event: class, severity gates crossed, lane, applicable external thresholds and clocks, and the current causal confidence, updated as the picture changes.

Proof

The preserved evidence packet plus the staged record: initial facts, updates, root cause, corrective action, and the trend review that catches repeated low-level failures.

Where to start

Start with one tabletop: run an unsafe AI recommendation, an out-of-authorization agent action, and a vendor model compromise through the incident command you already have, and score time to classify, evidence completeness, and recipient accuracy. If the gaps are material, widen to a definition, trigger, and routing review across the AI portfolio. Build the full regime, protected lane included, only when the sponsor wants it run.

Claim ledger

20
Checked
citations traced to primary or canonical sources in four independent clusters on July 23, 2026
0
Fabricated
no invented or unsupported sources surfaced in verification
8
Corrected
standard versions, clocks, and protection scopes narrowed after source review
3
Demoted
database counts and policy advocacy reclassified below primary evidence
ConfirmedOECD AI Incidents Monitor methodology: supplies the base incident and hazard definitions and four harm classes used here.oecd.ai
ConfirmedEU AI Act, Articles 3(49) and 73: defines serious incident, the causal gate, recipients, staged deadlines, and permitted incomplete initial reports.eur-lex.europa.eu
CorrectedEU Digital Omnibus: the adopted text resets high-risk application dates but awaited Official Journal publication at verification; Article 73 timing stays in motion, and the Commission's serious-incident guidance remained draft.eur-lex.europa.eu
ConfirmedNIST AI Risk Management Framework: voluntary support for incident response, recovery, communication, and third-party contingency planning.airc.nist.gov
CorrectedNIST Generative AI Profile: calls for tracking errors and near misses and rehearsing response; it does not require public near-miss sharing.nvlpubs.nist.gov
ConfirmedNERC CIP-008-6: the current mandatory cyber-incident standard in July 2026; one-hour notification to E-ISAC and CISA once an incident is determined reportable, next-day for attempted compromises.nerc.com
ConfirmedNERC EOP-004-4: the current event-reporting baseline, via entity Operating Plans and the later of 24 hours or the next business day.nerc.com
CorrectedNERC version timeline: CIP-008-7.1 is approved but generally effective July 2028; EOP-004-5 was filed but not FERC-approved at the briefing date. Advice built on the successors misstates the current obligation.nerc.com
CorrectedDOE Form OE-417: a separate channel with one-hour and six-hour tiered clocks; form acceptance does not deliver the report to every other required recipient.doe417.energy.gov
CorrectedFAA advisory circular and NASA reporting system: confidential voluntary reporting with a conditional sanction waiver; not blanket immunity and not a substitute for mandatory reports.faa.gov
ConfirmedSEC Form 8-K Item 1.05: material cyber incidents generally due four business days after a materiality determination that must not be unreasonably delayed; qualitative and aggregated effects count.sec.gov
CorrectedFDA postmarketing rule, 21 CFR 314.80: serious and unexpected events use a 15-day alert and follow-up model; the seven-day clock belongs to a separate regime.ecfr.gov
ConfirmedWei and Heim, AAAI 2026: seven design dimensions across nine purposive U.S. reporting regimes; supports contingent design, not one proven universal architecture.ojs.aaai.org
ConfirmedLi et al., AIES 2025: manual study of 499 public reports; 301 involved use-related failures, many affecting people who never touched the system. Public-reporting bias limits prevalence claims.ojs.aaai.org
DemotedStanford AI Index 2026: reports 362 recorded incidents in 2025 versus 233 in 2024, from a voluntary, media-driven database with no exposure denominator; a documentation count, not an incidence census.hai.stanford.edu
DemotedPolicy and industry frameworks: the critical-infrastructure policy case for reusing existing ISACs and the real-time agent failure-detection framework are useful design inputs; advocacy and voluntary norms, not law or measured causal evidence.cset.georgetown.edu
Where the evidence stops

The verified base establishes the definitions, the current mandatory clocks, and the design pattern mature regimes share. Two things it does not establish. First, public incident databases measure documented reports, not the AI incident rate: the counts are voluntary, media-dependent, skewed toward high-visibility English-language cases, and lack any exposure denominator, so a rising count proves accumulation of documented cases, not a measured rise in harm. Second, the two-lane architecture rests on cross-regime analogy and a peer-reviewed design review, not on controlled evidence that any single reporting design reduces incidents. There is also an open institutional question the briefing names honestly: no one has yet built the protected cross-firm sharing layer that would let one enterprise's incident prevent another's.

What would change our mind
Deep Dive staged from verified Storm Research · nothing here asserts above the registry calibration