← Back to the Overview
Deep Dive · Mission Operations

The Authority Map for AI in Mission Operations

The action layer behind the core verdict: how to give AI a bounded, preauthorized job in the loop without granting authority the assurance case cannot carry.

Source  Storm Research v2 Verification  17 citations checked Prepared for  Leader discussion
How to use this

Use this as a screen for any AI claim in mission operations, flight or ground. The question is never whether the model is impressive. The question is what authority the function holds: what it may decide, what it may command, who owns the exception, which deterministic safeguard bounds it, and how the operation reverses it. Authority design comes first; model selection is downstream. Every deployed case in the ledger below earned its place that way.

First moves before any AI function holds authority

01
Build the authority map for one system first.

For each proposed AI function, write down the input, the decision, the permitted action, the prohibited action, the escalation owner, the deterministic safeguard, and the reversal path. Separate detection, recommendation, scheduling, and command authority; most autonomy arguments dissolve once those four are separated.

02
Point it at high-volume, reversible, constraint-rich loops.

Telemetry prioritization, science-data triage, contact and activity scheduling, and nominal health monitoring match the strongest deployed evidence: AEGIS target selection, the Perseverance onboard planner, TECO scheduling, and console anomaly alerts. Leave novel, coupled, or hazardous conditions on the human side of the map.

03
Make off-nominal behavior the acceptance test.

Score the function on false alarms, missing data, configuration drift, conflicting constraints, degraded communications, operator override, and reconstruction after an event. Nominal accuracy alone proves a demo, not an operating capability.

04
Treat the console workflow and configuration baseline as part of the product.

Require requirements traceability, versioned data and model inputs, operations-like simulation, shadow operation, named shift procedures, and a rollback plan before changing authority or staffing. Hubble ran eight months of 24x7 shadow operations and anomaly simulations before cutting to 8x5.

05
Reserve crew-critical and hazardous authority for a dedicated assurance case.

For any function that can cause, or fail to prevent, a catastrophic hazard or an abort, start from the applicable human-rating and safety requirements: fault tolerance, recovery, situational awareness, configuration control, crew override. A generic AI governance review is not a substitute, and a model score is not a safety case.

Owner, briefing, proof

Owner

A named operating owner for each AI function's authority envelope, accountable for the escalation path, the deterministic safeguards, the shadow-operation record, and the decision to suspend or roll back.

Briefing

A routine-use / demonstration / claim decision brief per capability, so a one-mission demo or a contractor-reported metric never buys the authority that only routine, verified operational use has earned.

Proof

The trail from telemetry to alert to decision to authorized action to verified outcome, plus off-nominal test results and shadow-operation records, rather than model accuracy scores.

Where to start

Start with owner, briefing, and proof for one system and one play. If the gap is material, widen to a readiness look across the operation's monitoring, scheduling, and command loops, and build the operating machinery only when the operator wants it run.

Claim ledger

17/17
Checked
citations traced to public primary sources
0
Fabricated
no invented figures found
4
Corrected
qualified after source review
3
Demoted
preprints and a prototype kept out of the headline
ConfirmedNASA human-rating guidance (NASA-STD-3001): current public guidance ties crewed operations to fault tolerance, recovery, situational awareness, configuration control, and crew control of critical functions, with a 1-in-500 maximum allowable loss-of-crew probability for ascent or descent.nasa.gov
CorrectedNPR 8705.2B (2012): the detailed crew-control, override, and critical-software requirements are real, but the document is superseded. Used as explanatory history only, never for formal compliance.nodis3.gsfc.nasa.gov
ConfirmedJPL AEGIS: autonomous science-target selection against scientist-set parameters; first on Opportunity in 2010, in routine Curiosity use since 2016, on Perseverance since 2022.ai.jpl.nasa.gov
ConfirmedJPL M2020 onboard scheduler: Perseverance's onboard planner became the primary operations method on 5 October 2023; 429 sols and more than 7,800 requested activities executed by 29 January 2025.ai.jpl.nasa.gov
ConfirmedParjan and Gaines (IEEE 2024): the planner shipped through a requirements-driven verification campaign, 154 requirements across 32 verification activities, including operations-like, stress, and off-nominal testing.ai.jpl.nasa.gov
DemotedSoderstrom et al., LSTM anomaly detection (2018): a preprint on expert-labeled SMAP and Curiosity telemetry, framed as reducing engineer monitoring burden. Supports monitoring assistance, never autonomous recovery authority.ntrs.nasa.gov
ConfirmedMukai et al., MSL telecom anomaly detection (2020): used in daily operations with a reported 90% team-workload reduction; human review and hard safety thresholds stay in the decision path.ntrs.nasa.gov
CorrectedIverson et al., ISS AMISS (2010): data-driven monitoring integrated into flight-controller consoles with alerts and certification; the source does not support a 2012 certification date.ntrs.nasa.gov
ConfirmedESA TECO scheduler: a reliable service since 2013; accepts requests from four payload-control centers, applies hundreds of constraints to weekly schedules of 100-plus actions, sometimes 500-plus, and escalates anomalous conflicts to humans.esa.int
DemotedBresina et al., MAPGEN (2005): historical mixed-initiative planning evidence, automated constraint-based planning with human direction and evaluation, not a current performance result.ntrs.nasa.gov
CorrectedJPL ST6 Autonomous Sciencecraft (2005): the reported 116x downlink savings is agency program validation, not independent human-rating assurance.jpl.nasa.gov
ConfirmedBurley et al., Hubble operations automation (2012): the 24x7-to-8x5 staffing transition followed an eight-month shadow period, anomaly simulations, and operating-model changes across planning, ground systems, staffing, and responsibilities.ntrs.nasa.gov
DemotedESA KETTY telemetry prototype: produced too many false alarms on real XMM telemetry at TRL 2. A separate contractor-reported OPS-SAT result publishes strong test metrics but is not independent or peer-reviewed (contested, confidence 4/10); neither establishes ML recovery authority.esa.int
CorrectedAPL Multimission Operations Center: prelaunch simulation and separately described staffed and unattended modes are supported; the unquantified reliability and savings language remains an APL claim.jhuapl.edu
Where the evidence stops

Three boundaries hold after verification. Telemetry-model accuracy does not establish autonomous diagnosis or recovery authority: the KETTY prototype produced too many false alarms on real telemetry, and the strongest recent metrics are contractor-reported rather than independent. Workload, staffing, and savings figures, including the 90% number, are agency program-reported results, not generalizable ROI. And the most detailed human-rating language comes from a superseded document; formal compliance work starts from the current authorities.

What would change our mind
Deep Dive staged from verified Storm Research v2 · nothing here asserts above the registry calibration