← Back to the Overview
Deep Dive · Employee-Built Agents

The Evidence Map for Employee-Built Agent Sprawl

The action layer behind the core verdict: how to discover, promote, and retire employee-built agents without overstating what the growth telemetry, the control planes, or the governance analogies actually support.

Source  Storm Research v2 Verification  29 citations checked Prepared for  Leader discussion
How to use this

Use this when employees are packaging assistants and agents faster than anyone can review them. The failure point is not creation; it is the moment a useful personal tool acquires shared users, sensitive data, write authority, or business reliance without a named owner and runtime evidence. The goal is not to stop the building. It is to make the visible path faster, safer, and more reusable than the invisible one.

The findings at full depth

9/10
Governance follows operational capability and risk, not the employee's job title.

Risk-management frameworks call for risk-proportionate inventory, accountable roles, testing, monitoring, incident response, and decommissioning across the lifecycle, and binding regulation adds oversight and log-retention duties for covered high-risk deployers. Inventory and basic guardrails start before production; stronger controls attach as data sensitivity, audience, tools, permissions, autonomy, reversibility, and reliance rise.

8/10
Agent-like use is growing quickly, but the autonomous shadow-agent count is unknown.

One platform reports 15x year-over-year growth in active agents; another reports roughly 19x growth in weekly users of reusable AI configurations and one bank running more than 4,000 of them. All are vendor telemetry without absolute counts or definitions. A 360-leader survey found broad agent activity but no more than 15% even considering fully autonomous agents. Creation, active use, system connection, write authority, and autonomy are separate adoption stages; report the ladder, not one number.

7/10
A federated sandbox-to-production path is operationally credible, but not proven superior.

Government RPA and citizen-developer programs, and the older end-user-computing control record, support broad low-risk creation with stronger gates for AI-driven automation, mission-critical use, and advanced connectors. No reviewed source compares this design's outcomes with blanket prohibition or universal central approval, so treat it as a design hypothesis to pilot and measure.

7/10
Major vendors ship real control planes, but none is automatically complete.

Registry, agent identity, policy gateway, observability, evaluation, and lifecycle controls now exist across multiple major platforms, several of them cross-platform. Every reviewed surface depends on supported connectors or scanners, credentials, registration, instrumentation, gateway routing, managed status, source-platform APIs, or preview features. Some inline safety layers inspect the first prompt and final response but not intermediate agent steps.

7/10
The durable ownership split separates enablement, independent challenge, and business accountability.

Current lifecycle documentation separates the accountable owner from independent risk review, supporting a three-way split: platform team on shared infrastructure and evidence, risk and compliance on policy and block authority, and a business or product owner on purpose, output acceptance, value, and retirement. This is a supported operating-model inference, not a proven universal standard.

First moves before hiring anyone

01
Define the inventory object before buying a control tower.

Require agent ID, creator, accountable owner, purpose, users, environment, data sources, tools, runtime identity, permissions, model and version, cost center, risk tier, last activity, last review, incident contact, and retirement state. Reconcile vendor registries with identity, gateway, endpoint, API, expense, and owner-attestation data.

02
Use a four-tier promotion path with explicit authority limits.

Personal sandbox with low-risk data and no production actions; internal read and draft with automated controls and a named owner; write or action with human approval, integration tests, rollback, and security review; narrow autonomous or high-impact use with formal risk review, independent evaluation, a kill path, an incident plan, and periodic recertification.

03
Make registration automatic and promotion service-level-bound.

Never ask makers to type metadata the platform can infer. Give low-risk reviews a short service target, publish approved connectors and templates, and let teams search for reusable agents before building another one. Track approval lead time and workaround rate as governance-product metrics.

04
Separate the three owners in writing.

The platform team owns the paved road, identity, shared controls, telemetry, evidence collection, and technical support. Risk and compliance own classification, policy, exceptions, independent validation, and block authority. The business or product owner owns the use case, output acceptance, value, user readiness, and retirement authorization.

05
Govern the run, not only the artifact.

For agents with action authority, capture who delegated what, which identity acted, data and tools touched, policy checks, approvals, outputs, side effects, cost, exceptions, and recovery. Prompt screening alone is not a substitute for authorization and intermediate-step visibility.

06
Put retirement on the same dashboard as adoption.

Review agents when the owner leaves, usage is absent for 90 days, the value threshold is missed, a duplicate exists, a connector or model changes materially, permissions expand, or an incident occurs. Revoke credentials, preserve required records, transfer dependencies, and close the inventory record.

Owner, briefing, proof

Owner

One accountable operational owner per promoted agent, plus independent block authority that sits outside the owning team. Orphaned agents with live permissions are the signature failure.

Briefing

The six-stage fleet count: created, active, shared, system-connected, write-capable, autonomous. Plus each agent's tier, what it cannot do, and which move needs a new approval.

Proof

Run evidence for action-capable agents: delegation, acting identity, tools, policy checks, approvals, side effects, cost, and recovery. Plus retirement records with the same rigor as launches.

Where to start

Start by asking for the fleet count in six stages: created, active, shared, system-connected, write-capable, and autonomous. Not being able to answer is the first finding, not an embarrassment. If the gap is material, widen to a discovery reconciliation and a 90-day promotion pilot on one platform, measured on review time, workaround rate, and complete run evidence. Build the full operating machinery only when the sponsor wants it run.

Claim ledger

29/29
Checked
citations checked against primary sources in four independent verification clusters
0
Fabricated
no invented or unsupported sources surfaced in verification
10
Corrected
wording narrowed after source review, mostly vendor-claim labeling and scope limits
7
Demoted
useful signals kept out of the headline, including the famous spreadsheet error rate
CorrectedMicrosoft Work Trend Index: 15x active-agent growth and 18x in large enterprises are vendor claims from platform telemetry; no absolute counts or active-agent definition.microsoft.com
CorrectedOpenAI enterprise report: roughly 19x growth in weekly users of Custom GPTs and Projects and 4,000+ GPTs at one bank are vendor claims; the categories do not establish autonomy or production criticality.openai.com
CorrectedGartner survey: 75% of 360 IT application leaders involved with some form of agent, but no more than 15% even considering fully autonomous agents.gartner.com
CorrectedVendor customer story: 169+ employee-built agents at one customer is a vendor customer-story claim; a previously cited license figure did not survive source review.microsoft.com
CorrectedGlobal workforce study: policy-contravening use, uploaded company data, unreviewed reliance, and self-reported mistakes are documented for general AI use, not employee-built agents.unimelb.edu.au
CorrectedGallup: U.S. employee AI use roughly doubled in two years while policy coverage sat near 30%; no measurement of agents, autonomy, or production impact.gallup.com
ConfirmedNIST AI RMF: risk-proportionate inventory, roles, testing, monitoring, incident response, deactivation, and decommissioning across the lifecycle; voluntary guidance.nist.gov
ConfirmedEU AI Act: covered high-risk deployers carry oversight, monitoring, and log-retention duties; these do not extend identically to every employee-built agent.eur-lex.europa.eu
CorrectedFederal RPA playbook: separate development, test, and production environments plus credentialing, monitoring, and remediation; guidance, not a universal requirement.digital.gov
CorrectedCitizen-developer program: broad standard low-code access with AI-driven automation and mission-critical use on a higher-control path; an operating example, not comparative outcome evidence.digital.va.gov
DemotedSpreadsheet-error literature: the famous 94% error figure is not a reliable universal prevalence estimate; the durable lesson is the control pattern, not the statistic.dartmouth.edu
ConfirmedMicrosoft inventory and DLP docs: real inventory and data-loss controls, alongside documented limits: agent creation cannot be disabled and tenant isolation is unsupported in the studio tool.microsoft.com
CorrectedCross-platform registry sync: external-agent onboarding supports four named platforms in preview, requires authenticated connections, and is currently manual.microsoft.com
ConfirmedServiceNow governance lifecycle: discovery, risk review, deployment gates, monitoring, value, and retirement exist; only managed assets receive the full lifecycle, and discovered assets need configuration.servicenow.com
CorrectedAWS agent controls: registry with curator approval remains public preview; policy authorization applies to gateway-routed interactions, not an ambient enterprise layer.aws.amazon.com
DemotedGoogle inline safety integration: prompt and sensitive-data screening covers the initial user prompt and final response at configured points, not intermediate agent steps.cloud.google.com
ConfirmedSalesforce tracing and agent fabric: session tracing, registry, scanners, gateway policy, and telemetry are shipped; coverage is strongest for discovered, registered, or gateway-routed assets.salesforce.com
DemotedPricing and consumption evidence: published per-user, per-conversation, and per-assist pricing establishes metered exposure, not measured duplicated-agent waste.vendor docs
Where the evidence stops

The verified base establishes that agent-like use is growing inside major ecosystems, that lifecycle controls exist, and precisely where each control plane's coverage depends on integration. It does not establish how many employee-built agents hold production authority, because vendor telemetry lacks neutral denominators and workforce surveys measure general AI use. It does not quantify losses from duplicated agents, orphaned production agents, or agent-specific compliance failures; those stay risk scenarios and control objectives, not measured incident rates. And no comparative study yet shows that tiered promotion beats prohibition or universal central approval on incidents, speed, or value.

What would change our mind
Deep Dive staged from verified Storm Research v2 · nothing here asserts above the registry calibration