← Back to Insights
Briefing · Decision-Grade Data

The Data Is Reachable. The Answer Still Is Not.

A verified read on why access is the visible constraint in regulated enterprises while shared definitions, lineage, and valid joins often bind the defensible answer.

Conditional sign-off verdict

When the necessary population is already authorized, prioritize a semantic contract before a broad access program. The sponsor should be able to name the decision, grain, denominator, source lineage, valid join paths, entitlement filters, and owner for disputed meaning. If law, security, privacy, or source granularity blocks the necessary population, access remains binding.

The decision contract

Authorize

Define which data, population, and level of detail may enter this decision. Entitlement is a property of the join, not only a gate before analysis.

Define

Lock the metric name, grain, denominator, exclusions, time basis, and approved variants. A shared label is not a shared measure.

Join and trace

Show valid paths, entity keys, transformations, unmatched records, duplicates, and source-to-decision lineage.

Own and test

Name the authority that resolves disagreement. Test the resulting query, metric, and permitted population against known cases.

Owner, briefing, proof

Owner

A named authority for each consequential metric and cross-system decision, with the ability to approve variants and resolve denominator or join disputes.

Briefing

A decision-friction view separating authorization delay from definition, join, lineage, quality, and escalation delay, plus the exceptions still unresolved.

Proof

A reproducible metric contract and lineage record with test cases, unmatched records, transformations, entitlement filters, and a visible approval history.

What leaders should take from it

1
Access and semantics are nested controls.

Permission determines which data may enter the analysis. Meaning determines whether those authorized rows can support one defensible answer. Treating either as universally dominant loses the decision boundary.

2
Regulated firms still struggle after access controls exist.

Banking supervisors continue to report fragmented estates, lineage gaps, weak ownership, and incomplete aggregation more than a decade after formal data-governance requirements. Reachable data is not automatically decision-grade data.

3
The recurring work is reconciliation.

Analyst research repeatedly finds effort spent locating fields, learning definitions, integrating sources, checking quality, and validating assumptions. The evidence does not establish a universal ranking, but it does show why access alone fails to finish the job.

4
AI inherits unresolved meaning.

Enterprise text-to-SQL benchmarks remain difficult even with schemas and documentation, while one bounded regulated-health study improved sharply when experts supplied a business-context document. The safe operating rule is narrow: place AI downstream of authoritative context and deterministic validation.

5
Buy the capability before the category.

Semantic layers and metric stores provide useful primitives, but public outcome evidence is thin and largely vendor-published. Technical evidence also shows that join fanout and hidden aggregation assumptions can survive the layer. Fund the control objectives first; then test whether a product enforces them correctly.

Where the evidence stops

No reviewed study establishes that semantics is the binding constraint in every enterprise or directly compares an access expansion with a semantic-contract intervention on decision error, rework, or audit closure. No qualifying independent comparative study isolates the organization-level effect of a commercial semantic layer. Current AI benchmarks combine several difficulties and do not isolate conflicting definitions.

The findings in full

9/10
Access is necessary, but semantics often binds once the authorized boundary is sufficient.

Privacy, security, and regulatory constraints can define the lawful population, so access can remain the binding constraint. Inside an already authorized analytical boundary, permissions do not settle whether rows represent the same entities, whether measures share a grain and denominator, or whether the result is traceable. The most defensible model is nested: entitlement sets the feasible universe, then meaning and join validity determine whether a decision-grade answer exists.

9/10
Regulated-enterprise evidence points to reconciliation, lineage, and ownership failures after access exists.

Basel principles require integrated taxonomies, naming conventions, identifiers, reconciliation, and a concept dictionary. A 2023 supervisory review of 31 global systemically important banks found only two fully compliant with every principle. Later supervisory reporting still cites legacy and distributed estates, data lineage, fragmented responsibility, and weak management attention. These are qualitative supervisory findings, not a causal ranking against access.

8/10
Analytics work repeatedly stalls on field meaning, integration, quality, and assumption checking.

A qualitative study of 35 analysts across 25 organizations found recurring difficulty locating data and field definitions, integrating sources, checking quality, and validating assumptions. Health-data studies show the same mechanism at schema level: large reconciliation workloads, local extensions, and one source concept mapping to several target concepts. These studies explain the work; they do not establish an enterprise-wide prevalence rate.

7/10
AI cannot be trusted to invent the missing semantic contract.

Spider 2.0 reports 21.3% task success for its strongest o1-preview-based agent across 632 enterprise workflows. EntSQL reports 15.9% with long documents and 21.4% with concise expert-curated evidence. A peer-reviewed GSK study on 60 historical pharmacovigilance questions improved from 8.3% with schema alone to 78.3% with an expert business-context document. These bounded studies support a narrow rule: authoritative context and deterministic checks belong upstream of model use.

6/10
Semantic-layer primitives are useful, but the product category has not proved the enterprise outcome.

Metric stores can centralize definitions, encode valid join paths, and generate SQL for downstream tools. A technical study shows that join fanout, deduplication, null handling, and hidden table choices can still yield inconsistent aggregates. Public outcome evidence remains mainly vendor-primary, and no qualifying comparative study isolated effects on decision error, audit closure, or organization-level productivity.

First moves before buying more access or tooling

01
Run a decision-friction diagnostic.

For one priority decision, separate time lost to authorization, data location, definition conflict, cross-system joins, lineage reconstruction, quality repair, and owner escalation. Do not collapse them into one data-access label.

02
Write the metric contract.

Name the purpose, grain, numerator, denominator, exclusions, time basis, valid variants, source fields, approval owner, tests, and change process. Preserve legitimate plural definitions instead of forcing false uniformity.

03
Make lineage and join evidence visible.

Document source-to-decision lineage, entity keys, valid paths, unmatched-record rates, duplicate handling, transformations, and exclusions. A query that executes is not proof that its population is complete.

04
Carry entitlements into the join.

Test which rows and attributes the user may combine for this purpose, not only whether each table can be opened separately. Record the entitlement filters in the reproducible decision artifact.

05
Place AI downstream of the reconciled contract.

Give the model governed definitions, approved join paths, visible permissions, test cases, and deterministic validation. Use it to accelerate a settled process, not to adjudicate contested business meaning.

06
Evaluate tooling against the operating contract.

Test whether a semantic layer enforces the chosen definitions and entitlements across the actual toolchain. Do not ask a purchase to settle denominator, ownership, or purpose disputes the organization has not decided.

Where to start

Start with one decision that already has enough authorized data to be plausible and still takes too long to defend. Reproduce the answer from source to denominator, recording every definition choice, join, exclusion, entitlement, and manual reconciliation. The resulting friction map tells you whether to invest next in access, semantic governance, integration, quality, or ownership.

Claim ledger

17/17
Checked
citations traced to their sources on August 6, 2026
0
Fabricated
no invented source or unsupported citation survived verification
7
Corrected
dates, figures, attribution, and claim scope narrowed after source review
4
Demoted
indirect or vendor-primary evidence kept out of stronger conclusions
CorrectedBasel Committee, BCBS 239: integrated taxonomies, naming conventions, identifiers, reconciliation, and a concept dictionary support aggregation across legal entities and systems. The 2013 text does not use the report's exact entitlement-aware-join term.bis.org
ConfirmedBasel Committee 2023 progress report: review of 31 global systemically important banks found only two fully compliant with every principle; most still needed significant work. Supervisory assessment, not a causal experiment.bis.org
ConfirmedBasel Committee 2026 newsletter: cites data lineage, legacy and distributed estates, fragmented responsibilities, change resistance, and insufficient management attention. Qualitative and informational.bis.org
ConfirmedBank of England 2021: common identifiers and definitions are needed to combine information across proprietary and isolated standards. Explanatory policy evidence, not a quantified study.bankofengland.co.uk
CorrectedPA Consulting for the FCA and Bank of England, 2023: BIRD maps internal-bank data through logical and semantic rules into a harmonized dictionary. The selected initiatives were not representative.bankofengland.co.uk
ConfirmedKandel et al.: interviews with 35 analysts across 25 organizations found recurring work in locating fields, understanding definitions, integrating sources, checking quality, and verifying assumptions. Qualitative, not prevalence evidence.stanford.edu
CorrectedLee et al. 2010: 9,418 validated clinical terms, with 6,750 mapped lexically or semantically and 550 mapping to multiple standard concepts. One Korean cancer hospital.nih.gov
CorrectedBrown et al. 2016: 1,142 source elements and 1,175 matched pairs in the SMASH integration study. The former 87% arithmetic claim was removed after verification.nih.gov
CorrectedNIST SP 800-207: per-session, resource-specific least privilege supports entitlement-aware joins by engineering inference; the standard does not prescribe joins.nist.gov
CorrectedHHS HIPAA minimum necessary guidance: policies must identify user classes, data categories, and conditions of access. Sector-specific and does not prescribe analytical joins.hhs.gov
CorrectedSpider 2.0, ICLR 2025: final paper reports 21.3% for the strongest o1-preview agent on 632 enterprise workflows. The older 17.0% figure is stale and was removed.openreview.net
DemotedEntSQL 2026 preprint: 1,066 bilingual examples across five domains; 15.9% with long documents and 21.4% with concise expert-curated evidence. One enterprise setting and no decision-quality measure.arxiv.org
ConfirmedPainter et al. 2025: on 60 historical questions from one proprietary GSK database, schema-only prompting passed 8.3% and an expert business-context document passed 78.3%. Company-authored and funded, with limited external validity.doi.org
ConfirmedHuang, Damalapati, and Wu 2023: join fanout and heuristic deduplication can produce inconsistent aggregates in semantic layers. Technical workshop evidence, not an organization-level outcome comparison.arxiv.org
Demoteddbt Semantic Layer documentation: centralized definitions, entity-based join paths, generated SQL, and APIs are confirmed product mechanics. Vendor documentation is not outcome evidence.getdbt.com
DemotedPhiladelphia Inquirer customer account: qualitative claims of faster delivery and fewer errors lack baseline, comparison design, implementation cost, or independent audit.getdbt.com
DemotedJaakkola 2025 thesis: six tasks, three fully correct results, and mean completion of 2:30 versus 1:45 for the established workflow. No otherwise-identical no-semantic-layer arm.theseus.fi
Contested signals: held out of the verdict

The evidence does not support a universal claim that semantics always outranks access. Access remains binding when the lawful population, necessary granularity, or source itself is unavailable. It also does not show that a semantic-layer purchase improves enterprise outcomes, that every organization needs one canonical definition, or that AI failure on analytics is caused specifically by undocumented meaning.

What would change this conclusion

Related work

Verified research · 17 citations checked · 0 fabricated · 7 corrected · 4 demoted