The Data Is Reachable. The Answer Still Is Not.
A verified read on why access is the visible constraint in regulated enterprises while shared definitions, lineage, and valid joins often bind the defensible answer.
When the necessary population is already authorized, prioritize a semantic contract before a broad access program. The sponsor should be able to name the decision, grain, denominator, source lineage, valid join paths, entitlement filters, and owner for disputed meaning. If law, security, privacy, or source granularity blocks the necessary population, access remains binding.
The decision contract
Define which data, population, and level of detail may enter this decision. Entitlement is a property of the join, not only a gate before analysis.
Lock the metric name, grain, denominator, exclusions, time basis, and approved variants. A shared label is not a shared measure.
Show valid paths, entity keys, transformations, unmatched records, duplicates, and source-to-decision lineage.
Name the authority that resolves disagreement. Test the resulting query, metric, and permitted population against known cases.
Owner, briefing, proof
Owner
A named authority for each consequential metric and cross-system decision, with the ability to approve variants and resolve denominator or join disputes.
Briefing
A decision-friction view separating authorization delay from definition, join, lineage, quality, and escalation delay, plus the exceptions still unresolved.
Proof
A reproducible metric contract and lineage record with test cases, unmatched records, transformations, entitlement filters, and a visible approval history.
What leaders should take from it
Permission determines which data may enter the analysis. Meaning determines whether those authorized rows can support one defensible answer. Treating either as universally dominant loses the decision boundary.
Banking supervisors continue to report fragmented estates, lineage gaps, weak ownership, and incomplete aggregation more than a decade after formal data-governance requirements. Reachable data is not automatically decision-grade data.
Analyst research repeatedly finds effort spent locating fields, learning definitions, integrating sources, checking quality, and validating assumptions. The evidence does not establish a universal ranking, but it does show why access alone fails to finish the job.
Enterprise text-to-SQL benchmarks remain difficult even with schemas and documentation, while one bounded regulated-health study improved sharply when experts supplied a business-context document. The safe operating rule is narrow: place AI downstream of authoritative context and deterministic validation.
Semantic layers and metric stores provide useful primitives, but public outcome evidence is thin and largely vendor-published. Technical evidence also shows that join fanout and hidden aggregation assumptions can survive the layer. Fund the control objectives first; then test whether a product enforces them correctly.
No reviewed study establishes that semantics is the binding constraint in every enterprise or directly compares an access expansion with a semantic-contract intervention on decision error, rework, or audit closure. No qualifying independent comparative study isolates the organization-level effect of a commercial semantic layer. Current AI benchmarks combine several difficulties and do not isolate conflicting definitions.
The findings in full
Privacy, security, and regulatory constraints can define the lawful population, so access can remain the binding constraint. Inside an already authorized analytical boundary, permissions do not settle whether rows represent the same entities, whether measures share a grain and denominator, or whether the result is traceable. The most defensible model is nested: entitlement sets the feasible universe, then meaning and join validity determine whether a decision-grade answer exists.
Basel principles require integrated taxonomies, naming conventions, identifiers, reconciliation, and a concept dictionary. A 2023 supervisory review of 31 global systemically important banks found only two fully compliant with every principle. Later supervisory reporting still cites legacy and distributed estates, data lineage, fragmented responsibility, and weak management attention. These are qualitative supervisory findings, not a causal ranking against access.
A qualitative study of 35 analysts across 25 organizations found recurring difficulty locating data and field definitions, integrating sources, checking quality, and validating assumptions. Health-data studies show the same mechanism at schema level: large reconciliation workloads, local extensions, and one source concept mapping to several target concepts. These studies explain the work; they do not establish an enterprise-wide prevalence rate.
Spider 2.0 reports 21.3% task success for its strongest o1-preview-based agent across 632 enterprise workflows. EntSQL reports 15.9% with long documents and 21.4% with concise expert-curated evidence. A peer-reviewed GSK study on 60 historical pharmacovigilance questions improved from 8.3% with schema alone to 78.3% with an expert business-context document. These bounded studies support a narrow rule: authoritative context and deterministic checks belong upstream of model use.
Metric stores can centralize definitions, encode valid join paths, and generate SQL for downstream tools. A technical study shows that join fanout, deduplication, null handling, and hidden table choices can still yield inconsistent aggregates. Public outcome evidence remains mainly vendor-primary, and no qualifying comparative study isolated effects on decision error, audit closure, or organization-level productivity.
First moves before buying more access or tooling
For one priority decision, separate time lost to authorization, data location, definition conflict, cross-system joins, lineage reconstruction, quality repair, and owner escalation. Do not collapse them into one data-access label.
Name the purpose, grain, numerator, denominator, exclusions, time basis, valid variants, source fields, approval owner, tests, and change process. Preserve legitimate plural definitions instead of forcing false uniformity.
Document source-to-decision lineage, entity keys, valid paths, unmatched-record rates, duplicate handling, transformations, and exclusions. A query that executes is not proof that its population is complete.
Test which rows and attributes the user may combine for this purpose, not only whether each table can be opened separately. Record the entitlement filters in the reproducible decision artifact.
Give the model governed definitions, approved join paths, visible permissions, test cases, and deterministic validation. Use it to accelerate a settled process, not to adjudicate contested business meaning.
Test whether a semantic layer enforces the chosen definitions and entitlements across the actual toolchain. Do not ask a purchase to settle denominator, ownership, or purpose disputes the organization has not decided.
Start with one decision that already has enough authorized data to be plausible and still takes too long to defend. Reproduce the answer from source to denominator, recording every definition choice, join, exclusion, entitlement, and manual reconciliation. The resulting friction map tells you whether to invest next in access, semantic governance, integration, quality, or ownership.
Claim ledger
The evidence does not support a universal claim that semantics always outranks access. Access remains binding when the lawful population, necessary granularity, or source itself is unavailable. It also does not show that a semantic-layer purchase improves enterprise outcomes, that every organization needs one canonical definition, or that AI failure on analytics is caused specifically by undocumented meaning.
- A regulated-enterprise study compares access expansion with a governed semantic-contract intervention on time to auditable answer, rework, error, and audit closure.
- An independent multi-enterprise study measures outcomes before and after a metric store or semantic layer, including implementation and governance costs.
- A production benchmark tests AI analytics against conflicting metric definitions, entitlements, and undocumented joins rather than schema complexity alone.
- A decision-friction inventory across several enterprises shows permission delay consistently dominates definition, join, lineage, and reconciliation work after authorization.
- Regulators report sustained closure of lineage and aggregation gaps without stronger semantic ownership or shared definitions.