Guide · Platforms
Microsoft Purview and Fabric: what they cover, and what they leave to you
If you are on the Microsoft stack, someone has told you that Purview handles this. It handles a genuine and useful part of it — the part a machine can do by reading your estate — and the remainder is not a gap in the product. It is the half that is about your business rather than your platform.
The short answer
Purview and Fabric are strong at what can be derived automatically: scanning sources, harvesting technical metadata, classifying against pattern rules, tracing lineage through code they can read, and enforcing labels and access. They cannot supply the things no scan can see — who is accountable for a dataset, what your business means by trusted, what retention period is defensible, and the written procedure people follow when something changes.
What the platform genuinely does
- Discovery and scanning. Point it at your sources and it will inventory them, harvest schemas, and keep that current. This is real work you would otherwise do by hand, and it does it better.
- Technical metadata. Table sizes, column types, freshness, usage. Automatic and accurate.
- Pattern-based classification. A large library of built-in classifiers detects things that look like national insurance numbers, card numbers, email addresses. Useful as a first sweep and as a safety net.
- Lineage through code. Where transformations are expressed in SQL, pipelines or notebooks, it traces them. See the lineage guide for where the trail stops.
- Sensitivity labels and enforcement. Once a label exists, Microsoft Information Protection can carry it into Office, Power BI and access policy. This integration is the strongest argument for the Microsoft stack and there is no honest way to match it from outside.
- Access and identity. Entra ID, conditional access, privileged access — mature, and not something to reproduce elsewhere.
If you are on Fabric and not using this, the first move is to turn it on, not to buy anything.
What it leaves to you
Accountability
No scan reveals who is accountable for a dataset. Purview has fields for owners and stewards, and they are as good as what you put in them — which, in most tenants we have looked at, is the identity of whoever ran the scan. The conversation that produces a real owner happens in a meeting, and a tool can only record the outcome.
What "trusted" means here
Pattern classification finds what looks sensitive. It cannot tell you what your business considers board-ready, or what has to be true before a figure goes in front of a customer. Those criteria are written by people who understand the trade, and then something has to check that they were met. See trust tiers.
Retention with a basis
Microsoft Purview's data lifecycle management applies retention policies well. What it cannot supply is the period and the reason — the statute, the limitation period, or the documented business need. A policy applied without a recorded basis is a setting, and a setting is not a defence. See the retention guide.
Procedures people follow
When a schema changes, someone has to be told, someone has to assess the impact, and someone has to approve. The platform can trigger the alert. It has no opinion about who reviews it, what they check, or what happens if they disagree — and that is the artefact an auditor asks for.
Licensing and the practical shape
Purview's governance features are not uniformly included with a Fabric capacity, and the packaging has changed more than once — the data governance capabilities carry their own commercial model separate from the compliance features many organisations already have through Microsoft 365. Check the current position against your own agreement rather than against a blog post, including this one. What is worth knowing before the conversation is that "we already have Purview through M365" and "we have Purview data governance" are frequently different statements.
Two practical notes from deployments. Scanning cost scales with how often you scan and how much you scan, so a default of everything nightly gets expensive quietly. And the classification results need review: pattern classifiers produce false positives, and a column of eight-digit product codes flagged as something sensitive will erode trust in the whole sweep if nobody prunes it.
The same picture on AWS
Glue Data Catalog, Lake Formation, Macie and DataZone divide the work differently but the line falls in the same place. The catalogue harvests, Macie pattern-matches for sensitive data, Lake Formation enforces access, and none of them names an owner or writes a procedure. If you have not chosen yet, the platform comparison sets out what actually differs and what does not.
What to do with this
Use the platform for everything it can derive. Do not attempt to reproduce scanning, technical metadata, label enforcement or identity outside it — you will lose, and the integration is the reason you chose the stack.
Then handle the decision layer deliberately: ownership, trust criteria, retention basis, procedures, and the evidence that they were followed. Some organisations do that in documents and a spreadsheet, which works until the spreadsheet goes stale. Some use a product built for it. Either is defensible; assuming the platform is already doing it is the option that is not.
Common questions
Does Microsoft Purview handle data governance for us?
Is Purview included with Microsoft Fabric?
What is the difference between a data catalogue and a data operating model?
Is the picture different on AWS?
Where the product comes in
It reads what the platform knows, and holds what it cannot
Lake On Rails connects to Microsoft Fabric, Purview and AWS Glue and shows, capability by capability, what the platform covers natively, what it covers partly, and what still needs a human decision recorded. It deliberately does not reproduce scanning, label enforcement or identity — the platform is better at those and you are already paying for them.
The first step costs you nothing
Forty-five minutes with whoever runs your reporting
We tell you honestly whether this is worth doing at all, and roughly what it would take. If the answer is not yet, you will hear that. "Not for us" is a fine outcome, and a better one than a slow maybe.