Guide · Classification
Data classification: how many levels you actually need
Classification schemes fail in a recognisable way. Someone drafts five levels with elegant definitions, everyone agrees in the meeting, and six months later 90% of datasets are labelled with the second one because nobody could tell the difference between the second and the third at four o'clock on a Friday.
The short answer
Three levels for most businesses of 50 to 1,000 people: Public, Internal and Restricted, with a fourth only if you genuinely handle material that would cause serious harm. What makes a scheme work is not the number of levels but whether each one has a test a non-specialist can apply alone, and a stated consequence — who may see it, where it may go, how long it is kept.
Start from the consequence, not the label
A level is worth having only if something different happens because of it. So write the consequences first — who may see this, may it leave the building, may it go in an email, how long do we keep it, what happens if it leaks — and then work out how many distinct sets of consequences you actually have.
Most mid-market businesses have three. Some have four. Nobody we have worked with has needed six, and the organisations that adopted six ended up using two.
The three levels
Public
Disclosure causes no harm. Published prices, marketing material, the filed accounts, anything already on the website. Small category, but worth having, because it stops people treating everything as confidential by default and gives the scheme a floor.
Test: could this appear on the website tomorrow without anyone being consulted?
Internal
For staff, not for the world. Most operational data, most internal reporting, project documents, supplier terms. Disclosure would be embarrassing or commercially unhelpful rather than damaging.
Test: would you be comfortable with any employee seeing this, and uncomfortable with a competitor seeing it?
This will be your largest category, and that is correct. A scheme where the middle level holds most of the data is working, provided the level above it is genuinely reserved.
Restricted
Access is granted individually, on a need to know, and refusal is the default. Personal data about staff or customers, payroll, anything under NDA, security material, unpublished financials, legal advice.
Test: would a named person have to approve each individual who sees this?
The discipline is in refusing to let this level grow. If more than a fifth of your registered datasets end up Restricted, either your business is unusual or the level is being used as a way of expressing that data is important.
A fourth level, if you need one
Some businesses genuinely hold material where disclosure would cause serious harm to a person or would end a commercial relationship — health data, safeguarding records, an acquisition in progress. If so, add a top level with a small named list of who holds it and a rule that it never leaves a specified system. Do not add it speculatively.
Where special category data sits
UK GDPR's special category data — health, biometrics, genetics, race or ethnicity, political opinions, religious beliefs, trade union membership, sex life, sexual orientation — is not a rung on your sensitivity ladder. It is a legal overlay with its own requirements, and the commonest scheme design error is to model it as "level four".
Treat it as a separate flag on the dataset. Processing it needs an Article 9(2) condition on top of an ordinary lawful basis, and most UK conditions also require an appropriate policy document. Criminal offence data sits under its own regime and is handled the same way in practice.
The practical rule: a dataset can be Internal and carry a special category flag, and the flag is what triggers the extra review. Collapsing the two loses that.
Sensitivity is not trust
Two different questions get bundled into the word "classification" and separating them is worth more than any refinement of the levels.
- Sensitivity answers who may see this.
- Trust answers how much may I rely on it.
They are independent. Raw event data straight off a website is untrusted and often highly sensitive — full of device identifiers and IP addresses, all of which are personal data. A polished board-ready revenue summary is highly trusted and merely Internal. Run them as two fields. See trust tiers and medallion architecture for the second axis.
Write criteria a tired person can apply
The measure of a good scheme is whether a steward classifying their eleventh dataset on a Thursday afternoon gets the same answer as the policy author would. That needs examples, not definitions.
Against each level, list five real datasets from your own business by name. The list does more work than the paragraph above it, and it settles arguments in seconds. Add the two or three edge cases you argued about while drafting, with the answer you reached and one sentence of reasoning — those are the cases that will come up again.
Screening for personal data properly
Classification is when you find out what you are actually holding, and column-name scanning is how organisations miss it. Personal data under Article 4(1) is any information relating to an identified or identifiable living individual, directly or indirectly. Device identifiers, session IDs, IP addresses, cookie IDs and pseudonymised customer keys all qualify, and none of them looks like a name.
So read the column descriptions and a sample of the values, not only the headers, and ask the source owner what each key joins to. A clickstream table with no field called email can identify every visitor through a key that joins to the CRM.
Record the answer even when it is no. A recorded "no personal data, checked by name on this date" tells the next reviewer you asked. A blank tells them nothing at all.
Applying it without a six-month project
- Draft the three levels with tests and five named examples each. Half a day.
- Test the draft against ten real datasets with two stewards. Expect to change the wording; that is what the test is for.
- Classify the shortlist — the datasets whose failure would be visible outside the company — rather than everything. Fifteen to forty entries.
- Attach one consequence per level in your actual systems: a permissions group, a retention default, a rule about external sharing. A label with no mechanism behind it decays within a year.
- Review annually, and whenever a dataset changes purpose. Purpose change is the event that quietly moves something from Internal to Restricted.
Common questions
How many data classification levels should we have?
Is special category data just the highest classification level?
What is the difference between data classification and trust tiers?
How do we find personal data we do not know we have?
Where the product comes in
Sensitivity and trust are two fields, and both are on the record
Every dataset in Lake On Rails carries a classification and a trust tier separately, with the criteria for each published alongside so a steward can check rather than guess. Changes to either are written to the audit trail with the previous value, which is the version an auditor asks for.
The first step costs you nothing
Forty-five minutes with whoever runs your reporting
We tell you honestly whether this is worth doing at all, and roughly what it would take. If the answer is not yet, you will hear that. "Not for us" is a fine outcome, and a better one than a slow maybe.