Guide · Where to start

A data operating model for a business of 50 to 1,000 people

The published frameworks — DAMA-DMBOK, DCAM, the vendor pillar diagrams — were written for organisations with a data office. If you have 300 staff and one analyst who knows where everything is, they will tell you what good looks like and nothing about how to get there from here.

9 minute read For: whoever has been handed this Reviewed September 2026

The short answer

A workable framework at this size is five things written down: a register of what you have and who owns it, a classification scheme, a set of procedures for the handful of moments that go wrong, a responsibility matrix, and a measure that can go down. Everything in DAMA-DMBOK's eleven knowledge areas is real, but the order matters more than the coverage, and ownership comes first because nothing else can be agreed without it.

Five stacked layers labelled register, classification, procedures, responsibilities and measurement, sitting above a platform layer marked Fabric or AWS, showing the operating model as independent of the platform beneath it.
The five parts, and the line under them: everything above the line stays true when the platform below it changes. Illustration

Why the published frameworks stall

DAMA-DMBOK is a good book. It describes eleven knowledge areas, each with its own maturity path, and it assumes a reader who can staff them. DCAM assumes a bank. The vendor diagrams — four pillars, six pillars, seven pillars, depending on who is selling — describe a destination without a route, which is why they are easy to agree with in a meeting and impossible to act on the following Monday.

The gap is not intellectual. It is that a business of this size has roughly one to three person-days a week to spend on this in total, spread across people who have other jobs. A framework that needs a working group per knowledge area does not fail; it just never starts.

So the question is not "what does a complete framework contain?" It is "what is the smallest thing that changes behaviour, and what order do we build it in?"

The five parts

1. A register of what you have

One list of the datasets and reports the business depends on. Source system, department, a named owner, a sensitivity label, a trust level, a status. Nothing more on the first pass — a register with eight columns that is complete beats one with thirty that is a quarter filled.

Expect 60 to 300 entries. Expect a third of them to have nobody against them. That number is the single most useful thing you will produce in the first month, because it converts an argument about principles into a list.

2. A classification scheme

Two axes, kept separate. Sensitivity: who may see this, and what happens if it leaks. Trust: how much may a decision-maker rely on it. Three levels on each is enough, four is a stretch, five is a scheme people will guess at rather than apply.

The commonest mistake is collapsing the two into one bronze/silver/gold ladder, which produces the confident error of treating a refined table as a trustworthy one. See trust tiers and medallion architecture for why those are different questions.

3. Procedures for the moments that go wrong

Not a manual. Seven or eight short documents covering the events that actually cause damage: onboarding a new dataset, handling an access request, managing a schema change, promoting something to trusted, retiring a dataset, responding to a quality incident, running a periodic review.

Each one wants the same shape — purpose, scope, prerequisites, numbered steps with a role and an expected outcome against each, a checklist, an approver. One of ours is published in full so you can judge the level of detail rather than imagine it.

4. A responsibility matrix

A grid: activities down the side, roles across the top, one accountable name per row. It takes an afternoon to draft and about three weeks to agree, and the three weeks are the work. See the RACI guide for a filled-in example and the two rows that always cause an argument.

5. A measure that can go down

Pick four or five things you can count without a project: the share of registered datasets with a named owner, the share classified, the share of procedures approved and in date, the share of owners who have done the training. Publish the composite monthly.

The requirement that it can fall is not a nicety. A score that only rises is a marketing artefact, and everyone in the room knows it within two months.

A six-week sequence that works

This is roughly the shape of the engagements we run, and it is deliberately front-loaded with the uncomfortable part.

  1. Week 1 — inventory. Two or three sessions with the people who build reports. Produce the register with owners left blank. Do not attempt to be complete; attempt to be honest.
  2. Week 2 — the shortlist and the classification scheme. Mark the entries where a wrong number reaches a customer, a regulator, a lender or the board. Draft the two classification axes and test them against ten real datasets.
  3. Week 3 — ownership conversations. Take the shortlist to the named individuals. This is the week the whole thing lives or dies, and it cannot be done over email.
  4. Week 4 — procedures. Adapt the drafts rather than authoring from nothing. Argue about the approval steps; that argument is where the real policy gets set.
  5. Week 5 — the matrix and the measure. Fill in the RACI, agree the four or five measures, take the baseline. It will be a bad baseline. Publish it anyway.
  6. Week 6 — approval and handover. A sponsor signs the procedures. Owners get their ninety minutes. Diarise the monthly review before anyone leaves the room.

Then it becomes about half a day a week for whoever holds it, an hour a month per owner, and an hour a month for the sponsor. The maintenance figure is the one to check people can actually commit, because the six weeks are survivable on adrenaline and the twelve months afterwards are not.

What to leave out on the first pass

Tempting and premature: automated lineage, a full business glossary, master data management, quality rules engines, anything with the word platform in its budget line. Each of them is a reasonable second-year investment and each of them will absorb the entire first year if you let it.

Two specific traps. A glossary project that tries to define every term produces four hundred definitions nobody reads; twenty terms that appear in board papers is the version that works. And a tool selection run before ownership is agreed will pick the tool with the best catalogue features, which is not the problem you have.

Doing it before the platform decision

If you are weighing Microsoft Fabric against AWS, or have not started, this work is worth more now than it will be later. Everything in the five parts is about your business rather than your cloud: who owns customer data, what "trusted" means here, what is sensitive, how long you keep things. All of it becomes an input to the migration rather than a retrofit after it — and retrofitting ownership onto a platform that has already been built is where the cost genuinely lands.

Common questions

What should a data governance framework contain?
At mid-market scale: a register of datasets with named owners, a classification scheme covering sensitivity and trust separately, seven or eight written procedures for the events that cause damage, a responsibility matrix with one accountable name per activity, and four or five measures published monthly that are capable of falling.
How long does it take to set up a data operating model?
About six weeks of concentrated work to get the first version agreed and approved, on roughly two days a week from whoever holds it, plus ninety minutes per data owner. After that it settles to around half a day a week, an hour a month per owner and an hour a month for the sponsor. Configuring software is quicker than that; agreeing accountability is not.
Do we need DAMA-DMBOK or DCAM?
Not to start. They are reference works describing what a complete data management capability looks like, written for organisations with dedicated data offices. They are useful for checking coverage in year two. They are a poor sequencing guide for a company with one analyst and a Power BI estate.
Should we do this before or after choosing a cloud platform?
Before, if the decision is still open. Every decision in the operating model is about your business rather than your platform, so it survives the choice, and it makes the platform evaluation sharper — you will be comparing the two clouds against a written list of what you actually need governed rather than a feature grid.

Where the product comes in

You start from drafted procedures, not an empty system

Lake On Rails ships the register, the two classification axes, a responsibility matrix and seven written procedures already drafted, so the six weeks are spent approving and adjusting rather than authoring. The maturity score is built from measures that can fall, and the executive summary prints to one page for a board pack.

What it does not do is have the ownership conversation. That is week three, and it is yours.

The first step costs you nothing

Forty-five minutes with whoever runs your reporting

We tell you honestly whether this is worth doing at all, and roughly what it would take. If the answer is not yet, you will hear that. "Not for us" is a fine outcome, and a better one than a slow maybe.