What you get
First, what you don't need us for
Your cloud platform already covers storage, pipelines, permissions and a technical catalogue. Lake On Rails adds nothing to those rows, so you do not pay twice. What it holds is the part that needs a human decision recorded: owners, accountability, written procedures, approvals, and the evidence that they happened.
Storage, pipelines, permissions
Classification and sensitivity labels
An accountable owner per dataset
A procedure with an approval
Seven more it covers in part, and four that need a person to decide and be recorded deciding. The full map, per platform.
A register of every dataset and report, with a name against each
Owner and steward by name. Where it comes from, which department it belongs to, how sensitive it is, how long you keep it, who consumes it, and which procedures apply. Filter by coverage gaps to see what still has nobody against it. Bulk-import what you already have in a spreadsheet.
Each record carries a readiness check: owner assigned, steward assigned, classification set, retention defined, description written, consumers identified. Six ticks means the analyst's knowledge about that dataset is now the business's.
Trust tiers, with written criteria and a promotion rule
Three tiers by default: raw as it landed, cleaned and checked, and ready to report from. The criteria for each are drafted for you and editable. Moving a dataset up a tier needs somebody to approve it, which is the whole difference between "the trusted version" being a fact and being an opinion.
Sensitivity and retention sit alongside: what is personal, what is commercially sensitive, what must be kept for seven years and what must not be kept at all.
If you already use the "medallion" vocabulary, these map onto bronze, silver and gold. We say trust tiers because your department heads have never heard of a medallion.
| Gold | Ready to report from. Owner, steward, definition and quality checks in place. The board sees this tier and nothing below it. |
|---|---|
| Silver | Cleaned, checked, conformed. Fit for analysis, not yet fit for a number someone has to stand behind. |
| Bronze | As it landed from the source. Kept because you cannot rebuild it, not because anyone should report from it. |
Defaults, shown with the medallion names. Every criterion is yours to edit.
The five roles
| Data Owner | Accountable for what a dataset means and who may see it. A department head, not a technician. |
|---|---|
| Data Steward | Keeps the record true day to day. Often the analyst. |
| Producer | Runs the system or pipeline the data comes from. |
| Consumer | Uses the data. Asks for access through the workflow rather than by email. |
| Custodian | Looks after the platform it lives on. |
Roles, and a responsibility matrix you adjust rather than invent
Who is responsible, accountable, consulted and informed for each of the activities the operating model involves: onboarding a dataset, approving access, promoting a tier, handling a schema change, running the monthly review. Fifty cells, already filled with sensible defaults. You change the ones that are wrong for you.
When a role is assigned, the person is told, with a short guide to what it asks of them. One person can hold several roles; the product is built for a business where one person is most of the data team.
Seven procedures, written and ready to edit
Every tool in this category ships an empty template. Lake On Rails ships the procedures: onboarding a new dataset, classifying data, granting and reviewing access, promoting between tiers, handling a schema change, retiring a dataset, and running the periodic review. Each has a purpose, a scope, prerequisites, a named role per step and an expected outcome per step.
They are yours to edit, and editing an approved procedure withdraws its approval until someone signs it off again. Every version is kept.
Read one and you will know inside a minute whether it fits your business. That is the fastest test of this product there is.
Approvals that run through the product, not through email
Four workflows out of the box: dataset onboarding, access requests, schema changes and tier promotion. A request goes to the person the responsibility matrix says should decide it. They get an email and an in-app notification, open it, and decide in a browser. The decision, the person and the time are recorded.
A steward who spots a problem with a dataset — a broken source, a wrong owner, stale metadata — flags it, and that becomes a tracked item with an owner rather than a message that scrolls away.
Professional plan and above. Enterprise adds custom workflow templates.
An access request, end to end
- 1An analyst finds the dataset in the register and clicks Request Access, saying why.
- 2The Data Owner is notified. They see who is asking, for what, and the dataset's sensitivity.
- 3They approve or decline in a browser, in a minute, with a reason.
- 4The custodian grants the access on the platform. The whole chain is in the audit trail.
One score, six parts, and it goes down when something slips
Ownership coverage, procedure sign-off, responsibility matrix completeness, classification coverage, platform gap closure and training completion, each shown as the fraction it is measured from, averaged into a single number from zero to one hundred. Snapshots are taken weekly, so the trend is real rather than remembered.
Alerts fire when a dataset lands without an owner, a procedure's review date passes, or a role goes unstaffed. A maturity roadmap turns the gaps into a sequence of milestones, so the next thing to do is never a mystery.
Training that reaches people when they are appointed
Short courses and role guides, built in: what a Data Owner is for, how to classify a dataset, how to handle an access request, what the trust tiers mean. A learning path per role, with completion tracked, so "0 of 28 required modules completed" is a fact on the dashboard rather than a suspicion.
You can author your own courses too, for the procedures and definitions that are specific to your business.
Professional plan and above.
Built-in content
- Role guides for each of the five roles
- Courses on ownership, classification, access, promotion and the review cycle
- A glossary in plain English, so "steward", "tier" and "retention" mean one thing
- Templates for the decisions that get deferred: retention schedules, tier criteria, role definitions
- A help centre that explains what every screen is for
Above Fabric, above AWS, and useful before either
The capability map shows, for each platform, what it covers natively and what still needs a human decision recorded. Switch the tab from Fabric to AWS and the right-hand column is the same: the gaps are about your business, not your cloud.
Connect a platform and Lake On Rails reads technical metadata — dataset names, schemas, owners, refresh times — from Microsoft Fabric and Purview, or from the AWS Glue Data Catalog, on a schedule. It never moves, transforms or stores the contents of your data. Or run it with no connection at all, on metadata you enter yourself.
Fabric or AWS? The honest comparison, and what survives the decision
This is what "prove it" looks like
Who changed a classification, from what to what, when, and on whose authority. Field-level, and attributed to a person. The application has no edit or delete path for the trail: every change writes a new event and nothing is overwritten. Filter by entity, action, person or date, and export to CSV.
The executive summary turns the same data into one page for the board: a score, five measures, three plain-English talking points. On the Enterprise plan, controls are mapped against a framework catalogue with quarterly evidence snapshots.
Ask us the obvious follow-up, whether a database administrator could still alter it. We would rather answer that now than have you find it later.
Your assistant can do the typing
The reason these programmes stall in month three is not disagreement. It is that somebody has to write four hundred dataset descriptions, and nobody has the fortnight. Lake On Rails ships an MCP server and a versioned REST API so an AI assistant can do that drafting against your own register — and so a person still approves every word of it before it counts.
Connect it to Microsoft Copilot, ChatGPT, Claude or an agent you have built yourself. It is the same model the screens use: the register, roles, procedures, work items, courses, workflows and the audit trail. Any of it is optional — the product is complete without ever turning this on, and it ships turned off.
Professional and Enterprise. Authorisation is OAuth 2.1 with PKCE, or an API key. What the plans include.
What people use it for
- Draft the description and classification for forty new tables, then have a steward approve them in one sitting instead of typing them for a fortnight
- Ask, in a chat window, "which of finance's datasets have no retention set?" and get the answer from the register rather than an opinion
- Block a deployment when the dataset it creates has no owner
- Have an agent open a work item when something drifts, with its reasoning attached
- Pull the maturity score into your own reporting, or into the agent that writes your board pack
What stops it going wrong
- Off until you turn it on. Both API keys and OAuth default to disabled. Enabling either is an explicit, audited act by one of your admins.
- Proposals, not edits, by default. Writes land in a review queue with the assistant's stated reason as the first column and a field-by-field diff. Applying writes directly is a setting you choose, not the default.
- It cannot quietly overwrite a person. Every proposal records what it was based on; if a human has edited the same field in the meantime, approval stops and shows a three-way diff instead of clobbering them.
- Nothing is lost. Every change is snapshotted with the reason given, shown on the record as an AI-assisted change, and revertible to the previous version in one click.
- Narrow, named permissions. Twelve scopes of the form registry:read, sops:write. No wildcards, no delete, and each is described in a sentence on the consent screen before anyone approves it.
- The person's role still caps it. A credential can never exceed what its user may do today — demote them and the credential loses the same power on its next request.
- Bounded. Rate and daily quota limits per credential so a looping agent cannot run away with your register, an optional IP allowlist, and search results that return pointers rather than the contents of records.
Every one of those is enforced in the product rather than described in a policy document, and every action an assistant takes is attributed in the same audit trail as a person's — which client, which credential, when, and why.
The first step costs you nothing
Forty-five minutes with whoever runs your reporting
We tell you honestly whether this is worth doing at all, and roughly what it would take. If the answer is not yet, you will hear that. "Not for us" is a fine outcome, and a better one than a slow maybe.