001✓ copiedspec

CS-03 · HelloFresh SE

Putting $150M of data risk on the register

Founding a data management department at HelloFresh, and the difference between exposure you have mitigated and exposure you can finally see.

Organisation
HelloFresh SE
Period
2020-06 to 2023-05
Engagement
Employed role, not a client engagement
Layers
data-foundation · ai-context
more than$150.0MUSDdata risk exposure brought under management
Organisation
HelloFresh SE
Period
2020-06 to 2023-05
Basis
Exposure carried by data assets that were previously unowned and untraced, brought under management by mapping lineage end to end, enabling an audit trail, consolidating to a single source of truth with verified and named-owner data assets, and enriching the metadata across them. Stated as exposure placed under management, not as loss avoided.

002✓ copieddetail

This was an in-house role, not a client engagement.

The situation

HelloFresh was operating globally, at scale, with no data management department.

That sentence is easy to read past. What it meant in practice was that nobody could answer three questions about most of the company’s data: where did this number come from, who is responsible for it, and what breaks if it is wrong. Not because anyone was careless. Because the company had grown faster than the discipline, which is the normal way this happens.

The risk in that state is not that something goes wrong. Things go wrong everywhere. The risk is that when something goes wrong you cannot tell how far it reached, so every incident becomes an unbounded investigation.

The constraint

Governance has a reputation problem inside fast companies, and the reputation is earned. Most governance programmes arrive as an approval queue, teams learn to route around them within a quarter, and the programme survives as paperwork.

I also could not stop the business to do this. HelloFresh was growing through the period.

What I did

I founded a data management department and gave it governance and process automation as its mandate, with dedicated people rather than a committee. Four things, in this order.

Lineage, end to end. For any number, you can trace backwards through every transformation to the system it originated in.

In plain terms: you can point at a figure in a report and follow it all the way back to where it came from, step by step, without asking anyone.

An audit trail. Every change to a data asset recorded, so the state at any past moment can be reconstructed.

In plain terms: if an auditor, a regulator, or a nervous CFO asks what this number said in March and who changed it, there is an answer, and the answer does not depend on anyone’s memory.

A single source of truth, with verified and owned assets. One agreed version of each important thing, checked, with a named person accountable for it.

In plain terms: one definition of revenue instead of three, and a specific human being whose job it is when that definition is wrong. Not “the data team”. A name.

Metadata enrichment across the estate. Every asset described well enough to be understood without asking its author.

In plain terms: you can tell what a dataset is, who uses it, and whether you should trust it, by reading about it rather than by finding the person who built it and hoping they still work here.

The sequencing matters and it is the transferable part. Lineage first, because you cannot assign ownership of something you cannot trace. Ownership before metadata, because descriptions written by nobody in particular rot. Automation throughout, because a control that depends on a person remembering is not a control.

The result

More than $150M of data risk exposure brought onto a register, with lineage behind it, an audit trail under it, and named owners against it.

I want to be precise about what that figure is and is not, because it is the number on this site most likely to be challenged and it deserves the challenge.

It is the exposure carried by data assets that were previously unowned and untraced, which the department brought under management. It is a measure of scope: how much of the company’s data risk moved from invisible to visible and owned.

It is not a claim that $150M of losses were avoided. That would require knowing what would have happened otherwise, and nobody has that number honestly. If you see a consultant quote a risk figure as loss avoided, ask them for the counterfactual and watch what happens.

The version I will defend is the smaller, duller, checkable one: before the department there was no register, and afterwards there was one, with that much on it.

What I would do differently

I built the register before I built the appetite for it. A register is only useful if somebody reviews it and acts on what it says. For the first year it was an artefact my team maintained and the rest of the leadership acknowledged politely. The forum where risk actually gets discussed should have existed before the register that feeds it.

I automated the controls before I simplified the processes. Some of what we automated should have been deleted instead. Automating a bad process makes it permanent and gives it a maintenance cost, and a few of those are probably still running.