001✓ copiedspec

CS-05 · HelloFresh SE

The Data Academy: teaching a company to stop breaking its own data

A certification programme at HelloFresh that reached over a hundred people, and the sequencing mistake that cost it a year of momentum.

Organisation
HelloFresh SE
Period
2020-06 to 2023-05
Engagement
Employed role, not a client engagement
Layers
ai-enablement · data-foundation
Increase of more than100peopleemployees certified through the Data Academy
Organisation
HelloFresh SE
Period
2020-06 to 2023-05
Basis
Count of employees who completed the Data Academy certification track.
Decrease of 35%reductiondata errors after the Data Academy
Organisation
HelloFresh SE
Period
2020-06 to 2023-05
Basis
Reduction in recorded data incidents, achieved through automated incident detection, hardened processes at the steps that historically failed, and systematic remediation of the underlying systems. Measured against the recorded incident count before those controls were introduced.

002✓ copieddetail

This was an in-house role, not a client engagement.

The situation

HelloFresh was producing far more data than the business could correctly interpret. That is a good problem to have and a dangerous one to leave alone.

The failure mode was not what people assumed. When a number was wrong, the instinct in the organisation was to blame the pipeline. When we traced the incidents, most of them started with a person: a filter applied to the wrong column, a metric read against the wrong grain, a dashboard interpreted confidently by someone who had never been told what its denominator was.

You cannot engineer your way out of that. The data team could keep making the plumbing better and the errors would keep arriving from upstream of the plumbing.

The constraint

I had no authority to mandate training. There was no budget line for education. And the audience considered data somebody else’s job, which is a reasonable position to hold if nobody has ever shown you otherwise.

What I did

I built the Data Academy as a chapter inside the data organisation, with dedicated people rather than as a volunteer effort by whoever had capacity. Volunteer training programmes die in the second quarter.

The design decision that mattered was certification rather than attendance. A training log records that someone sat in a room. A certification records that they can do the thing. Once you have a certification you can make specific roles require it, which converts a nice-to-have into a gate, and a gate is the only version of this that changes behaviour.

The curriculum was built from our own incidents. Every module existed because something had actually gone wrong. That made it specific, and it made it defensible when someone asked why they had to sit through it.

The result

More than a hundred people certified, and a measurable drop in data errors and data operation incidents over the period.

The second-order effect was the one I did not plan for. Once enough people could read the data themselves, the data team stopped being a bottleneck for basic questions, and the questions that did arrive got sharper. The team spent less time answering and more time building.

What I would do differently

Certify the managers first. I certified individual contributors while their managers stayed uncertified, and the old behaviour kept getting requested from above. An analyst who has learned to ask what the denominator is still has a director asking for the number by Friday. Training the people who make the requests would have been worth more than training the people who fulfil them, and it would have cost less because there are fewer of them.

I should have published the incident that motivated each module. We taught the lesson and withheld the story, out of a reasonable instinct not to embarrass anyone. The stories were the memorable part. Anonymised, they would have done more work than the slides.