Copy a link to this pageThe canonical URL, on the clipboard↵ run
Copy this page as a Markdown linkTitle and URL, ready to paste into notes↵ run
Switch between the dark and paper themesFollows your system setting until you change it here↵ run
Clear what this palette remembersForgets which pages you open most. Stored only in this browser↵ run
Print this pageThe print stylesheet drops the navigation and expands links↵ run
Loading the indexIf this does not finish, the full index is a page.
move open> actions? helpAll shortcuts

Clemence W. Chee

I build the layer AI actually runs on.

The context, the retrieval, the governance, and the agents on top of it. Interim CTO when the product is the AI system. Chief Data and AI Officer when the scale is there and the operating model is not. Forward deployed engineer when what is missing is someone to build it in your codebase.

I do not run general platform engineering, infrastructure, security or SRE.

whoami → clemence w. chee · interim CTO / CDAO / FDE

Built insideHelloFresh SEBabbel GmbHRocket Internet SE

10downminutestime from question to delivered insightBabbel GmbH2023-05 to 2026basis
Organisation
Babbel GmbH
Period
2023-05 to 2026
Baseline
two to four weeks
Basis
Elapsed time from a business question being asked to the answer being available, before and after the platform work. Range of five to ten minutes after; two to four weeks before. Measured on recurring business questions, not on one-off analyses.
more than$150mUSDdata risk exposure brought under managementHelloFresh SE2020-06 to 2023-05basis
Organisation
HelloFresh SE
Period
2020-06 to 2023-05
Basis
Exposure carried by data assets that were previously unowned and untraced, brought under management by mapping lineage end to end, enabling an audit trail, consolidating to a single source of truth with verified and named-owner data assets, and enriching the metadata across them. Stated as exposure placed under management, not as loss avoided.
0systemsAI agents in production when I arrived. I shipped the first.Babbel GmbH2023-05basis
Organisation
Babbel GmbH
Period
2023-05
Basis
Count of AI agent systems running in production at Babbel at the start of the role. The internal application was the first, and the customer-facing agents followed it.

The ontology

This is the provenance graph an AI system needs to be auditable: twelve entity types and the relations between them, from the owner of a dataset through to the outcome a decision produced. It is the model behindAI Control Plane, and it is how I work whether or not that software is in the room. Read it bottom to top and the thesis reads off the picture. Most AI programmes fail one layer below where they get blamed: the agent that hallucinates gets blamed on the model, and the cause is usually that nobody curated the knowledge it could reach.

The model has its own page at /map, versioned and published as data at /map.json, so it can be cited rather than paraphrased.

Most AI programmes fail one layer below where they get blamed. Twelve capabilities across four layers, joined by fourteen arrows that each point from a capability to the thing it makes possible. The same map is written out in full in the list below.Owner is accountable for the metric definition.01Dataset is indexed into the retrievable corpus.02Feature records what it was derived from.03Dataset resolves the metric.04Metric is governed by policy.05Knowledge grounds the model at query time.06Lineage makes every agent hop traceable.07Lineage records what the model was trained on.08Policy constrains the prompt before execution.09Model produces the decision.10Agent action is measured as an outcome.11Model performance is measured as an outcome.12Dataset supplies the context a prompt is given.13PromptAn unversioned promptmakes every downstreamdecision unreproducible.DecisionA decision nobody recordedcannot be audited andcannot be improved.OutcomeDecisions nobody measuredare opinions with a logline.PolicyA control that depends ona person remembering isnot a control.ModelA model without an evalharness degrades quietly,and you find out from acustomer.AgentAn agent you cannot traceis an agent you cannotship.MetricA semantic layer withoutnamed owners becomes asecond source of truth.KnowledgeYour model is only as goodas the corpus it canreach, and nobody curatesa corpus by accident.LineageWithout lineage everyincident becomes anunbounded investigation.OwnerAn asset owned by "thedata team" is an assetowned by nobody.DatasetA dataset without acontract is a sharedmutable variable.FeatureFeatures computed twicediverge, and the seconddefinition is always theone in production.

Read the picture from the bottom up. An arrow points from a capability to the thing it makes possible. The numbered squares key into the edge list below.

04AI enablementThe layer that gets the blame.
03Agentic systemsThe layer everyone photographs.
02AI contextThe layer most programmes skip.
01Data foundationThe layer where the failure usually started.
Read this map as a list

The same map in text. Each node states what the term means and what breaks without it. Each edge names one failure mode. Every figure carries the organisation, the period, the basis of measurement, and the case study behind it. Figures without a stated basis are named and withheld rather than rounded into something safer.

01 Data foundation

Data foundation is the layer that decides whether a number can be trusted at all: named ownership, tested contracts, and metered cost.

The layer where the failure usually started.

Owner

Owner is the named person accountable for a data asset, not the team it sits in.

Failure modeAn asset owned by "the data team" is an asset owned by nobody.

EvidenceThe Data Academy: teaching a company to stop breaking its own data·AI governance for a Berlin deep-tech manufacturer·Putting $150M of data risk on the register

Dataset

Dataset is a governed collection with a declared schema, an owner, and a contract with the teams that consume it.

Failure modeA dataset without a contract is a shared mutable variable.

  • 95%

    availability of the standardised revenue data models

    Babbel GmbH·2023-05 to 2026·Basis: Availability of the standardised revenue models against data contracts written per consuming use case, enforced by computational governance rather than by manual review. Measured as the share of contract obligations met.·Seven data product managers, and a company that turned EBITDA positive

Withheld2 figures claimed against this capability (product margin under the supply chain models; cost saved by the finance and procurement data systems) stay off the page until the basis of measurement is written down. A number without a method is a number a buyer can take apart in thirty seconds.

EvidenceFrom one month to one week·Seven data product managers, and a company that turned EBITDA positive·Putting $150M of data risk on the register

Feature

Feature is a derived signal computed from datasets and reused across models, with its definition versioned.

Failure modeFeatures computed twice diverge, and the second definition is always the one in production.

  • 60%down

    data platform storage cost

    Babbel GmbH·2023-05 to 2026·Basis: Reduction in cloud storage spend, achieved by sunsetting the monolith and the legacy systems around it rather than running them alongside the replacement, and by migrating onto a stack that bills storage separately from compute. Measured as absolute storage spend against the pre-migration run rate.·Seven data product managers, and a company that turned EBITDA positive

Withheld2 figures claimed against this capability (RETIRED — reattributed to babbel-maintenance-cost; annual cost saved by platform consolidation) stay off the page until the basis of measurement is written down. A number without a method is a number a buyer can take apart in thirty seconds.

EvidenceSeven data product managers, and a company that turned EBITDA positive·From one month to one week

02 AI context

AI context is the layer that gives a model the vocabulary, the sources, and the provenance of the business it is answering about.

The layer most programmes skip.

Metric

Metric is an agreed business definition: one revenue, one lifetime value, each owned and versioned.

Failure modeA semantic layer without named owners becomes a second source of truth.

  • 1weekdown

    time from question to delivered insight, from one month

    HelloFresh SE·2018-12 to 2020-06·Basis: Elapsed time from an operations question being asked to the answer being delivered, before and after the architecture rebuild. Measured on the recurring operational questions the BI team handled, not on one-off analyses.·From one month to one week

EvidenceFrom one month to one week·AI governance for a Berlin deep-tech manufacturer·Putting $150M of data risk on the register·AI Control Plane

Knowledge

Knowledge is the retrievable corpus a model can reach at query time, with permissions that hold at retrieval rather than at the interface.

Failure modeYour model is only as good as the corpus it can reach, and nobody curates a corpus by accident.

No published figure yet. This capability is current work, and the evidence is the engagement and the software underneath it.

EvidenceAI governance for a Berlin deep-tech manufacturer·AI Control Plane

03 Agentic systems

Agentic systems is the layer where software acts on its own: traceability, policy enforcement, and evaluation of every autonomous step.

The layer everyone photographs.

04 AI enablement

AI enablement is the layer where people and money meet the technology: literacy, product management, and the standing operating model.

The layer that gets the blame.

Prompt

Prompt is a versioned instruction bound to the context it was given and the decision it produced.

Failure modeAn unversioned prompt makes every downstream decision unreproducible.

EvidenceThe Data Academy: teaching a company to stop breaking its own data·Seven data product managers, and a company that turned EBITDA positive

Decision

Decision is the recorded output of an agent or a person, carrying the context that produced it.

Failure modeA decision nobody recorded cannot be audited and cannot be improved.

  • 7peopleup

    data product managers, from 0 at role start

    Babbel GmbH·2023-05 to 2024·Basis: Filled data product management positions in the data organisation, counted at the end of the build-out. The function did not exist before the role started.·Seven data product managers, and a company that turned EBITDA positive

  • 6months

    time to the first launched data product

    Babbel GmbH·2023-05 to 2023-11·Basis: Elapsed months from the creation of the data product management function to the first product release.·Seven data product managers, and a company that turned EBITDA positive

  • more than2,000hours per yeardown

    manual work removed by the first data product

    HelloFresh SE·2018-12 to 2020-06·Basis: Annualised manual hours removed by the automation the data product replaced, counted from the task inventory it was built against.·From one month to one week

EvidenceSeven data product managers, and a company that turned EBITDA positive·From one month to one week

Outcome

Outcome is what the decision actually changed, measured, and fed back into the features that produced it.

Failure modeDecisions nobody measured are opinions with a log line.

  • 20peopleup

    people in the business intelligence centre of excellence

    HelloFresh SE·2017-04 to 2018-12·Basis: Full-time employees in the business intelligence centre of excellence at its largest point during the role.·From one month to one week

  • 4peopleup

    people in the Australian business intelligence team, from 0 at role start

    HelloFresh Australia·2015-11 to 2017-04·Basis: Full-time employees hired into the Australian business intelligence team. There was no such team before the role started.·From one month to one week

EvidenceFrom one month to one week·AI governance for a Berlin deep-tech manufacturer

14 edges, each one a failure mode

The nodes above are vocabulary. The sentences below are the argument. Each number matches a numbered edge in the diagram.

  1. 01OwnerMetricOwner is accountable for the metric definition.
  2. 02DatasetKnowledgeDataset is indexed into the retrievable corpus.
  3. 03FeatureLineageFeature records what it was derived from.
  4. 04DatasetMetricDataset resolves the metric.
  5. 05MetricPolicyMetric is governed by policy.
  6. 06KnowledgeModelKnowledge grounds the model at query time.
  7. 07LineageAgentLineage makes every agent hop traceable.
  8. 08LineageModelLineage records what the model was trained on.
  9. 09PolicyPromptPolicy constrains the prompt before execution.
  10. 10ModelDecisionModel produces the decision.
  11. 11AgentOutcomeAgent action is measured as an outcome.
  12. 12ModelOutcomeModel performance is measured as an outcome.
  13. 13DatasetPromptDataset supplies the context a prompt is given.The thesis edge. It runs the full height of the stack because that is the distance between where the failure starts and where the blame lands.
  14. 14OutcomeOwnerOutcome feeds back to the owner who acts on it.The return edge. The stack is a loop, and the funding decisions taken at the top decide whether the bottom holds.

Map version 1.0.0. Last updated . The model is also published as data at /map.json.

What I actually believe

  1. Data work is worth doing when it changes a decision.

    That is the test I apply to everything I build, and most data work does not pass it. A report nobody acts on is a cost with a dashboard attached.

  2. An agent you cannot trace is an agent you cannot ship.

    The interesting engineering is rarely the prompt. It is the context the system can reach, the permissions that hold at query time, and the evals that tell you when it quietly got worse.

  3. Governance belongs in the pipeline, not in a review meeting.

    A control that depends on a person remembering is not a control. Engineers should be able to do the right thing without asking permission.

  4. The person who designs the system should be able to build it.

    An architecture nobody on the team can implement is a diagram. I deploy into the codebase and ship the first working version myself, because the decisions that matter get made where the thing gets built.

  5. Every number should carry its measurement basis.

    If I cannot tell you how a figure was calculated, I do not put it on a slide. That applies to my own numbers on this site, which is why some of them are missing.

Track record

Six engagements, written at length, including the parts that did not work and the numbers I cannot fully source.

  1. HelloFresh SE2020-06 to 2023-05

    Putting $150M of data risk on the register

    Founding a data management department at HelloFresh, and the difference between exposure you have mitigated and exposure you can finally see.

  2. HelloFresh SE2017-04 to 2020-06

    From one month to one week

    Rebuilding the architecture and the team behind operational decisions at HelloFresh, and shipping the architecture before the ownership model.

All case studies

What I ship

Five systems I have built. Each one declares its status and prints its own known defects, including the ones that are parked and the reason they are parked.

  1. parked

    AI Control Plane

    A semantic control layer that makes every AI decision traceable, policy-enforced and auditable, without changing the code that calls the model.

    Policy-as-code · Knowledge graphs · Audit trails · LLM gateways

  2. live

    Prompt Master

    An interactive trainer for prompt engineering. Practise named frameworks against live feedback and see which structures change a model's answer.

    React · Node.js · Natural language processing

  3. live

    PeakMind

    A habit, task and focus manager with AI assistance, covering focus sessions, habit formation and shared workspaces in one interface.

    React · Node.js · Real-time analytics

  4. parked

    Finance AI

    A financial analysis prototype combining machine learning with conventional metrics. Parked. The deployed site still serves a default template title.

    Python · React · Node.js · TensorFlow

  5. live

    AgentFlow

    AgentFlow monitors AI agent systems, detects failures and sends alerts. A TypeScript monorepo of six packages, MIT licensed, with a documentation site.

    TypeScript · Node.js · OpenTelemetry · React · Docusaurus

All systems, with their defects

Ways in

Five shapes. Most people start with the teardown, because it is the one that ends in a decision rather than a commitment.

  • AI readiness teardown

    A three-week fixed-fee assessment of your data and AI stack that ends in a decision, with a written verdict you keep whatever you decide next.

    Duration
    3 weeks
    Commitment
    Roughly 8 days of my time, spread across three weeks
  • Forward deployed engineer

    I join your engineering team, work in your repository, and ship the AI system myself. Six to twelve weeks, ending in something running in production.

    Duration
    6 to 12 weeks
    Commitment
    3 to 4 days a week, inside your engineering team
  • Interim CTO, or Chief Data & AI Officer

    I take the seat, own the layer AI runs on end to end, hire the person who replaces me, and spend the last month handing over.

    Duration
    6 to 12 months
    Commitment
    2 to 3 days a week, inside your leadership team
  • Data product management, installed

    Twelve to sixteen weeks to build the data product management function: the role, the hiring loop, the intake process, proven by shipping real products.

    Duration
    12 to 16 weeks
    Commitment
    2 days a week
  • Data governance and AI readiness

    Eight to twelve weeks to a written position that holds up in an audit, a data room, or a regulator conversation, with controls built into the pipelines.

    Duration
    8 to 12 weeks
    Commitment
    2 days a week

Start a conversation