001✓ copiedspec
CS-02 · Babbel GmbH
The first AI agents at Babbel, and who got them first
Shipping the first internal AI agent at Babbel, then customer-facing ones, by giving engineers the tooling first and running compliance with Legal as a partner.
- Organisation
- Babbel GmbH
- Period
- 2023-05 to 2026
- Engagement
- Employed role, not a client engagement
- Layers
- ai-context · agentic-systems · ai-enablement
0systemsAI agents in production before this work
- Organisation
- Babbel GmbH
- Period
- 2023-05
- Basis
- Count of AI agent systems running in production at Babbel at the start of the role. The internal application was the first, and the customer-facing agents followed it.
002✓ copieddetail
This was an in-house role, not a client engagement.
The situation
Babbel had no AI agents in production. It had, like every company in 2023 and 2024, a great deal of interest in having some, and a set of proposals that mostly started at the customer.
Starting at the customer is the obvious move and it is the wrong one. A customer-facing agent is the highest-stakes possible place to discover that your retrieval is poor, your evaluation harness does not exist, and nobody has decided who is accountable when it says something wrong. You find out in public, with a brand attached.
The constraint
Two real ones, and a third that was not obvious.
Legal had not been asked to form a view yet, and the default trajectory was that they would be asked at the end, as an approval gate on something already built. That relationship, once established, is very hard to fix.
The regulatory picture was moving underneath us. The EU AI Act was arriving with real obligations, and GDPR questions about sending customer data through third-party model providers had answers that needed to exist before anyone needed them.
The third constraint was cultural. Engineers had already started using these tools informally, because engineers do. Any policy written as though adoption had not started would be a policy about a company that did not exist.
What I did
Engineers first. I put AI agents in the hands of the engineering organisation before anything went near a customer. The reasoning is not generosity. Engineers are the population who will find the failure modes fastest, describe them precisely, and tolerate a rough edge while it gets fixed. A customer will do none of those three.
It also meant that by the time we shipped externally, the organisation had a working intuition for what these systems are bad at, which is not a thing you can transfer by writing it down.
The internal application first, then the customer-facing agents. Same logic, one layer up. The first agent solved an internal problem where the blast radius was a colleague’s afternoon. What we learned there is what made the customer-facing ones defensible.
Legal as a partner, not as a gate. I brought Legal in at the design stage and kept them there. Internal policy, EU AI Act obligations, and GDPR posture were worked out alongside the build rather than presented to Legal as a finished thing to bless.
The practical effect is that the answer to “can we do this” arrived while it was still cheap to change the design. The cultural effect mattered more: Legal stopped being the department that says no at the end, because they were never put in that position.
Evaluation before scale. An agent you cannot trace is an agent you cannot ship. Before anything customer-facing, the question was always what we would measure, how we would know it had quietly got worse, and who would be paged when it did.
The result
Babbel went from no AI agents in production to an internal agent and then customer-facing agents, with a compliance position developed alongside them rather than retrofitted.
The sequencing is the transferable part, and it is what I now sell. Engineers before internal users, internal users before customers, and Legal in the room from the first design review.
What I would do differently
I should have written the evaluation harness before the first agent, not alongside it. We got away with it because the first application was internal and low stakes. If the first one had been customer-facing, running without a proper eval harness for those early weeks would have been an unforced error, and I would have deserved the consequences.
I underestimated how quickly informal adoption would outrun the inventory. By the time we did a proper count of what was in use across engineering, the list was longer than anyone expected. That is the normal finding, and it is why the inventory now comes first in every governance engagement I run. I learned it the expensive way here.
Compliance scope needs stating precisely, and I was loose about it internally. “We are compliant” is not a sentence. Compliant with what, assessed by whom, as of when. The frameworks that actually applied to us were the EU AI Act and GDPR, and being casual in conversation about which regimes were in scope created confusion I had to walk back later.