108 emails: what an eval is and why to use real cases

ai-consulting

There were 25 emails from people asking not to be written to again, cancelling or hurling insults, and the AI would have replied to 24 of them, as kindly as you can imagine.

It never got to, and that's the story. I had been asked for an automatic email reply and in testing it was perfect: good tone, accurate, better written than a lot of humans on a Monday. It was one switch away from going out to answer on its own.

Before switching it on I did something that has nothing sophisticated about it. I took the 108 real emails that had come in over the previous two months and had it answer them as if it were already live, without sending anything. Almost all of them would have gone out with nobody reviewing them. Those 25 were among them, and of the ones carrying insults it didn't recognize a single one as abuse.

Inbox of emails asking to stop, with replies marked 'would send'

We didn't switch it on like that. The AI was doing its job well, it simply didn't know which conversation it was in, like when you're a teenager and think you know it all until you have to book a doctor's appointment or file a formal complaint, and you get tangled up.

Evals: what someone imagined versus what will actually arrive

The tests it was acing had been written by another AI, and that was the problem. An AI writes the cases it imagines, which are the reasonable ones: somebody asks for a price, somebody wants an appointment. Real people write tired, write angry, write "I've told you three times to stop emailing me", and none of those emails were on the list because whoever put it together would never have thought of them.

An eval is running your system against what it will really receive, and not against what somebody imagined it would receive.

18,000 cups of water

Taco Bell put an AI on drive-thru orders in more than 500 restaurants and it handled over two million of them. Until somebody ordered, very calmly, 18,000 cups of water, and the video went viral. After that the company said it would coach its franchises on when to use voice AI and when to monitor it and step in. McDonald's had been testing its own with IBM since 2021, in more than a hundred restaurants, and ended the test in 2024, after its own viral videos of tangled orders; the company only said it wanted to explore other solutions.

Receipt: cup of water times 18,000, order confirmed
Source: public incident, Taco Bell drive-thru AI, AI Incident Database 1274

Two million orders protected nobody from the order that wasn't in the script, because that one doesn't get tested with invented cases but with the ones that already reached you.

If you tell me you have no data to build evals, my answer is that you do, it's in your inbox, in your support chat or in your form history. Start with twenty real cases, the last twenty, without picking the nice ones. And before you switch on what you're building, ask yourself what it would have answered.

This article is part of the series from my AI Week 2026 talk: talk resources and slides. Next: A date doesn't need an agent, it needs a calendar.

Subscribe

Get posts about AI, development, and the solo founder journey.