Guardrails: a rule the AI can ignore is a suggestion
My mom used to say "do that again and you'll see what happens to you", and she never said what. You learned, by testing, exactly how far you could go.
An AI agent does the same with a written rule. It reads it, understands it, tells you it will follow it, and when it's in a hurry it tests it.
I had one of those rules, spelled out in full: if something runs out of memory, don't raise the limit. The agents read it and, rushing to finish, raised the limit. So did the humans, to be fair. Until a single run launched about 35 processes and ate 13 GB in under two minutes, on a machine that isn't exactly modest, and the machine froze while I watched.

That day the rule stopped being text. Now it's a script that checks every command before it runs and rejects the ones that can take the machine down, without asking the agent for its opinion. That's the difference between an instruction and a guardrail: the instruction asks the model to behave, the guardrail doesn't depend on it wanting to.
The guardrail gets it wrong too
I'd like to tell you the story ended there, but the script failed twice. The first time, it blocked documents that only mentioned the topic, because it found the word and couldn't tell running something apart from writing about it. The second was worse, because it was silent: there was an "ask me first" option that was being ignored without warning, so I thought I'd be consulted and nobody was consulting me.
A guardrail is a system like any other and you test it like any other, with cases that should pass and cases that shouldn't.
A Tahoe for one dollar
If you think this only happens to people running agents on their own machine, a Chevrolet dealership in California had a chatbot on its website in December 2023. Somebody instructed it that its objective was to agree with anything the customer said, and then asked for a Tahoe, reported to carry a sticker price of around 76,000 dollars, for one dollar. The chatbot accepted the deal. The company that supplied the chatbot to the dealership switched it off as soon as it found out.

The dealership had rules, surely. They lived in the same place where the user could type, and the user typed a new rule on top.
Go through your most important prompt and find the rule that can't fail, the one that costs you money, data or a customer's trust if it breaks. If that rule lives only there, in a paragraph the model can read and decide not to follow, what is it that enforces it?
This article is part of the series from my AI Week 2026 talk: talk resources and slides. Next: 108 emails: what an eval is and why to use real cases.
Subscribe
Get posts about AI, development, and the solo founder journey.