The tool is not the problem.
Almost every business that calls me already has tools. Often too many. What is missing is the order around them: who decides what, what is allowed to happen automatically, and how you notice when something has quietly stopped. This page explains exactly that, without jargon.
The same order, four times.
An online store, a customer support desk, a hiring search and a booking system. Four things with nothing in common, where the first move was the same every time. What never changes is the order.
Online store for grave decoration
Collect 16 months of real search demand first, then decide what is worth writing about.
Customer support for a jewellery brand
Analyse 660 real tickets first, then write the rules that answers follow.
Hiring for vehicle inspection engineers
Read what these people actually complain about in forums first, then write the job ad.
Booking system for a flight simulator
Check what a customer actually sees first, then believe the setting.
Measure what actually happens. Then automate.
Nine habits I apply every single time.
None of it is magic and none of it is my invention. These are nine habits I picked up from mistakes that cost me money. They have names so they can be talked about.
Cold-start test
New chat, no history. Does it still work?
The most honest test for any process: start from zero, as if nobody remembered what this was about. Everything you have to re-tell by hand is a gap. The test costs ten minutes and shows exactly the places where something only works because you happen to know what it needs. Those are the same places where it stops the moment you are away for two weeks.
Autonomy tiers
How much rope the AI gets is settled up front.
Every task gets one of four tiers, sorted by four questions: Can it be undone? Who does it hit when it goes wrong? What does one failed attempt cost? Does it touch brand, money or law? Anything that reaches a customer stays at the most cautious tier until there are good reasons for something else. When torn, the more cautious tier wins.
Gap probe
Can the source even show a gap?
Many systems only return the fields that are actually filled once you move to their newer interface. That is cleaner and usually an improvement. For a completeness check it is fatal: built on that source, it can never find a gap by construction. And zero gaps is exactly the result you were hoping for, so nobody questions it. Before using a source for that kind of check, hold a case against it that you know is incomplete and see whether the gap shows up at all.
Rules first
The rule decides. The model only writes it down.
The decision is made by a rule you have read and approved. The model receives it already made and phrases it. It never sees the rules and is never asked what the policy is. If no rule matches, the system says exactly that and hands over. That is the difference between a process you can check and one you have to believe.
Replay
Run it against your own history before it runs live.
Replay past cases with the outcome hidden, compare against what actually happened at the time, and read every disagreement one by one. Half of them are bugs in the rules. The other half are cases you once decided differently yourself, and that is exactly where the good rules come from. It costs a few days and not a single real customer.
Source receipt
Every agent names what it read. Whoever does not say, did not read.
A plan was meant to be torn apart by five independent AI reviewers. The text reached none of them, the hand-off was empty. Not one agent stopped. All five produced readable, well-written, confident verdicts on a document they had never seen. A model without context does not fall silent, it gets plausible. Since then every run has to show what actually arrived as input before anyone trusts the result.
Result check
Verify what comes out, not what is configured.
After every change the result gets verified, and specifically at the edge rather than in the middle. Errors sit almost always right at the boundary between two cases. This one habit would have saved me 26 days of silent outage and around 1,100 euros: the setting looked correct, the interface reported it back correctly, and still nothing reached the customer.
Freshness check
Monitor how old the data is, not whether the run went green.
Green does not mean anything is arriving. Three of my processes ran for weeks reporting success while writing not a single row. A watchdog also needs a third state: not just "running" and "broken" but "not measured at all". Without that third state every access error reports a fault that does not exist, and you end up adjusting a setting that was fine.
Alarm drill
You test an alarm by setting it off on purpose.
A watchdog that reports nothing looks exactly like a clean system. Both produce the same output: none. One of my own automated checks ran for six weeks as a program that did nothing and reported success anyway. It was found neither by a false alarm nor by a missing one, but because somebody tested the test. Monthly since then: build a throwaway input carrying a known violation and expect the alarm. Takes five minutes.
If you already automate things yourself: this is the part people skip, because it is no fun. It is also the reason processes keep running without them.
How much is the AI allowed to decide on its own?
That is not a matter of belief, it is a classification. Every task gets a tier, and the tier says how much rope it has. Drag the dial.
Plan first, then approval
The AI proposes and shows exactly what it would do beforehand. It only runs after your yes. The normal case for anything touching money or public perception.
- ✕Running without approval
- ✕Several steps in a row
The classification runs on four questions: Can it be undone? Who does it hit when it goes wrong? What does one failed attempt cost? Does it touch brand, money or law? And when you are torn between two tiers, the more cautious one always wins.
Three failures that looked like success from outside.
All three really happened, to me or to my clients. They have one thing in common: there was no error message. The setting was saved, the process reported success, and still nothing happened.
The setting was right. The result was not.
In my own store, nobody could complete a basket over 99 euros for 26 days. The setting looked correct, and the interface even reported it back correctly. What revealed it was looking at what a customer actually gets shown. Around 1,100 euros of revenue and 13 blocked baskets later. Since then every change is verified by its result, not by its setting, and specifically at the edge rather than in the middle.
Green does not mean it is running.
Three processes reported success for weeks while writing not a single row. The status was green, the result was empty. A process only counts as active for me once something is verifiably arriving, not once it has been started. What gets monitored now is how old the data is, not whether the run completed.
Saved, and still doing nothing.
At one client, the booking system accepted the restriction "+2 Tage", saved it, and silently ignored it, because internally that field only understands English. The vendor’s own German documentation showed the German example. The restriction was never active, and nobody would ever have noticed.
Rules decide. The AI only writes it down.
This is the difference between a system you trust and one you switch off again after two weeks. The decision is made by a rule you have read and approved. The AI receives that decision already made and only phrases it. It never sees the rules and is never asked what the policy is. If no rule matches, the system says exactly that and hands over to a human. An empty result is a correct result.
Build the process and run it.
Build on the left, watch the result on the right. The same request goes in every time. Switch between the two ways of building it and hit Run a few times. The difference ends up in the log, not in a claim of mine.
Input: return on day 34, customer since 2023
The nodes can be dragged. The connections hold.
- Nothing has run yet. Hit Run.
Explaining it all over again, every time.
You open a new chat and type it all out again: what the company does, how you write, what was decided last time, which rules apply. It works. Then the conversation gets long and the answers get worse. And the one prompt that was genuinely good sits in a thread you cannot find again.
That is not you using it wrong. The AI gets whatever you happen to remember in that moment. I make sure it always gets exactly what this one step needs instead, without you assembling it. How that is built underneath is my problem, not yours. What you notice is this: you stop explaining from scratch, the answers stay as good as they were, and they stay that way even when somebody else is asking.
Only climb to the next rung once the one below genuinely runs. For something you do twice a year you do not need a system, a saved prompt will do. If I talk you out of it on the intro call, that is an answer too, and it costs you nothing.
Do not know everything. Know the right thing.
Every square is a piece of what your business knows: price list, return rules, brand voice, supplier contacts, holiday rota. A single step almost never needs more than a handful of them. Flip it.
Pouring everything in feels thorough and is the opposite. The more unrelated material rides along, the likelier the model reaches for the wrong thing, and you pay for every square, grey ones included.
Token: the chunks a language model breaks text into. Roughly one token per four characters. It is what you are billed for, and at some point the window is full.
Let's talk about one process for 25 minutes.
Free, no strings, by video or in person in Berlin. You get an honest assessment at the end, including if that assessment is that you do not need anyone for this.
My commitment: if a handed-over process does not run the way we described it beforehand, I rebuild it until it does. No new invoice, and no argument about whose fault it was.
Pick a slot.
Choose a time that suits you. The few questions in the form are there so I arrive prepared.
Clicking loads the Cal.com calendar, which sends data to Cal.com. Nothing is loaded before you click. Privacy