How I work

The tool is not the problem.

Almost every business that calls me already has tools. Often too many. What is missing is the order around them: who decides what, what is allowed to happen automatically, and how you notice when something has quietly stopped. This page explains exactly that, without jargon.

Why this is not industry-specific

The same order, four times.

An online store, a customer support desk, a hiring search and a booking system. Four things with nothing in common, where the first move was the same every time. What never changes is the order.

My own business

Online store for grave decoration

Collect 16 months of real search demand first, then decide what is worth writing about.

Client project

Customer support for a jewellery brand

Analyse 660 real tickets first, then write the rules that answers follow.

Internship semester

Hiring for vehicle inspection engineers

Read what these people actually complain about in forums first, then write the job ad.

Client project

Booking system for a flight simulator

Check what a customer actually sees first, then believe the setting.

Measure what actually happens. Then automate.

The techniques

Nine habits I apply every single time.

None of it is magic and none of it is my invention. These are nine habits I picked up from mistakes that cost me money. They have names so they can be talked about.

Before building
01

Cold-start test

New chat, no history. Does it still work?

The most honest test for any process: start from zero, as if nobody remembered what this was about. Everything you have to re-tell by hand is a gap. The test costs ten minutes and shows exactly the places where something only works because you happen to know what it needs. Those are the same places where it stops the moment you are away for two weeks.

02

Autonomy tiers

How much rope the AI gets is settled up front.

Every task gets one of four tiers, sorted by four questions: Can it be undone? Who does it hit when it goes wrong? What does one failed attempt cost? Does it touch brand, money or law? Anything that reaches a customer stays at the most cautious tier until there are good reasons for something else. When torn, the more cautious tier wins.

03

Gap probe

Can the source even show a gap?

Many systems only return the fields that are actually filled once you move to their newer interface. That is cleaner and usually an improvement. For a completeness check it is fatal: built on that source, it can never find a gap by construction. And zero gaps is exactly the result you were hoping for, so nobody questions it. Before using a source for that kind of check, hold a case against it that you know is incomplete and see whether the gap shows up at all.

While building
04

Rules first

The rule decides. The model only writes it down.

The decision is made by a rule you have read and approved. The model receives it already made and phrases it. It never sees the rules and is never asked what the policy is. If no rule matches, the system says exactly that and hands over. That is the difference between a process you can check and one you have to believe.

05

Replay

Run it against your own history before it runs live.

Replay past cases with the outcome hidden, compare against what actually happened at the time, and read every disagreement one by one. Half of them are bugs in the rules. The other half are cases you once decided differently yourself, and that is exactly where the good rules come from. It costs a few days and not a single real customer.

06

Source receipt

Every agent names what it read. Whoever does not say, did not read.

A plan was meant to be torn apart by five independent AI reviewers. The text reached none of them, the hand-off was empty. Not one agent stopped. All five produced readable, well-written, confident verdicts on a document they had never seen. A model without context does not fall silent, it gets plausible. Since then every run has to show what actually arrived as input before anyone trusts the result.

In operation
07

Result check

Verify what comes out, not what is configured.

After every change the result gets verified, and specifically at the edge rather than in the middle. Errors sit almost always right at the boundary between two cases. This one habit would have saved me 26 days of silent outage and around 1,100 euros: the setting looked correct, the interface reported it back correctly, and still nothing reached the customer.

08

Freshness check

Monitor how old the data is, not whether the run went green.

Green does not mean anything is arriving. Three of my processes ran for weeks reporting success while writing not a single row. A watchdog also needs a third state: not just "running" and "broken" but "not measured at all". Without that third state every access error reports a fault that does not exist, and you end up adjusting a setting that was fine.

09

Alarm drill

You test an alarm by setting it off on purpose.

A watchdog that reports nothing looks exactly like a clean system. Both produce the same output: none. One of my own automated checks ran for six weeks as a program that did nothing and reported success anyway. It was found neither by a false alarm nor by a missing one, but because somebody tested the test. Monthly since then: build a throwaway input carrying a known violation and expect the alarm. Takes five minutes.

If you already automate things yourself: this is the part people skip, because it is no fun. It is also the reason processes keep running without them.

The question everyone gets stuck on

How much is the AI allowed to decide on its own?

That is not a matter of belief, it is a classification. Every task gets a tier, and the tier says how much rope it has. Drag the dial.

A0A1A2A3

Plan first, then approval

The AI proposes and shows exactly what it would do beforehand. It only runs after your yes. The normal case for anything touching money or public perception.

Stays off-limits at this tier
  • ✕Running without approval
  • ✕Several steps in a row

The classification runs on four questions: Can it be undone? Who does it hit when it goes wrong? What does one failed attempt cost? Does it touch brand, money or law? And when you are torn between two tiers, the more cautious one always wins.

Why the checking gets built in

Three failures that looked like success from outside.

All three really happened, to me or to my clients. They have one thing in common: there was no error message. The setting was saved, the process reported success, and still nothing happened.

26
days unnoticed

The setting was right. The result was not.

In my own store, nobody could complete a basket over 99 euros for 26 days. The setting looked correct, and the interface even reported it back correctly. What revealed it was looking at what a customer actually gets shown. Around 1,100 euros of revenue and 13 blocked baskets later. Since then every change is verified by its result, not by its setting, and specifically at the edge rather than in the middle.

3
dead processes

Green does not mean it is running.

Three processes reported success for weeks while writing not a single row. The status was green, the result was empty. A process only counts as active for me once something is verifiably arriving, not once it has been started. What gets monitored now is how old the data is, not whether the run completed.

0
times triggered

Saved, and still doing nothing.

At one client, the booking system accepted the restriction "+2 Tage", saved it, and silently ignored it, because internally that field only understands English. The vendor’s own German documentation showed the German example. The restriction was never active, and nobody would ever have noticed.

The blueprint

Rules decide. The AI only writes it down.

This is the difference between a system you trust and one you switch off again after two weeks. The decision is made by a rule you have read and approved. The AI receives that decision already made and only phrases it. It never sees the rules and is never asked what the policy is. If no rule matches, the system says exactly that and hands over to a human. An empty result is a correct result.

Hands on

Build the process and run it.

Build on the left, watch the result on the right. The same request goes in every time. Switch between the two ways of building it and hit Run a few times. The difference ends up in the log, not in a claim of mine.

Input: return on day 34, customer since 2023

Request comes inCheck the rulesAI phrases itHuman checksReply goes out

The nodes can be dragged. The connections hold.

Log
  • Nothing has run yet. Hit Run.
The problem nobody quite names

Explaining it all over again, every time.

You open a new chat and type it all out again: what the company does, how you write, what was decided last time, which rules apply. It works. Then the conversation gets long and the answers get worse. And the one prompt that was genuinely good sits in a thread you cannot find again.

That is not you using it wrong. The AI gets whatever you happen to remember in that moment. I make sure it always gets exactly what this one step needs instead, without you assembling it. How that is built underneath is my problem, not yours. What you notice is this: you stop explaining from scratch, the answers stay as good as they were, and they stay that way even when somebody else is asking.

Chat→Saved prompt→A system that feeds itself

Only climb to the next rung once the one below genuinely runs. For something you do twice a year you do not need a system, a saved prompt will do. If I talk you out of it on the intro call, that is an answer too, and it costs you nothing.

Context, loaded sparingly

Do not know everything. Know the right thing.

Every square is a piece of what your business knows: price list, return rules, brand voice, supplier contacts, holiday rota. A single step almost never needs more than a handful of them. Flip it.

Loaded
42,000tokens
Actually used
3,200tokens

Pouring everything in feels thorough and is the opposite. The more unrelated material rides along, the likelier the model reaches for the wrong thing, and you pay for every square, grey ones included.

Token: the chunks a language model breaks text into. Roughly one token per four characters. It is what you are billed for, and at some point the window is full.

Intro call

Let's talk about one process for 25 minutes.

Free, no strings, by video or in person in Berlin. You get an honest assessment at the end, including if that assessment is that you do not need anyone for this.

01You describe one task that costs you time every week
02We work out whether it is worth automating and roughly what it costs
03You get a recommendation, even when it argues against a project

My commitment: if a handed-over process does not run the way we described it beforehand, I rebuild it until it does. No new invoice, and no argument about whose fault it was.

Pick a slot.

Choose a time that suits you. The few questions in the form are there so I arrive prepared.

Clicking loads the Cal.com calendar, which sends data to Cal.com. Nothing is loaded before you click. Privacy