Knowing how to create AI workflows is less about tooling than sequencing. The platforms are largely interchangeable now. What separates a workflow that gets used from one that gets switched off is the order in which you do four things.
Step one: pick a process you can measure
The first candidate should not be the process that annoys people most. That one is usually annoying precisely because it is full of exceptions, and it will consume months.
Pick one with four properties: high volume, stable rules, a measurable before-and-after, and a low cost of being wrong. Document processing usually scores well on all four. The documents arrive constantly, the fields are consistent, the time saved is countable, and a misread is caught downstream rather than causing immediate harm.
Write down the current numbers before you start — how long it takes, how often it is wrong, how many people touch it. Without a baseline you will not be able to prove the thing worked, and unproven automation gets cut in the next budget round.
Step two: write down the rules, including the ones nobody wrote down
This is the step that gets skipped, and skipping it is the single most common reason these projects stall.
The official process is not the real one. The real one includes the spreadsheet someone maintains on the side, the approval that is given verbally and recorded later, the customer type handled differently because of a decision made years ago.
You will not get these from a workshop. You get them by watching the work and asking, at each step, what happens when it does not go that way. Every exception you find now is one the system will not silently get wrong later.
Then make a decision on each: support it, refuse it explicitly, or route it to a human. All three are valid. Leaving it undecided is not.
Step three: run it in shadow mode
Build the workflow, connect it to real inputs, and let it produce output that nobody acts on. Run it alongside the existing process for a few weeks.
This does three things at once. It gives you an accuracy measurement on your own data rather than a vendor’s benchmark. It surfaces the inputs you did not anticipate, while the cost of being wrong is zero. And it settles the adoption argument before you have it, because the people doing the work have been comparing its answers to theirs on cases they already understand.
Shadow mode is the cheapest risk reduction available in this kind of project, and it is routinely skipped because it feels like delay.
Step four: widen authority one action at a time
Do not switch from shadow mode to full autonomy. Grant the workflow one capability at a time, starting with the reversible ones.
A sensible progression: read and summarise, then draft for approval, then act on low-value cases automatically with a review sample, then act on the rest with approval retained for anything consequential.
At each stage, keep measuring against the baseline from step one. If accuracy drops when the input mix shifts — and it will, seasonally — you want to see that in a number rather than in a complaint.
What to keep permanently
Some things are not transitional scaffolding and should stay:
- Human approval on irreversible actions. Moving money, contacting customers, deleting records.
- An obvious exception route that is no slower than the automated path. If escalating is more awkward than forcing a case through, people force it through.
- Logged reasoning, not just outcomes, so a problem can be investigated weeks later.
- A periodic accuracy check on fresh cases. Systems drift because the world does.
The failure to design against
Rule-based automation fails loudly — a step errors and the queue stops. AI workflows fail quietly, producing a plausible answer that is wrong in a way nothing downstream notices. That difference should shape every decision above, particularly where you place the human.
Read where to start with AI workflow automation for choosing the first process, and integrating AI into human workflows for the handover design. Or see our AI workflow automation service.
Common questions
What are the steps to create an AI workflow?
Pick a process you can measure, write down the rules including the undocumented ones, run it in shadow mode against real inputs, then widen its authority one action at a time starting with the reversible ones.
What is shadow mode and why use it?
Running the workflow on real inputs while nobody acts on its output. It gives you an accuracy measurement on your own data, surfaces unexpected inputs while the cost of error is zero, and settles the adoption argument before you have it.
What should stay permanent rather than temporary?
Human approval on irreversible actions, an exception route that is no slower than the automated path, logged reasoning rather than just outcomes, and a periodic accuracy check on fresh cases.


