Three engagements. Each one ends with something you can check and a decision that is yours: whether to spend more money, add more people, or give the system more autonomy.
Two weeks. We help you choose the workflow, map it as it runs today, and measure it.
Ends with: An agreed baseline and a recommendation
We redesign the workflow, build the system, connect it to the tools you use, and test it.
Ends with: A working system and its test results
Your team uses it on live work, with training, monitoring, and named owners.
Ends with: Results against the baseline and your decision
| Step | Stage | Question answered | What must be true to move on |
|---|---|---|---|
| Scope | 1 Select | Which work matters most, and is it ready? | Named owner, measurable outcome, usable data |
| Scope | 2 Diagnose | How does the work really happen, and where is the bottleneck? | Baseline agreed with the business owner |
| Build | 3 Design | What is the lowest level of automation that solves it? | Design approved by the owner |
| Build | 4 Prove | Does it meet the acceptance criteria agreed in advance, and beat the non-AI option? | Criteria met on the test set, zero unacceptable errors |
| Run | 5 Pilot | Does it work for real users under enforceable controls? | Used independently, no unresolved critical incidents |
| Run | 6 Decide | Did the value arrive, and what happens next? | Scale, iterate, or stop, with reasons |
Each gate is a decision point. You can stop after any stage, with the evidence in hand.
During Scope we measure how the work performs today and agree on that baseline with the person who owns the result. Before any testing, we propose acceptance criteria from it: what counts as good enough, and which errors are unacceptable. You know the work and the risk. We know what can be tested. Neither side sets the standard alone.
What a pass means. A pass shows the system met the agreed standard on the agreed test set. It is not a guarantee that it will never be wrong on live work. AI output can be inaccurate, which is why our builds keep a person in review on consequential actions, and why monitoring continues through Run.
Every step is tested against four questions before any AI is considered. Most stop at the first or second. Across our four documented builds, 59% of steps run on rules.
What counts as good enough is written down before testing starts, so the test result means something.
Business, AI operations, engineering, and risk each own a part. Any of them can pause the system.
| When | What lands on your desk |
|---|---|
| End of week 2 | Ranked shortlist, map of the work, agreed baseline |
| Design | Signed-off design, with each step marked rules, AI-assisted, or human |
| Prove | Test results against the agreed acceptance criteria |
| Pilot | Your team running it, with training and monitoring in place |
| Decide | Results against the baseline and a decision memo |
A working example of what Scope produces. Five tools take you from how people say the work happens to how it actually happens, and then to a step-by-step call on rules, AI, or a person.
Half an hour with Andrew. If AI is not the answer, you will hear that too.