Omegaswift
AI

What a Pilot Should Prove Before You Roll It Out

A pilot has one job, which is to produce a reason to stop or continue that you agreed on before you saw any output. Most of them settle for everybody liked it.

The Omegaswift engineering teamAI and product engineering9 min read

Write the Success Test First

Decide in writing what result would make you roll this out, and what result would make you stop. Do it before the first configuration call. Get the person who will pay for the rollout to agree to those two sentences, and store them somewhere you cannot quietly edit later.

The reason is unflattering and it holds anyway. Once you have seen the output, your sense of what counts as good moves to accommodate it. Everybody involved has spent weeks on this by then. Nobody wants to be the person who says the bar was meant to be higher, and if the bar was never written down there is nothing to say it against.

The test does not need to be clever. It needs to be specific enough that two people reading it a month later would agree on whether it was met. Anything vaguer will be met.

A workable test names the work, the measure and the comparison. Something like this: on the invoices that arrive as scanned documents, the supplier and the total come out matching the paperwork more often than the current process manages, and a person spends less time per invoice checking the result. Dull to read. Almost impossible to argue about afterwards.

What a Pilot Is Actually For

A pilot answers a much narrower question than most buyers expect. The technology works somewhere, on somebody's data, with examples somebody chose. What you still do not know is whether it works here: on your records, inside your process, with your staff, at the volume you actually run on a busy week.

That is why a pilot on tidied sample data proves nothing. The interesting failures live in your messy cases. The customer who exists twice under two spellings. The scanned document that came in sideways. The policy page that contradicts the other policy page. Exclude those and you have watched the vendor's demo a second time with your own logo on it.

Volume matters for the same reason. Something that behaves well for one attentive user can behave very differently when a department leans on it at once, and the cost per request only becomes real at that point too.

The Pilot That Quietly Fails

A pilot that fails loudly is useful. Something breaks, everybody sees it, you learn something and you stop. The dangerous one runs like this: it gets set up, a few people use it for a fortnight, nobody complains, usage tails off, and at the end everybody agrees it was promising and asks for more time.

The signature is that nobody can say what changed. Ask what work stopped being done by hand. Ask which specific decisions came out differently. Ask what would break if you switched it off this afternoon. When the answers are vague, the honest reading is that nothing happened.

Usage decay is the clearest early signal and the easiest thing to check. Look at who used it in the first week and who was still using it in the last one. People abandon tools that are not helping them silently, without filing a complaint, because complaining is more work than going back to the old way.

Listen to the language in the closing meeting as well. Promising. Exciting. Early days. A good foundation to build on. Those words turn up when there is no result to describe. A pilot that genuinely worked gets described in specifics, by the people who used it, without anybody having to prompt them for examples.

Compare It Against How You Work Now

A result on its own means very little. What you need is a comparison against the current process, measured the same way, over the same work. That usually means running both for a period and accepting the duplication. Nobody enjoys it. Skip it and you end up comparing a measured new process against a remembered old one, and memory flatters whichever side the speaker is arguing for.

Measure the current process before anybody is excited. How long does the task take today. How often does it go wrong today. How much of it gets escalated today. Teams are regularly surprised by their own baseline, and occasionally the exercise ends there, because the existing process turns out to be better than anyone remembered.

Where running both is impossible, push the same set of past examples through each. It is weaker than a live comparison and far stronger than opinion.

Check where the work went, not only whether it went. Effort has a habit of moving rather than disappearing. Drafting gets faster and checking gets slower. The queue shrinks and the cases left in it get harder. A comparison that measures only the step you automated will report a saving the department as a whole never sees, and somebody will eventually notice that.

Everybody Liked It Is Not Evidence

Enthusiasm is real and it is worth something, mostly as a signal about adoption rather than about value. People like new tools. They like being on the team that was picked for the trial. They also like the version of the tool that is being watched closely, configured attentively, and supported by the vendor's own engineers in a shared channel that will be gone by autumn.

Ask the enthusiastic user a harder question. What did you use it for yesterday. What did you stop doing because of it. Show me the last thing it produced that you sent on without changing a word. Concrete answers are common among genuine users and rare among people who are being polite.

The reverse is worth watching too. A pilot everybody hated can still be the right call, and a pilot everybody loved can still cost more than it returns. Feelings tell you how the rollout will go. They tell you nothing about the economics.

Keep a Record While It Runs

Set up the record keeping before the first day, because reconstructing it afterwards produces a story rather than data. Three things are worth capturing without fail: every case the system got wrong, every case a person had to take over, and how long the checking took.

The failure log is the most valuable thing a pilot produces and it is usually thrown away at the end. Keep the input, the wrong output, and what the right answer was. That collection becomes the evaluation set you run before every change for years afterwards, and it cannot be manufactured later from memory.

Log the escalations separately, with a reason attached to each. A system that hands over often and always for a good reason is in decent shape. One that hands over rarely and wrongly is the expensive kind. On a usage chart the two look identical.

Narrow Scope, Real Deadline

One team, one job, one clearly stated period. A pilot spread across several departments and a handful of workflows produces a mixed outcome that everybody reads their own way, and no part of it can be attributed to anything in particular.

Pick the team carefully. The keenest team gives you the best possible case and tells you nothing about the ordinary one. A team that is busy, mildly sceptical and doing work you genuinely care about gives you a result you can generalise from.

Give it an end date and hold it. Pilots that get extended are usually being extended because the answer was no and nobody wants to say so out loud. One extension with a stated reason is fine. A second extension is a decision, whatever anybody calls it.

Deciding at the End

There are three honest outcomes and only one of them is a rollout. Roll it out. Stop, and spend the money elsewhere. Or change one specific thing and run it again, with the change and the new test written down before you restart.

If you are rolling out, decide who owns it in production before you go. The pilot ran with the vendor paying attention and your best people watching. Neither condition holds next quarter. Name the person who keeps the source content current, the person who watches the running cost, and the person who reads the failure log every week.

Write down what would make you turn it off, too. A feature nobody has permission to remove accumulates quietly, gets worse as the world around it changes, and eventually somebody discovers it has been producing wrong output for a long time with nobody watching.

Written by

The Omegaswift engineering team

AI and product engineering at Omegaswift. Filed under AI.

Ask us about this

Ready to talk about your IT?

We are happy to answer any questions you have and help you work out which of our services fit your needs.