Measuring results
How to measure an automation pilot
Use a simple baseline and worked example to measure time saved, quality and the effort that remains after automation.

The useful part, upfront
- Compare equivalent work and include review, rework and exception handling.
- Measure quality and turnaround alongside the time your team spends.
- Treat released capacity as capacity until you know how the business uses it.
Agree what better means before the build
A pilot should answer a business question: does this workflow now take less effort, move faster or produce a more reliable result? A polished demonstration cannot answer that on its own. You need a baseline, a clear definition of success and a way to see the work that remains.
Choose one primary measure, such as active handling time per completed item. Add quality checks so that a faster result is still a useful one. Record turnaround time separately if waiting is part of the problem.
Decide what evidence would support extending the pilot, what would require another iteration and what would make you stop. These decisions should be agreed with the person who owns the workflow, before the team becomes attached to a particular solution.
Measure the work people actually do
Observe a representative period of the existing process. Include routine items and the exceptions that take longer. Record volume, hands-on time, waiting time and corrections. If work varies by season or job type, keep those categories visible rather than blending them into one average.
Use the same start and finish points in both versions. If the baseline ends when a record is approved, the pilot must also include approval. Measuring only how quickly a draft appears would leave out the work needed to make it usable.
| Measure | What to record |
|---|---|
| Volume | Eligible items received and completed |
| Handling time | Doing, reviewing, correcting and resolving exceptions |
| Turnaround | Elapsed time from the agreed trigger to completion |
| Quality | Items meeting the agreed standard and types of defect |
| Reliability | Failed runs, repeated actions and unresolved items |
| Upkeep | Time spent checking and maintaining the workflow |
A worked example: count the effort that remains
The figures below are hypothetical. They illustrate the calculation and are not an Amtec client result or a forecast. Assume a team handles 100 comparable requests each week. Today, each request takes an average of 12 minutes of active work, including its share of corrections.
In the pilot, all 100 requests need four minutes of review and handling each. Twenty also need an extra five minutes to resolve an exception. The team spends one additional hour each week checking and maintaining the workflow. Count all three parts.
| Calculation | Weekly time |
|---|---|
| Before: 100 requests × 12 minutes | 1,200 minutes / 20 hours |
| After: 100 requests × 4 minutes | 400 minutes |
| Plus exceptions: 20 requests × 5 extra minutes | 100 minutes |
| Plus checking and upkeep | 60 minutes |
| Total effort after the change | 560 minutes / 9 hours 20 minutes |
| Illustrative capacity released: 1,200 − 560 minutes | 640 minutes / 10 hours 40 minutes |
Make sure the comparison is fair
A quieter week or a simpler mix of requests can make a pilot look better than it is. Compare similar work, show the number of items observed and report exceptions separately. Where practical, compare equivalent batches during the same period. A small pilot is evidence for the next decision, not proof of a permanent business-wide result.
Define quality using observable checks: the correct customer, complete required fields, accurate source details and the right approval. Record how many items passed, how many needed correction and what went wrong. Check unresolved work too; excluding unfinished items can hide the hardest cases.
For AI steps, keep representative test cases and rerun them after changes. Anthropic’s evaluation guidance emphasises explicit success criteria and checking task outcomes. A business pilot should also assess the finished work, rather than relying only on whether the system produced an answer.
Turn the evidence into a decision
Review the result with the people doing the work. Show the baseline and pilot side by side, include software and support costs, and keep one-off setup effort visible separately from ongoing operation. Avoid counting the same benefit twice under both time saved and extra output.
- Extend when quality meets the agreed standard, time improves and an owner can support ongoing use.
- Revise when the result is promising but a specific source problem, review step or exception causes avoidable work.
- Pause when errors are unacceptable, the team cannot control the workflow or the effort has simply moved elsewhere.
Further reading
Glow Saunas
Explore the Glow Saunas project, where manual effort and order lead times were the measures that mattered.
Explore the project