How to measure an automation pilot

Use a simple baseline and worked example to measure time saved, quality and the effort that remains after automation.

On this page

The useful part, upfront

  • Compare equivalent work and include review, rework and exception handling.
  • Measure quality and turnaround alongside the time your team spends.
  • Treat released capacity as capacity until you know how the business uses it.

Agree what better means before the build

A pilot should answer a business question: does this workflow now take less effort, move faster or produce a more reliable result? A polished demonstration cannot answer that on its own. You need a baseline, a clear definition of success and a way to see the work that remains.

Choose one primary measure, such as active handling time per completed item. Add quality checks so that a faster result is still a useful one. Record turnaround time separately if waiting is part of the problem.

Decide what evidence would support extending the pilot, what would require another iteration and what would make you stop. These decisions should be agreed with the person who owns the workflow, before the team becomes attached to a particular solution.

Measure the work people actually do

Observe a representative period of the existing process. Include routine items and the exceptions that take longer. Record volume, hands-on time, waiting time and corrections. If work varies by season or job type, keep those categories visible rather than blending them into one average.

Use the same start and finish points in both versions. If the baseline ends when a record is approved, the pilot must also include approval. Measuring only how quickly a draft appears would leave out the work needed to make it usable.

MeasureWhat to record
VolumeEligible items received and completed
Handling timeDoing, reviewing, correcting and resolving exceptions
TurnaroundElapsed time from the agreed trigger to completion
QualityItems meeting the agreed standard and types of defect
ReliabilityFailed runs, repeated actions and unresolved items
UpkeepTime spent checking and maintaining the workflow

A worked example: count the effort that remains

The figures below are hypothetical. They illustrate the calculation and are not an Amtec client result or a forecast. Assume a team handles 100 comparable requests each week. Today, each request takes an average of 12 minutes of active work, including its share of corrections.

In the pilot, all 100 requests need four minutes of review and handling each. Twenty also need an extra five minutes to resolve an exception. The team spends one additional hour each week checking and maintaining the workflow. Count all three parts.

CalculationWeekly time
Before: 100 requests × 12 minutes1,200 minutes / 20 hours
After: 100 requests × 4 minutes400 minutes
Plus exceptions: 20 requests × 5 extra minutes100 minutes
Plus checking and upkeep60 minutes
Total effort after the change560 minutes / 9 hours 20 minutes
Illustrative capacity released: 1,200 − 560 minutes640 minutes / 10 hours 40 minutes

Make sure the comparison is fair

A quieter week or a simpler mix of requests can make a pilot look better than it is. Compare similar work, show the number of items observed and report exceptions separately. Where practical, compare equivalent batches during the same period. A small pilot is evidence for the next decision, not proof of a permanent business-wide result.

Define quality using observable checks: the correct customer, complete required fields, accurate source details and the right approval. Record how many items passed, how many needed correction and what went wrong. Check unresolved work too; excluding unfinished items can hide the hardest cases.

For AI steps, keep representative test cases and rerun them after changes. Anthropic’s evaluation guidance emphasises explicit success criteria and checking task outcomes. A business pilot should also assess the finished work, rather than relying only on whether the system produced an answer.

Turn the evidence into a decision

Review the result with the people doing the work. Show the baseline and pilot side by side, include software and support costs, and keep one-off setup effort visible separately from ongoing operation. Avoid counting the same benefit twice under both time saved and extra output.

  • Extend when quality meets the agreed standard, time improves and an owner can support ongoing use.
  • Revise when the result is promising but a specific source problem, review step or exception causes avoidable work.
  • Pause when errors are unacceptable, the team cannot control the workflow or the effort has simply moved elsewhere.

Further reading

See it in practice

Glow Saunas

Explore the Glow Saunas project, where manual effort and order lead times were the measures that mattered.

Explore the project

Keep going

Getting started 5 min read

Which workflow should you automate first?

Read the guide

Making AI useful 5 min read

AI or automation: what does your workflow need?

Read the guide

Bring us the workflow.Find your next step.

Tell us where work gets stuck and which tools your team uses. We’ll help you identify a practical way forward.