Define what “finished” means
Before timing anything, write down the output you need and the checks it must pass. A customer reply might need correct facts, an appropriate tone and no unsupported promises. A report might need every total to reconcile. Keep that standard the same in both versions of the process.
Decide whether you are measuring staff effort or elapsed waiting time. They answer different questions. Ten minutes of work spread across a day is not a day of labor. If you track both, give them separate columns. The method below measures active human effort.
Research on AI-assisted customer support found different effects for different workers in one specific setting. A published average is useful context, not a forecast of the result in your business.
Brynjolfsson, Li and Raymond: Generative AI at WorkRecord the current process first
Time at least five ordinary examples before changing the process. This is a practical starting sample for the worksheet, not a statistical threshold. Record the active effort needed to produce a usable result, including the checking and corrections you already do. Note the type and difficulty of each task.
Do not compare five complicated baseline cases with five easy trial cases. Choose a similar mix, and note differences in operator experience, busy periods or available information. With a small sample, one unusual case can change the average substantially. Keep the individual measurements so you can see that.
The downloadable CSV is a blank log, with no automatic formulas. For each baseline and trial case, use separate work, review and correction columns. In the single SETUP row, put the one-time setup total in work_minutes; exclude that row from the case count. Use yes or no for quality_pass, and keep personal details out of notes.
- Record actual minutes rather than guessing from memory at the end of the week.
- Keep failed and abandoned attempts in your record, including the time spent finishing them manually.
- Leave missing measurements blank. A blank value is unknown; zero means you measured no effort.
Count the whole AI-assisted job
For each trial case, record the time spent preparing inputs and producing the draft, checking it, and fixing it. Avoid counting the same minute twice. If a task fails, include both the failed attempt and the manual work needed to deliver an acceptable result.
| Category | Include |
|---|---|
| Work | Preparing input, prompting, handling output and any necessary manual completion. |
| Review | Checking facts, source material, calculations, tone and required approvals. |
| Corrections | Fixing errors and redoing work after review. |
| Setup | One-time preparation, training and configuration for this trial; count once, separately. |
Record subscription or usage charges separately as money. Record quality separately as well: whether the output met the standard, what went wrong, and whether a customer or colleague had to come back for a correction. Quicker handling does not compensate for unacceptable quality.
Compare the same amount of work
The comparison
Baseline equivalent = average baseline minutes × number of trial cases Trial effort = all trial work + review + corrections + one-time setup Net time difference = baseline equivalent − trial effort
Illustrative example: five baseline cases take 10 minutes each. Five comparable trial cases each take 5 minutes of work, 1 minute of review and 1 minute of corrections. Setup takes another 5 minutes. The baseline equivalent is 50 minutes; the trial takes 40. That is 10 fewer minutes for this sample, assuming quality remains acceptable.
If setup instead takes 30 minutes, this same trial takes 65 minutes overall: 15 more than the baseline. Routine handling is faster, but setup has not yet been recovered. Report both facts. Do not hide the setup cost to make a pilot look successful.
Make a decision the evidence can support
- Keep testing when comparable handling takes less effort and quality is at least as good. Track whether later work continues to support that result.
- Adjust when the problem is specific and fixable, such as an unclear source document or a costly review step. Change one part and test again.
- Stop when quality is unacceptable, review removes the benefit, or the workflow introduces a risk you cannot manage.
If the result is close or inconsistent, gather more comparable cases. A second person’s quality check can catch weaknesses you have become used to. Recheck after a material change to the tool, prompt, source information or type of work.
Original practical guidance by Future of My Business. Linked sources support the adjacent research or risk context; our checklists and examples are editorial tools, not validated benchmarks or endorsements.
How our assessment works