Free guide · Francisco Arrieta · 7 min

Measure whether AI is actually saving you time

Run a two week test on your own work, because the best controlled study found people were slower while believing they were faster

Sixteen experienced developers worked on real issues in their own repositories, some with AI tools and some without, randomly assigned. They took 19% longer with AI. Afterwards they said it had made them about 20% faster.

Sit with that for a second, because it isn’t a story about AI being bad. Those developers weren’t fooled by marketing and they weren’t new to the tools. They did the work, they had the experience to judge it, and their sense of their own speed was off by nearly forty points against the clock.

Meanwhile the surveys say something else entirely. Hours saved per week, minutes saved per day, take your pick of the number. They’re mostly self-reported, and self-report is precisely the thing that study found to be unreliable. They’re also averages of other people doing other work, which tells you very little about the specific task you were going to do this afternoon.

So the interesting question isn’t whether AI saves time in general. It’s whether it saves you time on this task, and that’s measurable in about two weeks with a stopwatch and a note.

You can’t feel this one. You have to count it.

Before you start

Two weeks of ordinary work, and somewhere to write down numbers.

Pick one task type to start. Not your whole job. Something you do at least a few times a week, that has a recognizable end, and that you’d plausibly use AI for: drafting client emails, writing a first pass of a proposal, summarizing a call, producing a chunk of code, editing copy.


Step 1

Write down your guess first, and seal it

Before you measure anything, write down what you expect. A number, not a feeling. “I think AI makes this about 30% faster.”

Put it somewhere you won’t edit it. This is the whole experiment. Without a guess recorded up front, you’ll read the result and quietly remember having expected it, which is exactly what happened to the developers in that study.

Date it. You’re going to compare against it in step 5.


Step 2

Decide what finished means

Write down what “done” is for this task, in one sentence, before you start timing.

This matters more than it sounds. If done means “a draft exists,” AI wins nearly every task by definition. If done means “sent to the client,” you’ve included the reading, the correcting, and the bit where you rewrite the opening because it sounds like nobody.

Use the honest definition: the point at which you’d stop touching it. Anything short of that measures the fast half and hides the slow half.


Step 3

Alternate, don’t choose

For the next two weeks, do the task both ways. Alternate, or flip a coin. Do not pick per task.

Choosing is what breaks the experiment. You’ll reach for AI on the ones that feel suited to it and skip it on the ones that don’t, and at the end you’ll have measured your instinct rather than the tool.

Log four things each time: the date, which way you did it, the total minutes, and one line on whether the output was any good. Paper is fine.

You want at least five of each before the numbers mean anything. Fewer than that and you’re measuring what kind of week you had.


Step 4

Count all of the time, not just the part with the AI in it

When you total up, include everything the task consumed:

  • Writing the prompt, and rewriting it when the first answer missed
  • Reading what came back
  • Checking anything factual in it
  • Fixing the parts that were wrong or sounded wrong
  • The second attempt, when there was one

That’s the number. The stretch where you were typing to the model is not the task, it’s part of the task, and measuring only that part is how a tool looks like it saves an hour while the afternoon still disappears.


Step 5

Open your guess

Now compare. Average each column, and put your step 1 number next to them.

Three things you might find, and all three are useful:

Faster, and you knew roughly by how much.
Good. Keep going, and you now have a number rather than a vibe.
Faster, but much less than you guessed.
The common result. The tool works and your sense of the size was inflated, which matters if you’ve made plans on the strength of it.
Slower.
Also common, especially on work you’re already fluent at. Not an argument against AI. An argument against AI for this task.

Look at the quality column too. Faster and worse is a real outcome, and so is slower and better. Neither shows up in a time average.


Step 6

Decide per task, then leave it alone

The answer is not global and it does not transfer. Fluent work tends to lose. Unfamiliar work, blank pages, and things you’d have procrastinated on tend to win.

Write one line per task type: use AI, don’t, or use it for the first draft only. Then stop measuring and go back to work. This is a two-week experiment, not a practice.

Re-run it in six months on the same task. The tools change fast enough that today’s answer has a shelf life, and you’ll have your own baseline to compare against rather than somebody’s survey.


A boundary worth knowing about

Time is one measure and it’s the easiest one, which is why everybody uses it.

It says nothing about whether the work got better, whether you learned anything, or whether you’ll still understand the thing in six months. A developer who ships faster and understands their codebase less has traded something, and no stopwatch shows the trade. Same for a writer whose drafts arrive quicker and sound more like everyone else’s.

The other edge is that this measures you, on your tasks, at your level of skill with the tools. Somebody two months into using them properly may get a different answer, and so might you. That’s an argument for re-running it, not for distrusting your own result.


If you have staff

Ask for the measurement, never the feeling. “Is AI helping?” gets you a yes, because nobody wants to say the tool they were given is slowing them down.

Then be careful what you do with the answer, because this is where it goes wrong. There are surveys suggesting a majority of employees deliberately conceal the time AI saves them, staying visibly online rather than surfacing the capacity. Whether or not the exact figure holds, the incentive behind it is real and obvious: if saved time reliably becomes more work, it stops being reported.

If you want honest numbers, say in advance what the time is for. A team that believes the answer is “you finish earlier” measures differently from one that suspects it’s “you get more tickets.”


The short version

  1. Write your guess down as a number and seal it
  2. Define what done means, honestly, before timing anything
  3. Alternate for two weeks. Never choose per task. At least five each way
  4. Count the whole task: prompting, reading, checking, fixing, second attempts
  5. Open the guess and compare. Look at the quality notes as well as the average
  6. Decide per task type, write one line, and stop measuring

Sources

Written August 2026. The METR result is from July 2025 and the tools have moved since, which is an argument for measuring your own rather than trusting either their number or a survey’s. The method here does not date.

Prints to PDF from your browser — colours and all.