Two eras of computing, held to the same 400 millisecond standard.

Doherty’s threshold set the rules and has now moved from the answer to the acknowledgment. The fix is a week of front-end work and a definition of done, not cheaper inference.

Productivity soars when a computer and its users interact at a pace — under 400 milliseconds — that ensures neither has to wait on the other. Below the threshold, work feels like a conversation; above it, each pause breaks the user’s train of thought and errors climb.
 — Walter Doherty

Ask a product team why their AI feature feels slow and you get a shrug wrapped in a technical explanation.

The model takes eight seconds.

Inference is expensive.

Everyone’s model is slow.

Buried in the shrug is an assumption: that the 400 millisecond threshold belongs to green screens and the entire single transactions, and generative AI retired it. It actually rewrote the threshold’s job, and the new job sits entirely on the designer’s side of the request.

The reality is that Walter Doherty and Ahrvind Thadani were measuring what happens to a person’s attention when a machine makes them wait, and they attached a dollar value to it.

What generative AI changed is the shape of the thing measured. It split one event into two — the acknowledgment and the answer — and the threshold followed the acknowledgment.

The stakes changed too, and this is the part worth arguing about.

In 1982, a slow response cost throughput. In an AI product it costs supervision. Every AI feature assumes a person is paying close enough attention to catch the model being wrong, and attention is exactly what Doherty’s threshold protects.

The acknowledgment budget is a supervision budget. Overspend it and the human in the loop check your product depends on quietly stops happening.

What Doherty Actually Measured

The 1982 paper, The Economic Value of Rapid Response Time, plotted transactions per hour against system response time, and the shape is not a line.

Below roughly 400 milliseconds, productivity climbs faster than the time saved can explain, because something changes in the operator rather than the machine. People stop batching their thinking around the pauses. They hold one thread instead of rebuilding context on every return.

That was an economic argument before it was a design one, which is why the number still works in a room full of executives. When system response time dropped, the operator’s own think time dropped with it.

The threshold was never a statement about how fast a computer can think. It was a statement about how long a person will hold a thought.

Milliseconds Make Millions. Speed matters.

It keeps getting re-measured.

Deloitte Digital found a 0.1 second speed improvement raised retail conversions 8.4% across 37 brands in Milliseconds Make Millions, 2020. And it runs on the same budget conversation does: across ten languages, Universals and cultural variation in turn-taking in conversation found a median gap between turns of about 100 milliseconds.

The threshold has also been rewritten before. Jakob Nielsen restated the same physics for the web as three limits — a tenth of a second feels instant, one second holds the flow of thought, ten seconds is the edge of attention — and those numbers have governed interface work since.

Rewriting What The Threshold Governs

In 1982, request and response were one event, and one clock covered all of it.

A generative interaction is two events wearing one costume.

  • The turn, which asks whether the system heard you
  • The work, which asks whether the answer is any good

Josh Clark and Veronika Kindred call the result an intelligent interface in Say Hello to Sentient Design: a system that composes its response in the moment rather than serving a prepared one.

Artificial Analysis. Higher is better. Napster, still bad.

Artificial Analysis reports time to first answer token in seconds, and for reasoning models that figure includes the thinking the model does before it says anything at all.

Teams read that number, conclude the interaction is seconds-scale, and design accordingly. They are treating a measurement of the model as a measurement of the interface. The interface has its own budget, no dependency on inference, and can echo an input in 50 milliseconds on a bad phone.

The model is allowed to take eight seconds to return the answer. The initial time after selecting the send button is not.

In enterprise software the mistake compounds, because the person using the tool has a queue.

Microsoft’s 2025 telemetry in Breaking down the infinite workday has the average Microsoft 365 user receiving 117 emails and 153 Teams messages a weekday, interrupted every two minutes. Put an AI assist on a list that size and a one-second acknowledgment delay is four minutes of unanswered clicks a day, across 270 moments where the person was already being pulled elsewhere.

Consumer chat gets one slow turn and a curious user. Work software multiplies it by the queue.

This describes what’s in an interaction from Interaction to Next Paint. There are ways we can mitigate presentation delay.

Google’s bar is tighter than Doherty’s and applies squarely to the send button: Interaction to Next Paint treats 200 milliseconds or less, at the 75th percentile of real visits, as good. Your prompt box is already held to it, whatever the model does next.

Streaming Is Not An Acknowledgment

Streaming is the pattern every team reaches for, and it earns its reputation. Words arriving at the speed you read them make a long answer feel like a conversation instead of a queue.

It solves the wrong problem.

Luke Wroblewski was making the case against the gap in front of the first token back in 2013, in Mobile Design Details: Avoid The Spinner. The failure people actually feel happens before that: the press that produces nothing, the input that clears with nothing in its place, the thread that sits empty while a reasoning model spends fifteen seconds deciding how to begin.

Streaming turns a wait into reading. It does nothing for the half second when there is nothing to read.

The gap also moves without warning. Turn on extended reasoning, route to a larger model, add a retrieval step, and the time before the first token stretches while the interface stays the same. The acknowledgment is the only part that holds still when the routing changes.

A second failure hides inside the first. A stream that shows motion while the system is stalled, or a progress line narrating steps it is not taking, buys patience with a claim that turns out to be false.

Users forgive slowness. They do not forgive being managed.

Two clocks, two owners, two sets of rules.

Two Clocks, Two Budgets

Here is the frame I would put in the spec: every AI interaction runs two clocks.

Clock one

Clock one starts when the user commits — enter, click, tap — and stops when the interface proves it heard. Echo the prompt into the thread. Change the button. Put the skeleton where the answer will land. Target 100 milliseconds, treat 400 as the ceiling. This is front-end work with a hard number attached, and it is where the supervision budget gets spent or saved.

Most teams run one clock: the final response. There really should two clocks: acknowledgement and the final response.

Clock two

It starts when the request leaves and stops when the answer is usable, and people will wait here, sometimes for a long time, when the wait buys something. Nobody abandons a four-minute research run that returns a report worth four minutes.

Duration is not what breaks the exchange. Silence is: a wait with no receipt, no visible progress, and no sense of what it is purchasing. Speed is not the governing rule on this clock — honesty is.

Ben Shneiderman put informative feedback among his eight golden rules in 1985, and the rule survives the rewrite intact: show real state, keep cancel available throughout, and if you commit a change optimistically, as Luke Wroblewski described in Performing Actions Optimistically, have a rollback that tells the truth when the write fails.

Most teams instrument clock two, because that is what the model dashboard reports, and leave clock one unowned.

That is how a product gets a respectable p95 latency chart and a feature that feels dead on the first tap. Clock two already belongs to the platform or applied AI team, the people with leverage over prompts, routing, and caching. Clock one belongs to whoever builds the surface, and nobody wrote it down as a requirement.

Put both numbers in one document with a name on each, and the argument about whose problem the slowness is ends.

Latency Debt Shows Up In Agents

A run with four model calls, two tool calls, and a retry inherits the acknowledgment problem at every hop, and the pauses add up to more than any single delay predicts.

Take a contract review agent which I have done before:

  • Pull the clauses
  • Compare each against the playbook
  • Flag the gaps
  • Draft the redlines

Four steps, a few seconds each, none alarming alone. Most interfaces show one spinner for the run, so nothing tells the user at step two that the agent grabbed the wrong document version.

They find out at the end, in a redline they have to unwind. Wroblewski has been cataloging the fix in Showing the Work of Agents in UI: the run’s steps, visible as they happen, so the person can stop it.

Watch what people do and you see the 1982 behavior. They batch: fire off the run, switch to Slack, come back later. The mainframe adaptation, wearing a nicer interface.

Users who leave the tab do not come back to review your output. They come back to accept it which is not human in the loop.

That is the supervision budget, overspent. A collaborative tool becomes a background job, and the review step you designed becomes a rubber stamp. This is the out-of-the-loop problem arriving through the side door, and latency is what opened it. Nobody decided to trust the model more. The interface made supervision expensive.

Latency debt does not stay a performance problem. It redesigns the product, and the surface it demands — notifications, run histories, a way back into a task nobody was supposed to leave — is work you never scoped.

The Budget You Already Own

Doherty’s threshold survived the command line, the graphical interface, the web, and mobile. Each transition produced people arguing the old numbers no longer applied, and each time the number outlasted them, because it describes attention rather than hardware.

Chat is not the exception. The rewrite splits the response a user waits for from the answer a model produces. Teams holding both numbers ship interfaces that feel fast while running models that are not. Teams holding one wait for inference to get cheaper and call that a roadmap.

The 400 millisecond line is still the most useful performance number in design, because it is achievable today, without a single change to the model. It used to protect productivity. Now it protects the person you are counting on to check the model’s work, and that is far more expensive to lose.

The model’s latency is a vendor’s problem. The first 400 milliseconds are yours.

Action Items

  • Split the budget in the spec. Two numbers beside the visual requirements: acknowledgment under 400 milliseconds, answer unbounded but narrated. One team owns each.
  • Echo the input first, always. The user’s text lands in the thread before any request leaves the client. No inference, no network, no excuse.
  • Replace thinking animations with real state. Searching, reading, drafting, calling a tool. If the interface cannot name the step, that is a missing event, not a copy problem.
  • Measure the first tap in the field. Track acknowledgment latency on real devices and networks, the way you track Interaction to Next Paint, not on the demo laptop.
  • Give every wait an exit, and a reason. Cancel, edit and resend, or start over — and say what the wait is buying. People sit through long waits with a visible payoff and abandon short ones without.
  • Audit the agent run hop by hop. Find the silent step and give it a visible state before you spend a quarter chasing faster inference.

Stop blaming the model for slow AI. Using Doherty’s threshold as a guideline. was originally published in UX Collective on Medium, where people are continuing the conversation by highlighting and responding to this story.