Low latency is difficult to evaluate when the timer is not defined. A voice system can measure time to the first model event, first audio chunk, first audible playback, or completion of the entire response. Those are different metrics and answer different questions.
The voice-agent project in my portfolio reports sub-200 ms response latency, but its public record does not specify which interval was timed or publish raw measurements. A reproducible benchmark needs that context before the number can be compared or independently checked.
Choose the metric before testing
For conversational responsiveness, teams may measure the interval from the end of a user turn to the first audible response. They may separately track time to the first returned audio chunk and time until the complete response finishes. Document whether turn detection, buffering, network transport, and playback are inside each measurement window.
Control and describe test conditions
Record the client device, audio input, network conditions, geographic region, model or service configuration, and whether the connection was already open. Run repeatable scripted turns and representative natural speech. Do not mix cold connection setup with warm streaming turns without labeling the difference.
Repeat each scenario enough times to show variability. Report the sample count and a distribution such as median and p95 rather than the single fastest run. Preserve timestamps or traces that allow the team to reproduce the calculation.
Separate component timing and quality
Measure capture and buffering, network round trips, time to first model output, and playback startup separately when instrumentation supports it. This can distinguish model delay from client-side audio delay. Avoid presenting estimated component timings as a direct end-to-end measurement.
Latency is not the only acceptance measure. Check interruption behavior, intelligibility, unstable connections, errors, and dropped turns. A fast but incomplete response may not be useful.
A credible result states its start and end events, test conditions, sample size, and distribution. The [documented voice-agent project](/case-studies/ai-voice-agent-realtime/) provides context for its reported result.