AI API benchmarks
AI API response times, measured daily
Time to first token from real servers around the world
- AI APIs
- 5
- Test locations
- 6
- Latest data
- 27 Sept 2026
Current AI latency benchmarks
Typical and tail latency for popular AI APIs, measured from 6 global locations.
- ChatGPT (Low reasoning effort)
gpt-5.6-sol
Typical TTFT1,025 msp95 TTFT2,055 msGeographic spread1.8×14-day trendFastest locationMontreal - Claude (Low reasoning effort)
claude-opus-5
Typical TTFT1,343 msp95 TTFT2,105 msGeographic spread1.2×14-day trendFastest locationMontreal - DeepSeek (Reasoning turned off)
deepseek-v4-flash
Typical TTFT698 msp95 TTFT937 msGeographic spread1.4×14-day trendFastest locationSingapore - Grok (Low reasoning effort)
grok-4.6
Typical TTFT2,495 msp95 TTFT4,136 msGeographic spread1.2×14-day trendFastest locationSan Francisco - Mistral (Provider default settings)
mistral-medium-latest
Typical TTFT344 msp95 TTFT525 msGeographic spread1.8×14-day trendFastest locationAmsterdam
*Low reasoning effort: ChatGPT, Claude, Grok.
**Provider default settings: Mistral.
***Reasoning turned off: DeepSeek.
Listed alphabetically. APIs with different marks were sent different reasoning settings, so their times aren't a like-for-like comparison.
Fast from where?
Darker cells are slower than that API's fastest location.
| Amsterdam | San Francisco | Montreal | Singapore | Tokyo | Mumbai | |
|---|---|---|---|---|---|---|
| 1,050 | 970 | 827 | 1,173 | 999 | 1,485 | |
| 1,295 | 1,294 | 1,245 | 1,500 | 1,390 | 1,486 | |
| 735 | 742 | 709 | 546 | 569 | 687 | |
| 2,593 | 2,229 | 2,330 | 2,612 | 2,471 | 2,519 | |
| 242 | 349 | 295 | 380 | 430 | 338 |
Where the wait comes from
Reaching the server is at most 6% of the wait. The rest is the model.
Latest findings
Comparison · 27 Sept
ChatGPT vs Claude vs Grok, who's fastest?
ChatGPT was usually first to respond, Grok usually last.
See comparison →Monthly report · 27 Sept
State of AI API latency, September 2026
Provider choice mattered more than location.
Read the report →Comparison · 22 Sept
DeepSeek direct or via OpenRouter?
Asia favored direct. Montreal favored OpenRouter.
See comparison →
How we measure
Every day, each API gets the same one-line prompt as a streaming request from every test location. We time how long the first visible word of the reply takes. Hidden reasoning doesn't count, and we don't measure how fast the rest of the reply arrives or how good it is.
Read the full methodology →How fast is your own API?
Run a free test from the same 6 cities and see where the time goes.
Free, no account needed.