AI API latency report · Sept 2026
Grok took more than 2× as long as ChatGPT to start responding
Across five cities and 1,050 streaming requests, median time to first visible token was 1.1 s for ChatGPT, 1.4 s for Claude and 2.5 s for Grok.
- ChatGPT
- 1.1 s
- Claude
- 1.4 s
- Grok
- 2.5 s
What the data shows
Finding 1
ChatGPT was fastest overall
1.1 s median time to first token, ahead of Claude at 1.4 s and Grok at 2.5 s.
- 58/70 city-days vs Claude
- 70/70 vs Grok
Finding 2
Provider choice mattered more than location
In every city, the gap between providers was larger than the gap between locations.
- Provider gap: ≥1,291 ms
- City gap: ≤466 ms
Finding 3
Most of the delay was not network
Network setup took just 15–27 ms; most of the wait came before the first visible token.
- Network: 15–27 ms
- First token: 826–2,687 ms
Time to first token, day by day
A dot marks a day when every compared city was well above that provider’s usual level.
Show the numbers
| Day | ChatGPT | Claude | Grok |
|---|---|---|---|
| 27 August 2026 | 1164 ms | — | 2431 ms |
| 28 August 2026 | 930 ms | — | 2451 ms |
| 29 August 2026 | 935 ms | 1255 ms | 2728 ms |
| 31 August 2026 | 1007 ms | 1435 ms | 3061 ms |
| 1 September 2026 | 1127 ms | 1223 ms | 2675 ms |
| 2 September 2026 | 1016 ms | 1517 ms | 2669 ms |
| 3 September 2026 | 1047 ms | 1491 ms | 2198 ms |
| 4 September 2026 | 850 ms | 1463 ms | 2998 ms |
| 5 September 2026 | 939 ms | 1297 ms | 2758 ms |
| 6 September 2026 | 951 ms | 1400 ms | 2524 ms |
| 7 September 2026 | 1339 ms | 1498 ms | 2507 ms |
| 8 September 2026 | 943 ms | 1249 ms | 1631 ms |
| 9 September 2026 | 993 ms | 2085 ms ▲ | 2108 ms |
| 10 September 2026 | 1161 ms | 2073 ms ▲ | 3183 ms |
| 11 September 2026 | 1209 ms | 1276 ms | 2960 ms |
| 12 September 2026 | 1160 ms | 1402 ms | 1920 ms |
| 13 September 2026 | 959 ms | 1408 ms | 2854 ms |
| 14 September 2026 | 1291 ms | 1323 ms | 2294 ms |
| 15 September 2026 | 1346 ms | 1366 ms | 2462 ms |
| 16 September 2026 | 1121 ms | 1297 ms | 2685 ms |
| 17 September 2026 | 999 ms | 1793 ms | 2496 ms |
| 18 September 2026 | 1133 ms | 1810 ms | 2933 ms |
| 19 September 2026 | 1063 ms | 1272 ms | 2910 ms |
| 20 September 2026 | 910 ms | 1231 ms | 2320 ms |
Does your location matter?
For AI APIs, your provider matters more than your server location.
- ChatGPT
- Claude
- Grok
Mumbai has been collecting since 7 September 2026 and joins the chart once it clears the publication rules for every provider.
Show exact values
| City | ChatGPT | Claude | Grok |
|---|---|---|---|
| Amsterdam | 1207 msslower 1.6 s | 1415 msslower 1.6 s | 2715 msslower 4.1 s |
| San Francisco | 1099 msslower 1.9 s | 1326 msslower 2.3 s | 2427 msslower 3.9 s |
| Montreal | 976 msslower 1.1 s | 1338 msslower 2.1 s | 2267 msslower 3.5 s |
| Singapore | 1272 msslower 1.7 s | 1505 msslower 2.3 s | 2733 msslower 4.1 s |
| Tokyo | 967 msslower 1.5 s | 1453 msslower 1.5 s | 2466 msslower 3.4 s |
Bold is the lowest typical time in that city. "Slower" is the time 95% of requests in the controlled run finished within.
Where the waiting happens
How much of the wait a closer server could remove, and how much is the provider’s own queue and model.
- ChatGPTthe network is about 3% of the wait
- Network setup
- 25 ms
- Waiting for the model
- 801 ms
First visible token after 826 ms in total.
- Claudethe network is about 1% of the wait
- Network setup
- 15 ms
- Waiting for the model
- 1316 ms
First visible token after 1331 ms in total.
- Grokthe network is about 1% of the wait
- Network setup
- 27 ms
- Waiting for the model
- 2660 ms
First visible token after 2687 ms in total.
Network setup is finding the server, reaching it and setting up a secure connection. Waiting for the model is everything after that until the first visible token: queueing and the model’s own work.
Side by side
| 7 September 2026 to 20 September 2026 | ChatGPT | Claude | Grok |
|---|---|---|---|
| Model | gpt-5.6-sol | claude-opus-5 | grok-4.6 |
| Typical time to first token (median city) | 1099 ms | 1415 ms | 2466 ms |
| Slower requests, p95 (median city) | 1555 ms | 2132 ms | 3859 ms |
| Fastest city | Tokyo, 967 ms | San Francisco, 1326 ms | Montreal, 2267 ms |
| Slowest city | Singapore, 1272 ms | Singapore, 1505 ms | Singapore, 2733 ms |
| Spread between cities | 32% | 13% | 21% |
| Days when every city slowed together | 0 | 2 | 0 |
| Requests in this window | 350 over 14 daily runs | 350 over 14 daily runs | 350 over 14 daily runs |
| Latest controlled run | 4 September 2026 | 20 September 2026 | 20 September 2026 |
How we test
We send the same short streaming prompt to each provider from five cities and measure time to first visible token. Results shown here use pooled daily requests, with separate controlled runs for the slower (p95) figure.
Claude Opus 5 includes its adaptive thinking phase; ChatGPT and Grok run at their lowest reasoning effort.
Full testing methodology →How does your API compare?
Test your endpoint from the same 5 cities used here and see your response time next to these numbers.
No account required · Takes about 30 seconds.