AI API latency report · Sept 2026

Grok took more than 2× as long as ChatGPT to start responding

Across five cities and 1,050 streaming requests, median time to first visible token was 1.1 s for ChatGPT, 1.4 s for Claude and 2.5 s for Grok.

ChatGPT
1.1 s
Claude
1.4 s
Grok
2.5 s

What the data shows

  1. Finding 1

    ChatGPT was fastest overall

    1.1 s median time to first token, ahead of Claude at 1.4 s and Grok at 2.5 s.

    • 58/70 city-days vs Claude
    • 70/70 vs Grok
    See the day-by-day chart
  2. Finding 2

    Provider choice mattered more than location

    In every city, the gap between providers was larger than the gap between locations.

    • Provider gap: ≥1,291 ms
    • City gap: ≤466 ms
    See the table by city
  3. Finding 3

    Most of the delay was not network

    Network setup took just 15–27 ms; most of the wait came before the first visible token.

    • Network: 15–27 ms
    • First token: 826–2,687 ms
    See where the time goes

Time to first token, day by day

Daily median time to first token, 27 August 2026 to 20 September 2026. Each point is the median of the compared cities’ daily medians, 5 requests per city.
01000200030004000 ms27 Aug20 SeptClaude, 9 September 2026: 2085 ms, every compared city well above its usual level.Claude, 10 September 2026: 2073 ms, every compared city well above its usual level.23201231910

A dot marks a day when every compared city was well above that provider’s usual level.

Show the numbers
DayChatGPTClaudeGrok
27 August 20261164 ms2431 ms
28 August 2026930 ms2451 ms
29 August 2026935 ms1255 ms2728 ms
31 August 20261007 ms1435 ms3061 ms
1 September 20261127 ms1223 ms2675 ms
2 September 20261016 ms1517 ms2669 ms
3 September 20261047 ms1491 ms2198 ms
4 September 2026850 ms1463 ms2998 ms
5 September 2026939 ms1297 ms2758 ms
6 September 2026951 ms1400 ms2524 ms
7 September 20261339 ms1498 ms2507 ms
8 September 2026943 ms1249 ms1631 ms
9 September 2026993 ms2085 ms ▲2108 ms
10 September 20261161 ms2073 ms ▲3183 ms
11 September 20261209 ms1276 ms2960 ms
12 September 20261160 ms1402 ms1920 ms
13 September 2026959 ms1408 ms2854 ms
14 September 20261291 ms1323 ms2294 ms
15 September 20261346 ms1366 ms2462 ms
16 September 20261121 ms1297 ms2685 ms
17 September 2026999 ms1793 ms2496 ms
18 September 20261133 ms1810 ms2933 ms
19 September 20261063 ms1272 ms2910 ms
20 September 2026910 ms1231 ms2320 ms

Does your location matter?

For AI APIs, your provider matters more than your server location.

Typical time to first token by city: the median of every daily request from that city over the window.
0100020003000 msChatGPT from Amsterdam: 1207 ms typical time to first token; slower requests 1555 ms.Claude from Amsterdam: 1415 ms typical time to first token; slower requests 1554 ms.Grok from Amsterdam: 2715 ms typical time to first token; slower requests 4087 ms.AmsterdamChatGPT from San Francisco: 1099 ms typical time to first token; slower requests 1933 ms.Claude from San Francisco: 1326 ms typical time to first token; slower requests 2278 ms.Grok from San Francisco: 2427 ms typical time to first token; slower requests 3859 ms.San FranciscoChatGPT from Montreal: 976 ms typical time to first token; slower requests 1092 ms.Claude from Montreal: 1338 ms typical time to first token; slower requests 2132 ms.Grok from Montreal: 2267 ms typical time to first token; slower requests 3486 ms.MontrealChatGPT from Singapore: 1272 ms typical time to first token; slower requests 1732 ms.Claude from Singapore: 1505 ms typical time to first token; slower requests 2303 ms.Grok from Singapore: 2733 ms typical time to first token; slower requests 4073 ms.SingaporeChatGPT from Tokyo: 967 ms typical time to first token; slower requests 1537 ms.Claude from Tokyo: 1453 ms typical time to first token; slower requests 1512 ms.Grok from Tokyo: 2466 ms typical time to first token; slower requests 3403 ms.Tokyo
  • ChatGPT
  • Claude
  • Grok

Mumbai has been collecting since 7 September 2026 and joins the chart once it clears the publication rules for every provider.

Show exact values
CityChatGPTClaudeGrok
Amsterdam1207 msslower 1.6 s1415 msslower 1.6 s2715 msslower 4.1 s
San Francisco1099 msslower 1.9 s1326 msslower 2.3 s2427 msslower 3.9 s
Montreal976 msslower 1.1 s1338 msslower 2.1 s2267 msslower 3.5 s
Singapore1272 msslower 1.7 s1505 msslower 2.3 s2733 msslower 4.1 s
Tokyo967 msslower 1.5 s1453 msslower 1.5 s2466 msslower 3.4 s

Bold is the lowest typical time in that city. "Slower" is the time 95% of requests in the controlled run finished within.

Where the waiting happens

How much of the wait a closer server could remove, and how much is the provider’s own queue and model.

Median city in each provider’s controlled 20-request run (4 September 2026; 20 September 2026). Both bars share one scale.
  • ChatGPTthe network is about 3% of the wait
    Network setup
    25 ms
    Waiting for the model
    801 ms

    First visible token after 826 ms in total.

  • Claudethe network is about 1% of the wait
    Network setup
    15 ms
    Waiting for the model
    1316 ms

    First visible token after 1331 ms in total.

  • Grokthe network is about 1% of the wait
    Network setup
    27 ms
    Waiting for the model
    2660 ms

    First visible token after 2687 ms in total.

Network setup is finding the server, reaching it and setting up a secure connection. Waiting for the model is everything after that until the first visible token: queueing and the model’s own work.

Side by side

7 September 2026 to 20 September 2026ChatGPTClaudeGrok
Modelgpt-5.6-solclaude-opus-5grok-4.6
Typical time to first token (median city)1099 ms1415 ms2466 ms
Slower requests, p95 (median city)1555 ms2132 ms3859 ms
Fastest cityTokyo, 967 msSan Francisco, 1326 msMontreal, 2267 ms
Slowest citySingapore, 1272 msSingapore, 1505 msSingapore, 2733 ms
Spread between cities32%13%21%
Days when every city slowed together020
Requests in this window350 over 14 daily runs350 over 14 daily runs350 over 14 daily runs
Latest controlled run4 September 202620 September 202620 September 2026

How we test

We send the same short streaming prompt to each provider from five cities and measure time to first visible token. Results shown here use pooled daily requests, with separate controlled runs for the slower (p95) figure.

Claude Opus 5 includes its adaptive thinking phase; ChatGPT and Grok run at their lowest reasoning effort.

Full testing methodology →

How does your API compare?

Test your endpoint from the same 5 cities used here and see your response time next to these numbers.

Run a free latency test

No account required · Takes about 30 seconds.

Explore each API