AI API benchmarks

AI API response times, measured daily

Time to first token from real servers around the world

AI APIs
5
Test locations
6
Latest data
27 Sept 2026

What time to first token means →How we measure →

Current AI latency benchmarks

Typical and tail latency for popular AI APIs, measured from 6 global locations.

  • Typical TTFT1,025 ms
    p95 TTFT2,055 ms
    Geographic spread1.8×
    14-day trendDaily typical time, 14 Sept to 27 Sept: between 729 ms and 1,403 ms.
    Fastest locationMontreal
  • Typical TTFT1,343 ms
    p95 TTFT2,105 ms
    Geographic spread1.2×
    14-day trendDaily typical time, 14 Sept to 27 Sept: between 1,109 ms and 1,824 ms.
    Fastest locationMontreal
  • Typical TTFT698 ms
    p95 TTFT937 ms
    Geographic spread1.4×
    14-day trendDaily typical time, 14 Sept to 27 Sept: between 558 ms and 902 ms.
    Fastest locationSingapore
  • Typical TTFT2,495 ms
    p95 TTFT4,136 ms
    Geographic spread1.2×
    14-day trendDaily typical time, 14 Sept to 27 Sept: between 1,445 ms and 2,961 ms.
    Fastest locationSan Francisco
  • Typical TTFT344 ms
    p95 TTFT525 ms
    Geographic spread1.8×
    14-day trendDaily typical time, 14 Sept to 27 Sept: between 302 ms and 376 ms.
    Fastest locationAmsterdam

*Low reasoning effort: ChatGPT, Claude, Grok.

**Provider default settings: Mistral.

***Reasoning turned off: DeepSeek.

Listed alphabetically. APIs with different marks were sent different reasoning settings, so their times aren't a like-for-like comparison.

Fast from where?

Darker cells are slower than that API's fastest location.

Typical time to first token in milliseconds, by provider and test location
AmsterdamSan FranciscoMontrealSingaporeTokyoMumbai
ChatGPT1,0509708271,1739991,485
Claude1,2951,2941,2451,5001,3901,486
DeepSeek735742709546569687
Grok2,5932,2292,3302,6122,4712,519
Mistral242349295380430338
Typical TTFT in ms. Compared with each API's fastest location:

Where the wait comes from

Reaching the server is at most 6% of the wait. The rest is the model.

Latest findings

How we measure

Every day, each API gets the same one-line prompt as a streaming request from every test location. We time how long the first visible word of the reply takes. Hidden reasoning doesn't count, and we don't measure how fast the rest of the reply arrives or how good it is.

Read the full methodology →

How fast is your own API?

Run a free test from the same 6 cities and see where the time goes.

Test your API

Free, no account needed.