Benchmark report
ChatGPT API thumbnail

Is the ChatGPT API slow? GPT-5.6 time to first token from 5 cities, measured daily

We ran real, authenticated streaming completions against OpenAI's flagship gpt-5.6-sol, with reasoning effort set to "low", from Amsterdam, San Francisco, Montreal, Singapore and Tokyo, every day from 27 August to 4 September 2026, plus two controlled runs of 20 requests per city. We parsed the token stream ourselves and timestamped the moment the first visible token arrived: the time to first token (TTFT).

What we tested · AI response

Time to first visible token from gpt-5.6-sol on a one-word prompt: streaming, reasoning_effort low, 256 max tokens, paid key.

Doesn’t measure: The ChatGPT app, tokens per second, long prompts, other models, tool calls or the Batch API.

Quick take

The first token typically arrived in 0.97 to 1.05 seconds, measured 40 times per city over eight days: San Francisco and Tokyo quickest at 971 ms, Singapore slowest at 1053 ms, an 8% spread. Almost all of it is OpenAI-side: subtract each city's network floor and 0.8 to 0.9 seconds remain everywhere. Slower responses ranged from 1.1 seconds in Montreal to 1.9 in San Francisco. In the same daily collection, Claude typically took 1.4 seconds and Grok 2.6.

Where in the world is it fast?

San Francisco and Tokyo came out quickest at 971 ms, followed by Montreal (1007 ms), Amsterdam (1039 ms) and Singapore (1053 ms). Compare that to our unauthenticated probe of the same API, where regions spread 3.4× apart (70–235 ms): once the model's roughly constant thinking time is added on top, the spread compresses to 8%. Your region still sets a floor, and Singapore pays real physics on every call, but for a flagship reasoning model it is far from the dominant term. A city's own daily median moved more than that between days: Amsterdam ranged 786 to 1643 ms across the eight runs.

Amsterdam 1039ms typicalSan Francisco 971ms typicalMontreal 1007ms typicalSingapore 1053ms typicalTokyo 971ms typical
  • NetherlandsAmsterdam
    1039 msslower 1555 ms
  • United StatesSan Francisco
    971 msslower 1933 ms
  • CanadaMontreal
    1007 msslower 1092 ms
  • SingaporeSingapore
    1053 msslower 1732 ms
  • JapanTokyo
    971 msslower 1537 ms

If OpenAI is slow for you

Check OpenAI’s status page
What the number includes
Queueing and low-effort reasoning on OpenAI's side before the first token. The stream opens with the first token, so network setup is a small share.
Regions
The request starts at our test location. Network evidence does not tell us where OpenAI runs the model.
Comparison basis
Same prompt and low effort as the Claude and Grok surfaces, collected on the same days.

Where is the time actually going?

These timings come from parsing the token stream itself, not from transport phases alone. For OpenAI the stream only opens when the first token is ready: headers and first visible token arrived within 26 to 89 ms of each other in every city, so nearly all of the one-second wait is queueing plus the model's low-effort reasoning before the stream opens. Subtract the pure network floor we measured separately (70–235 ms by city) and the model-side share comes out at 0.81 to 0.92 s per region. Connection setup and TLS stay under 30 ms everywhere, which is Cloudflare's edge.

Click a city to see its stage-by-stage breakdown.

ConnectTLSServer waitDownload
NetherlandsAmsterdam
888 ms
Finding the server
1 ms
Reaching the server
8 ms
Setting up security
16 ms
Waiting for the server: most of the time
765 ms
Receiving the response
99 ms
Total
888 ms
Response opened
789 ms
First visible word
878 ms
Answer complete
888 ms
United StatesSan Francisco
766 ms
Finding the server
11 ms
Reaching the server
6 ms
Setting up security
17 ms
Waiting for the server: most of the time
703 ms
Receiving the response
40 ms
Total
766 ms
Response opened
685 ms
First visible word
752 ms
Answer complete
766 ms
CanadaMontreal
838 ms
Finding the server
2 ms
Reaching the server
21 ms
Setting up security
29 ms
Waiting for the server: most of the time
734 ms
Receiving the response
54 ms
Total
838 ms
Response opened
784 ms
First visible word
826 ms
Answer complete
838 ms
SingaporeSingapore
979 ms
Finding the server
3 ms
Reaching the server
5 ms
Setting up security
11 ms
Waiting for the server: most of the time
884 ms
Receiving the response
79 ms
Total
979 ms
Response opened
900 ms
First visible word
972 ms
Answer complete
979 ms
JapanTokyo
809 ms
Finding the server
1 ms
Reaching the server
3 ms
Setting up security
8 ms
Waiting for the server: most of the time
742 ms
Receiving the response
56 ms
Total
809 ms
Response opened
773 ms
First visible word
799 ms
Answer complete
809 ms

The bars and phases below come from the controlled run of 4 September 2026 (20 requests per city), so they add up within that run. The typical and slower times in the table above pool every daily run, which is why the two sets of numbers differ.

Finding the server is measured once per location, so it sits outside these bars. Open a city to see it.

Bars show the typical request, so the parts add up to its total. Why

Every city, side by side

One row per city, with the typical time highlighted. For the stage-by-stage detail behind any city, open it in the chart above.

ChatGPT API · 27 August 2026 – 4 September 2026

Test locationTypical first wordOn a bad daySlowest seen
Amsterdam
1039 ms1555 ms1876 ms
San Francisco
Fastest
971 ms1933 ms2590 ms
Montreal
1007 ms1092 ms1474 ms
Singapore
Slowest
1053 ms1732 ms2376 ms
Tokyo
971 ms1537 ms3958 ms

How it has changed over time

One panel per city, one point per run, all panels on the same scale. Daily runs and the controlled 20-request runs are shown as separate charts because they measure different things.

Time to first visible token from gpt-5.6-sol on a one-word prompt: streaming, reasoning_effort low, 256 max tokens, paid key.

Daily response time

Daily median of 5 requests, per city, 8 August 2026 to 7 September 2026. Open markers on the baseline are runs with fewer than 3 valid requests; they carry no value. A dashed line marks the day a city joined the measurements.

Amsterdam1370 ms
0100020008 Aug7 Septjoined 27 AugAmsterdam, 27 August 2026: 1643 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 28 August 2026: 862 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 29 August 2026: 1297 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 31 August 2026: 1156 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 1 September 2026: 786 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 2 September 2026: 1016 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 3 September 2026: 1047 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 4 September 2026: 850 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 5 September 2026: 939 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 6 September 2026: 963 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 7 September 2026: 1370 ms (daily median of 5 requests, 5 valid requests).
San Francisco1434 ms
0100020008 Aug7 Septjoined 27 AugSan Francisco, 27 August 2026: 1290 ms (daily median of 5 requests, 5 valid requests).San Francisco, 28 August 2026: 930 ms (daily median of 5 requests, 5 valid requests).San Francisco, 29 August 2026: 954 ms (daily median of 5 requests, 5 valid requests).San Francisco, 31 August 2026: 757 ms (daily median of 5 requests, 5 valid requests).San Francisco, 1 September 2026: 1431 ms (daily median of 5 requests, 5 valid requests).San Francisco, 2 September 2026: 1118 ms (daily median of 5 requests, 5 valid requests).San Francisco, 3 September 2026: 1002 ms (daily median of 5 requests, 5 valid requests).San Francisco, 4 September 2026: 1155 ms (daily median of 5 requests, 5 valid requests).San Francisco, 5 September 2026: 718 ms (daily median of 5 requests, 5 valid requests).San Francisco, 6 September 2026: 1058 ms (daily median of 5 requests, 5 valid requests).San Francisco, 7 September 2026: 1434 ms (daily median of 5 requests, 5 valid requests).
Montreal1031 ms
0100020008 Aug7 Septjoined 27 AugMontreal, 27 August 2026: 1012 ms (daily median of 5 requests, 5 valid requests).Montreal, 28 August 2026: 750 ms (daily median of 5 requests, 5 valid requests).Montreal, 29 August 2026: 635 ms (daily median of 5 requests, 5 valid requests).Montreal, 31 August 2026: 1152 ms (daily median of 5 requests, 5 valid requests).Montreal, 1 September 2026: 1165 ms (daily median of 5 requests, 5 valid requests).Montreal, 2 September 2026: 1012 ms (daily median of 5 requests, 5 valid requests).Montreal, 3 September 2026: 1145 ms (daily median of 5 requests, 5 valid requests).Montreal, 4 September 2026: 800 ms (daily median of 5 requests, 5 valid requests).Montreal, 5 September 2026: 1010 ms (daily median of 5 requests, 5 valid requests).Montreal, 6 September 2026: 700 ms (daily median of 5 requests, 5 valid requests).Montreal, 7 September 2026: 1031 ms (daily median of 5 requests, 5 valid requests).
Singapore1339 ms
0100020008 Aug7 Septjoined 27 AugSingapore, 27 August 2026: 1164 ms (daily median of 5 requests, 5 valid requests).Singapore, 28 August 2026: 1039 ms (daily median of 5 requests, 5 valid requests).Singapore, 29 August 2026: 744 ms (daily median of 5 requests, 5 valid requests).Singapore, 31 August 2026: 1007 ms (daily median of 5 requests, 5 valid requests).Singapore, 1 September 2026: 1127 ms (daily median of 5 requests, 5 valid requests).Singapore, 2 September 2026: 1296 ms (daily median of 5 requests, 5 valid requests).Singapore, 3 September 2026: 1233 ms (daily median of 5 requests, 5 valid requests).Singapore, 4 September 2026: 1027 ms (daily median of 5 requests, 5 valid requests).Singapore, 5 September 2026: 1136 ms (daily median of 5 requests, 5 valid requests).Singapore, 6 September 2026: 938 ms (daily median of 5 requests, 5 valid requests).Singapore, 7 September 2026: 1339 ms (daily median of 5 requests, 5 valid requests).
Tokyo959 ms
0100020008 Aug7 Septjoined 27 AugTokyo, 27 August 2026: 971 ms (daily median of 5 requests, 5 valid requests).Tokyo, 28 August 2026: 1029 ms (daily median of 5 requests, 5 valid requests).Tokyo, 29 August 2026: 935 ms (daily median of 5 requests, 5 valid requests).Tokyo, 31 August 2026: 986 ms (daily median of 5 requests, 5 valid requests).Tokyo, 1 September 2026: 996 ms (daily median of 5 requests, 5 valid requests).Tokyo, 2 September 2026: 967 ms (daily median of 5 requests, 5 valid requests).Tokyo, 3 September 2026: 999 ms (daily median of 5 requests, 5 valid requests).Tokyo, 4 September 2026: 795 ms (daily median of 5 requests, 5 valid requests).Tokyo, 5 September 2026: 881 ms (daily median of 5 requests, 5 valid requests).Tokyo, 6 September 2026: 951 ms (daily median of 5 requests, 5 valid requests).Tokyo, 7 September 2026: 959 ms (daily median of 5 requests, 5 valid requests).
Mumbai1476 ms
0100020008 Aug7 Septjoined 7 SeptMumbai, 7 September 2026: 1476 ms (daily median of 5 requests, 5 valid requests).

Collecting since 7 Sept

Show the numbers
RunAmsterdamSan FranciscoMontrealSingaporeTokyoMumbai
27 August 20261643 ms1290 ms1012 ms1164 ms971 ms
28 August 2026862 ms930 ms750 ms1039 ms1029 ms
29 August 20261297 ms954 ms635 ms744 ms935 ms
31 August 20261156 ms757 ms1152 ms1007 ms986 ms
1 September 2026786 ms1431 ms1165 ms1127 ms996 ms
2 September 20261016 ms1118 ms1012 ms1296 ms967 ms
3 September 20261047 ms1002 ms1145 ms1233 ms999 ms
4 September 2026850 ms1155 ms800 ms1027 ms795 ms
5 September 2026939 ms718 ms1010 ms1136 ms881 ms
6 September 2026963 ms1058 ms700 ms938 ms951 ms
7 September 20261370 ms1434 ms1031 ms1339 ms959 ms1476 ms

Slower response time in the controlled runs

Slower response time in each 20-request run, per city, 9 June 2026 to 7 September 2026. Open markers on the baseline are runs with fewer than 15 valid requests; they carry no value. A dashed line marks the day a city joined the measurements.

Amsterdam1410 ms
0125025009 Jun7 Septjoined 27 AugAmsterdam, 30 August 2026: 1699 ms (slower response time in each 20-request run, 20 valid requests).Amsterdam, 4 September 2026: 1410 ms (slower response time in each 20-request run, 20 valid requests).

Collecting since 27 Aug

San Francisco2179 ms
0125025009 Jun7 Septjoined 27 AugSan Francisco, 30 August 2026: 1686 ms (slower response time in each 20-request run, 20 valid requests).San Francisco, 4 September 2026: 2179 ms (slower response time in each 20-request run, 20 valid requests).

Collecting since 27 Aug

Montreal1187 ms
0125025009 Jun7 Septjoined 27 AugMontreal, 30 August 2026: 996 ms (slower response time in each 20-request run, 20 valid requests).Montreal, 4 September 2026: 1187 ms (slower response time in each 20-request run, 20 valid requests).

Collecting since 27 Aug

Singapore2270 ms
0125025009 Jun7 Septjoined 27 AugSingapore, 30 August 2026: 1194 ms (slower response time in each 20-request run, 20 valid requests).Singapore, 4 September 2026: 2270 ms (slower response time in each 20-request run, 20 valid requests).

Collecting since 27 Aug

Tokyo1678 ms
0125025009 Jun7 Septjoined 27 AugTokyo, 30 August 2026: 1395 ms (slower response time in each 20-request run, 20 valid requests).Tokyo, 4 September 2026: 1678 ms (slower response time in each 20-request run, 20 valid requests).

Collecting since 27 Aug

MumbaiNo measurements yet
0125025009 Jun7 Septjoined 7 Sept

Technical note: 95% of that run's requests finished within this time (technical: p95 of 20 requests).

Show the numbers
RunAmsterdamSan FranciscoMontrealSingaporeTokyoMumbai
30 August 20261699 ms1686 ms996 ms1194 ms1395 ms
4 September 20261410 ms2179 ms1187 ms2270 ms1678 ms

Global presence

Verdict

Served from one location

Every city reached the same address, and the far ones waited much longer. That looks like one origin server.

Based on 1 server address and a 215 ms spread between cities in the latest controlled run.

What these OpenAI results mean

The catalog contains two OpenAI requests. The primary is time to first visible token from gpt-5.6-sol at api.openai.com/v1/chat/completions: a one-word prompt, streaming, reasoning_effort low, 256 max tokens, paid key. It stands for the wait before the first word a user sees when calling the flagship model with light reasoning. The second is GET /v1/models with the same key, a small JSON list, which stands for the network and authentication path alone; no model runs, so it separates the front door from inference. Its results stay separate because a model-list response does not measure inference.

Not measured: the ChatGPT app at chatgpt.com, tokens per second after the first token, long prompts, other models, tool calls, the Batch API, and thinking at higher effort.

The published measurements cover 27 August to 4 September 2026. Typical first-token times range from 971 to 1053 ms across the five cities, while Amsterdam’s daily median ranges from 786 to 1643 ms. In this window the day-to-day variation in Amsterdam is larger than the difference between cities’ pooled typical times.

How this test was run

8 daily quick runs from 27 August 2026 to 4 September 2026 (40 requests per city), plus 2 controlled runs of 20 requests per city, the latest on 4 September 2026 · 27 August 2026 – 4 September 2026

Technical details
Request
Authenticated streaming POST to api.openai.com/v1/chat/completions
Measured from
Amsterdam · San Francisco · Montreal · Singapore · Tokyo
Model
gpt-5.6-sol · reasoning effort "low" · max_completion_tokens 256 · default temperature
Prompt
"Say 'ok' and nothing else." Identical across every city and every LLM API we benchmark
Network floor
Measured separately by an unauthenticated probe of the same hostname on 22 August 2026
Comparison figures
Claude (claude-opus-5, effort low) and Grok (grok-4.6, reasoning_effort low) come from the same daily collection over the same days
Timings taken
Parsed from the token stream: stream open, first visible token, completion, plus DNS, connect and TLS
Typical response time
Median of all 40 daily requests per city between 27 August 2026 and 4 September 2026; not an average of daily medians
Slower response time
Median of the 95th percentiles of the 2 most recent controlled 20-request runs per city
Day-to-day range (daily medians)
Amsterdam 786–1643 ms · San Francisco 757–1431 ms · Montreal 635–1165 ms · Singapore 744–1296 ms · Tokyo 795–1029 ms

ChatGPT API

https://api.openai.com/v1/chat/completions

Independent measurement by LatencyRadar. Not affiliated with OpenAI. · How we measure (v1.0)

Get the monthly AI latency report, by test location

One email a month: how fast the major AI providers respond, measured globally.

Check your inbox for the confirmation email. Nothing is sent until you confirm. Unsubscribe any time.

How fast is your API compared to OpenAI's?

Run a free speed test from multiple cities and find out where your users are waiting. No setup, no account required.

Test my API

Takes about 30 seconds.

More public benchmarks →