Benchmark report
Grok API thumbnail

Is the Grok API slow? grok-4.6 time to first token from 5 cities, measured daily

We ran real, authenticated streaming completions against xAI's grok-4.6, with reasoning effort set to "low", from Amsterdam, San Francisco, Montreal, Singapore and Tokyo, every day from 27 August to 4 September 2026, plus two controlled runs of 20 requests per city, parsing the token stream ourselves and timestamping the moment the first visible token arrived.

What we tested · AI response

Time to first visible token from grok-4.6 on a one-word prompt: streaming, reasoning_effort low, 256 max tokens, paid key.

Doesn’t measure: Grok inside X or at grok.com, tokens per second, long prompts or other xAI models.

Quick take

Grok's first visible token typically arrived in 2.45 to 2.73 seconds, measured 40 times per city over eight days, an 11% spread, so where you call from barely matters. The stream opens in 0.13 to 0.29 seconds and reasoning deltas start about half a second in; the rest is the model thinking. Plan on two and a half seconds before the first visible word and 3.3 to 3.9 seconds for slower responses. ChatGPT typically took 1.0 second and Claude 1.4 in the same daily collection.

Where in the world is it fast?

Montreal was quickest at 2.45 s to the first visible token, then San Francisco (2.57 s), Amsterdam (2.57 s), Tokyo (2.69 s) and Singapore (2.73 s). A 275 ms spread on a two-and-a-half-second wait means geography is nearly irrelevant for grok-4.6: when the model spends most of the time thinking, moving closer to the servers buys you very little. The day matters more than the city. Tokyo's daily median ranged from 1.2 to 3.2 s across the eight runs.

Amsterdam 2572ms typicalSan Francisco 2565ms typicalMontreal 2453ms typicalSingapore 2728ms typicalTokyo 2691ms typical
  • NetherlandsAmsterdam
    2572 msslower 3565 ms
  • United StatesSan Francisco
    2565 msslower 3809 ms
  • CanadaMontreal
    2453 msslower 3357 ms
  • SingaporeSingapore
    2728 msslower 3920 ms
  • JapanTokyo
    2691 msslower 3332 ms

If xAI is slow for you

Check xAI’s status page
What the number includes
Model queueing and low-effort reasoning. The stream opens early, so a fast stream open does not mean a fast first word.
Regions
The request starts at our test location. Network evidence does not tell us where xAI runs the model.
Comparison basis
Same prompt and low effort as the ChatGPT and Claude surfaces, collected on the same days.

Where is the time actually going?

Measured from the stream itself, one grok-4.6 call has three acts: the stream opens in 0.13–0.29 s (that part is network and edge), reasoning deltas start arriving around 0.5–0.65 s, and the first visible token lands at 2.1–2.6 s in the latest controlled run, with the trivial answer completing within a few milliseconds after it. Unlike the ChatGPT and Claude APIs, which hold the response until the first token is ready, Grok's wait happens inside an open stream. The first visible word is what a user sees appear, so that is the headline number here. The transport phases below are supporting detail, and the long "download" phase is the in-stream reasoning, not bytes on the wire.

Click a city to see its stage-by-stage breakdown.

ConnectTLSServer waitDownload
NetherlandsAmsterdam
2612 ms
Finding the server
1 ms
Reaching the server
10 ms
Setting up security
22 ms
Waiting for the server
381 ms
Receiving the response: most of the time
2199 ms
Total
2612 ms
Response opened
281 ms
First visible word
2609 ms
Answer complete
2612 ms
United StatesSan Francisco
2131 ms
Finding the server
15 ms
Reaching the server
4 ms
Setting up security
8 ms
Waiting for the server
99 ms
Receiving the response: most of the time
2020 ms
Total
2131 ms
Response opened
128 ms
First visible word
2130 ms
Answer complete
2131 ms
CanadaMontreal
2556 ms
Finding the server
1 ms
Reaching the server
22 ms
Setting up security
29 ms
Waiting for the server
156 ms
Receiving the response: most of the time
2349 ms
Total
2556 ms
Response opened
210 ms
First visible word
2555 ms
Answer complete
2556 ms
SingaporeSingapore
2477 ms
Finding the server
6 ms
Reaching the server
3 ms
Setting up security
8 ms
Waiting for the server
253 ms
Receiving the response: most of the time
2213 ms
Total
2477 ms
Response opened
286 ms
First visible word
2474 ms
Answer complete
2477 ms
JapanTokyo
2243 ms
Finding the server
4 ms
Reaching the server
3 ms
Setting up security
6 ms
Waiting for the server
190 ms
Receiving the response: most of the time
2044 ms
Total
2243 ms
Response opened
201 ms
First visible word
2242 ms
Answer complete
2243 ms

The bars and phases below come from the controlled run of 4 September 2026 (20 requests per city), so they add up within that run. The typical and slower times in the table above pool every daily run, which is why the two sets of numbers differ.

Finding the server is measured once per location, so it sits outside these bars. Open a city to see it.

Bars show the typical request, so the parts add up to its total. Why

Every city, side by side

One row per city, with the typical time highlighted. For the stage-by-stage detail behind any city, open it in the chart above.

Grok API · 27 August 2026 – 4 September 2026

Test locationTypical first wordOn a bad daySlowest seen
Amsterdam
2572 ms3565 ms3105 ms
San Francisco
2565 ms3809 ms4406 ms
Montreal
Fastest
2453 ms3357 ms3999 ms
Singapore
Slowest
2728 ms3920 ms4000 ms
Tokyo
2691 ms3332 ms3733 ms

How it has changed over time

One panel per city, one point per run, all panels on the same scale. Daily runs and the controlled 20-request runs are shown as separate charts because they measure different things.

Time to first visible token from grok-4.6 on a one-word prompt: streaming, reasoning_effort low, 256 max tokens, paid key.

Daily response time

Daily median of 5 requests, per city, 8 August 2026 to 7 September 2026. Open markers on the baseline are runs with fewer than 3 valid requests; they carry no value. A dashed line marks the day a city joined the measurements.

Amsterdam2507 ms
0250050008 Aug7 Septjoined 27 AugAmsterdam, 27 August 2026: 2431 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 28 August 2026: 2451 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 29 August 2026: 3216 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 31 August 2026: 3506 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 1 September 2026: 2703 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 2 September 2026: 2572 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 3 September 2026: 2170 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 4 September 2026: 3439 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 5 September 2026: 2883 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 6 September 2026: 2653 ms (daily median of 5 requests, 5 valid requests).Amsterdam, 7 September 2026: 2507 ms (daily median of 5 requests, 5 valid requests).
San Francisco2363 ms
0250050008 Aug7 Septjoined 27 AugSan Francisco, 27 August 2026: 1988 ms (daily median of 5 requests, 5 valid requests).San Francisco, 28 August 2026: 3093 ms (daily median of 5 requests, 5 valid requests).San Francisco, 29 August 2026: 2259 ms (daily median of 5 requests, 5 valid requests).San Francisco, 31 August 2026: 2390 ms (daily median of 5 requests, 5 valid requests).San Francisco, 1 September 2026: 2570 ms (daily median of 5 requests, 5 valid requests).San Francisco, 2 September 2026: 2669 ms (daily median of 5 requests, 5 valid requests).San Francisco, 3 September 2026: 2198 ms (daily median of 5 requests, 5 valid requests).San Francisco, 4 September 2026: 3345 ms (daily median of 5 requests, 5 valid requests).San Francisco, 5 September 2026: 1936 ms (daily median of 5 requests, 5 valid requests).San Francisco, 6 September 2026: 2323 ms (daily median of 5 requests, 5 valid requests).San Francisco, 7 September 2026: 2363 ms (daily median of 5 requests, 5 valid requests).
Montreal2608 ms
0250050008 Aug7 Septjoined 27 AugMontreal, 27 August 2026: 2331 ms (daily median of 5 requests, 5 valid requests).Montreal, 28 August 2026: 2038 ms (daily median of 5 requests, 5 valid requests).Montreal, 29 August 2026: 2841 ms (daily median of 5 requests, 5 valid requests).Montreal, 31 August 2026: 2016 ms (daily median of 5 requests, 5 valid requests).Montreal, 1 September 2026: 3289 ms (daily median of 5 requests, 5 valid requests).Montreal, 2 September 2026: 2419 ms (daily median of 5 requests, 5 valid requests).Montreal, 3 September 2026: 2383 ms (daily median of 5 requests, 5 valid requests).Montreal, 4 September 2026: 2502 ms (daily median of 5 requests, 5 valid requests).Montreal, 5 September 2026: 2161 ms (daily median of 5 requests, 5 valid requests).Montreal, 6 September 2026: 2311 ms (daily median of 5 requests, 5 valid requests).Montreal, 7 September 2026: 2608 ms (daily median of 5 requests, 5 valid requests).
Singapore3066 ms
0250050008 Aug7 Septjoined 27 AugSingapore, 27 August 2026: 2874 ms (daily median of 5 requests, 5 valid requests).Singapore, 28 August 2026: 2596 ms (daily median of 5 requests, 5 valid requests).Singapore, 29 August 2026: 2728 ms (daily median of 5 requests, 5 valid requests).Singapore, 31 August 2026: 3061 ms (daily median of 5 requests, 5 valid requests).Singapore, 1 September 2026: 2643 ms (daily median of 5 requests, 5 valid requests).Singapore, 2 September 2026: 2856 ms (daily median of 5 requests, 5 valid requests).Singapore, 3 September 2026: 2421 ms (daily median of 5 requests, 5 valid requests).Singapore, 4 September 2026: 2678 ms (daily median of 5 requests, 5 valid requests).Singapore, 5 September 2026: 3742 ms (daily median of 5 requests, 5 valid requests).Singapore, 6 September 2026: 2863 ms (daily median of 5 requests, 5 valid requests).Singapore, 7 September 2026: 3066 ms (daily median of 5 requests, 5 valid requests).
Tokyo2059 ms
0250050008 Aug7 Septjoined 27 AugTokyo, 27 August 2026: 2992 ms (daily median of 5 requests, 5 valid requests).Tokyo, 28 August 2026: 1218 ms (daily median of 5 requests, 5 valid requests).Tokyo, 29 August 2026: 2405 ms (daily median of 5 requests, 5 valid requests).Tokyo, 31 August 2026: 3150 ms (daily median of 5 requests, 5 valid requests).Tokyo, 1 September 2026: 2675 ms (daily median of 5 requests, 5 valid requests).Tokyo, 2 September 2026: 2877 ms (daily median of 5 requests, 5 valid requests).Tokyo, 3 September 2026: 2108 ms (daily median of 5 requests, 5 valid requests).Tokyo, 4 September 2026: 2998 ms (daily median of 5 requests, 5 valid requests).Tokyo, 5 September 2026: 2758 ms (daily median of 5 requests, 5 valid requests).Tokyo, 6 September 2026: 2524 ms (daily median of 5 requests, 5 valid requests).Tokyo, 7 September 2026: 2059 ms (daily median of 5 requests, 5 valid requests).
Mumbai3001 ms
0250050008 Aug7 Septjoined 7 SeptMumbai, 7 September 2026: 3001 ms (daily median of 5 requests, 5 valid requests).

Collecting since 7 Sept

Show the numbers
RunAmsterdamSan FranciscoMontrealSingaporeTokyoMumbai
27 August 20262431 ms1988 ms2331 ms2874 ms2992 ms
28 August 20262451 ms3093 ms2038 ms2596 ms1218 ms
29 August 20263216 ms2259 ms2841 ms2728 ms2405 ms
31 August 20263506 ms2390 ms2016 ms3061 ms3150 ms
1 September 20262703 ms2570 ms3289 ms2643 ms2675 ms
2 September 20262572 ms2669 ms2419 ms2856 ms2877 ms
3 September 20262170 ms2198 ms2383 ms2421 ms2108 ms
4 September 20263439 ms3345 ms2502 ms2678 ms2998 ms
5 September 20262883 ms1936 ms2161 ms3742 ms2758 ms
6 September 20262653 ms2323 ms2311 ms2863 ms2524 ms
7 September 20262507 ms2363 ms2608 ms3066 ms2059 ms3001 ms

Slower response time in the controlled runs

Slower response time in each 20-request run, per city, 9 June 2026 to 7 September 2026. Open markers on the baseline are runs with fewer than 15 valid requests; they carry no value. A dashed line marks the day a city joined the measurements.

Amsterdam3043 ms
0250050009 Jun7 Septjoined 27 AugAmsterdam, 30 August 2026: 4087 ms (slower response time in each 20-request run, 20 valid requests).Amsterdam, 4 September 2026: 3043 ms (slower response time in each 20-request run, 20 valid requests).

Collecting since 27 Aug

San Francisco4087 ms
0250050009 Jun7 Septjoined 27 AugSan Francisco, 30 August 2026: 3531 ms (slower response time in each 20-request run, 20 valid requests).San Francisco, 4 September 2026: 4087 ms (slower response time in each 20-request run, 20 valid requests).

Collecting since 27 Aug

Montreal3486 ms
0250050009 Jun7 Septjoined 27 AugMontreal, 30 August 2026: 3228 ms (slower response time in each 20-request run, 20 valid requests).Montreal, 4 September 2026: 3486 ms (slower response time in each 20-request run, 20 valid requests).

Collecting since 27 Aug

Singapore3766 ms
0250050009 Jun7 Septjoined 27 AugSingapore, 30 August 2026: 4073 ms (slower response time in each 20-request run, 20 valid requests).Singapore, 4 September 2026: 3766 ms (slower response time in each 20-request run, 20 valid requests).

Collecting since 27 Aug

Tokyo3403 ms
0250050009 Jun7 Septjoined 27 AugTokyo, 30 August 2026: 3260 ms (slower response time in each 20-request run, 20 valid requests).Tokyo, 4 September 2026: 3403 ms (slower response time in each 20-request run, 20 valid requests).

Collecting since 27 Aug

MumbaiNo measurements yet
0250050009 Jun7 Septjoined 7 Sept

Technical note: 95% of that run's requests finished within this time (technical: p95 of 20 requests).

Show the numbers
RunAmsterdamSan FranciscoMontrealSingaporeTokyoMumbai
30 August 20264087 ms3531 ms3228 ms4073 ms3260 ms
4 September 20263043 ms4087 ms3486 ms3766 ms3403 ms

Global presence

Verdict

Uneven around the world

More than one server answers, but some cities wait noticeably longer, so some visitors may be reaching a distant one.

Based on 2 server addresses and a 158 ms spread between cities in the latest controlled run.

What these xAI results mean

One xAI surface: time to first visible token from grok-4.6 at api.x.ai/v1/chat/completions, a one-word prompt, streaming, reasoning_effort low, 256 max tokens, paid key. xAI opens the stream before the first token arrives, so stream-open time and first-token time differ here more than for OpenAI or Anthropic. We report the first visible token as the measurement, never the stream opening.

Not measured: Grok inside X or at grok.com, tokens per second, long prompts, other xAI models, image or voice.

In the published measurements, typical first-token times range from about 2.45 to 2.73 seconds across five cities. The stream can open before visible text arrives, so an early connection alone would make the wait appear shorter than it is. The first-token measurement captures the delay a user sees before the answer starts.

How this test was run

8 daily quick runs from 27 August 2026 to 4 September 2026 (40 requests per city), plus 2 controlled runs of 20 requests per city, the latest on 4 September 2026 · 27 August 2026 – 4 September 2026

Technical details
Request
Authenticated streaming POST to api.x.ai/v1/chat/completions
Measured from
Amsterdam · San Francisco · Montreal · Singapore · Tokyo
Model
grok-4.6 · reasoning effort "low" · max_tokens 256 · default temperature
Prompt
"Say 'ok' and nothing else." Identical across every city and every LLM API we benchmark
Reasoning stream
xAI reasons inside the open stream, so we also timestamp the first reasoning delta; it never counts as the first visible token
Comparison figures
ChatGPT (gpt-5.6-sol, reasoning_effort low) and Claude (claude-opus-5, effort low) come from the same daily collection over the same days
Timings taken
Parsed from the token stream: stream open, first visible token, completion, plus DNS, connect and TLS
Typical response time
Median of all 40 daily requests per city between 27 August 2026 and 4 September 2026; not an average of daily medians
Slower response time
Median of the 95th percentiles of the 2 most recent controlled 20-request runs per city
Day-to-day range (daily medians)
Amsterdam 2170–3506 ms · San Francisco 1988–3345 ms · Montreal 2016–3289 ms · Singapore 2421–3061 ms · Tokyo 1218–3150 ms

Grok API

https://api.x.ai/v1/chat/completions

Independent measurement by LatencyRadar. Not affiliated with xAI. · How we measure (v1.0)

Get the monthly AI latency report, by test location

One email a month: how fast the major AI providers respond, measured globally.

Check your inbox for the confirmation email. Nothing is sent until you confirm. Unsubscribe any time.

How fast is your API compared to xAI's?

Run a free speed test from multiple cities and find out where your users are waiting. No setup, no account required.

Test my API

Takes about 30 seconds.

More public benchmarks →