Benchmark report

DeepSeek API thumbnail

Is the DeepSeek API slow?

DeepSeek V4 Flash time to first token from 6 cities, measured daily

Uneven around the world

First visible token typically arrived in 573 ms from Tokyo and 761 ms from San Francisco, with every city under 0.8 seconds. Asian cities wait less than North American and European ones, which suggests the model is served from Asia.

Fastest
573 ms
Tokyo
Slowest
761 ms
San Francisco
Successful requests
100%
6 locations

Response time around the world

Amsterdam 734ms typicalSan Francisco 761ms typicalMontreal 732ms typicalSingapore 612ms typicalTokyo 573ms typicalMumbai 664ms typical
  • JapanTokyo
    573 ms
  • SingaporeSingapore
    612 ms
  • IndiaMumbai
    664 ms
  • CanadaMontreal
    732 ms
  • NetherlandsAmsterdam
    734 ms
  • United StatesSan Francisco
    761 ms

The stream opens before the first token does

In the controlled run of 7 September DeepSeek opened the stream 130 to 245 ms before the first visible token arrived: 113 ms against 319 ms in Singapore, 275 ms against 500 ms in Amsterdam. OpenAI, Anthropic and Mistral hold the response until the first token is ready, so a client that times the first byte will read DeepSeek as faster than it is. Connecting took 2 to 4 ms and the secure handshake 4 to 5 ms; the wait before the stream opened was 119 ms in Singapore, 172 in Mumbai, 177 in Tokyo, 258 in Amsterdam, 316 in San Francisco and 319 in Montreal. Every city reached the same CloudFront address through a local point of presence, so that wait is the trip from the edge to DeepSeek, and it is shortest from Asia.

Click the location to see each stage.

ConnectTLSServer waitDownload
SingaporeSingapore
388 ms
Finding the server
2 ms
Reaching the server
3 ms
Setting up security
5 ms
Waiting for the server
119 ms
Receiving the response: most of the time
261 ms
Total
388 ms
Response opened
113 ms
First visible word
319 ms
Answer complete
388 ms

Finding the server is measured once per location, so it sits outside these bars. Open a city to see it.

Bars show the typical request, so the parts add up to its total. Why

Compare all locations

Click a city to see its stage-by-stage breakdown.

ConnectTLSServer waitDownload
NetherlandsAmsterdam
579 ms
Finding the server
1 ms
Reaching the server
4 ms
Setting up security
4 ms
Waiting for the server
258 ms
Receiving the response: most of the time
313 ms
Total
579 ms
Response opened
275 ms
First visible word
500 ms
Answer complete
579 ms
United StatesSan Francisco
592 ms
Finding the server
11 ms
Reaching the server
4 ms
Setting up security
5 ms
Waiting for the server: most of the time
316 ms
Receiving the response
267 ms
Total
592 ms
Response opened
288 ms
First visible word
524 ms
Answer complete
592 ms
CanadaMontreal
611 ms
Finding the server
5 ms
Reaching the server
3 ms
Setting up security
4 ms
Waiting for the server: most of the time
319 ms
Receiving the response
285 ms
Total
611 ms
Response opened
327 ms
First visible word
546 ms
Answer complete
611 ms
SingaporeSingapore
388 ms
Finding the server
2 ms
Reaching the server
3 ms
Setting up security
5 ms
Waiting for the server
119 ms
Receiving the response: most of the time
261 ms
Total
388 ms
Response opened
113 ms
First visible word
319 ms
Answer complete
388 ms
JapanTokyo
486 ms
Finding the server
1 ms
Reaching the server
3 ms
Setting up security
4 ms
Waiting for the server
177 ms
Receiving the response: most of the time
302 ms
Total
486 ms
Response opened
178 ms
First visible word
423 ms
Answer complete
486 ms
IndiaMumbai
404 ms
Finding the server
8 ms
Reaching the server
2 ms
Setting up security
5 ms
Waiting for the server
172 ms
Receiving the response: most of the time
225 ms
Total
404 ms
Response opened
173 ms
First visible word
307 ms
Answer complete
404 ms

The phases come from the controlled run of 7 September 2026. Typical times pool daily requests; slower times use controlled full runs. 6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai.

Finding the server is measured once per location, so it sits outside these bars. Open a city to see it.

Bars show the typical request, so the parts add up to its total. Why

DeepSeek still feels slow?

Check DeepSeek’s status page
Incidents and degraded performance are posted there. A slow first token during one is not your region.
Possibly slower at midday in China
Our daily runs land between 03:30 and 06:50 UTC, which is late morning to early afternoon in China. The controlled run on Sunday 7 September at 21:33 UTC came back 307 to 546 ms per city, about 200 ms faster than the daily typicals. We have one evening run, so treat it as a hint. 15 September was the slowest day everywhere; Tokyo's median that day was 1.4 seconds against 0.4 to 0.8 seconds on the other thirteen.
What this does not measure
DeepSeek with thinking on, the chat app, tokens per second, long prompts, or other DeepSeek models. The ChatGPT, Claude and Grok numbers on this page use a different request: the same prompt at low reasoning effort, which adds a short thinking phase this request turns off.

Response time over the last 16 days

Daily median of 5 requests per location, 5 September 2026 to 21 September 2026.
010002000 ms5 Sept21 SeptMumbai joinedTokyo, 15 September 2026: 1406 ms, well above its usual level.

A dot marks a day well above that location’s usual level.

Show the numbers
DayAmsterdamSan FranciscoMontrealSingaporeTokyoMumbai
5 September 2026776 ms714 ms853 ms396 ms604 ms
6 September 2026561 ms493 ms543 ms296 ms331 ms
7 September 2026690 ms623 ms612 ms360 ms621 ms396 ms
8 September 2026747 ms614 ms459 ms470 ms516 ms560 ms
9 September 2026443 ms508 ms644 ms389 ms424 ms413 ms
10 September 2026779 ms888 ms905 ms687 ms782 ms765 ms
11 September 2026745 ms774 ms626 ms536 ms596 ms557 ms
12 September 2026714 ms631 ms791 ms658 ms395 ms613 ms
13 September 2026789 ms583 ms825 ms590 ms688 ms664 ms
14 September 2026823 ms837 ms923 ms701 ms780 ms737 ms
15 September 2026891 ms912 ms942 ms781 ms1406 ms ▲867 ms
16 September 2026734 ms732 ms932 ms650 ms605 ms761 ms
17 September 2026670 ms810 ms662 ms492 ms419 ms525 ms
18 September 2026771 ms798 ms692 ms631 ms550 ms800 ms
19 September 2026670 ms690 ms726 ms459 ms526 ms589 ms
20 September 2026662 ms734 ms704 ms454 ms552 ms575 ms
21 September 2026735 ms876 ms802 ms665 ms569 ms656 ms

About this measurement

POST api.deepseek.com/chat/completions · 70 requests per location over 14 daily runs · 8 September 2026 to 21 September 2026

What we tested · AI response. Time to first visible token from deepseek-v4-flash on a one-word prompt: streaming, thinking disabled, temperature 0, 256 max tokens, paid key.

How LatencyRadar measures response time →

Technical details

Doesn’t measure: DeepSeek with thinking on, the chat app, tokens per second or long prompts. Not the same request as the low-effort direct providers.

One DeepSeek surface: time to first visible token from deepseek-v4-flash at api.deepseek.com/chat/completions, a one-word prompt, streaming, thinking disabled, temperature 0, 256 max tokens, paid key. Thinking is off because DeepSeek V4 Flash thinks at high effort by default, and the request is kept identical to the OpenRouter DeepSeek V4 Flash routes on the provider matrix, so you can read direct and routed side by side.

Not measured: DeepSeek with thinking on, the chat app, tokens per second, long prompts, or other DeepSeek models. The ChatGPT, Claude and Grok figures on this page are the same prompt on the same days at low reasoning effort. Their daily medians ran 0.6 to 2.4 seconds (ChatGPT), 1.2 to 2.3 seconds (Claude) and 1.1 to 3.7 seconds (Grok) while DeepSeek's ran 0.4 to 0.9 seconds outside one slow day. Part of that gap is thinking this request turns off, so the two sets of numbers describe different requests and are not a ranking.

Typical first-token times over 8 to 21 September 2026 are 573 ms in Tokyo to 761 ms in San Francisco, with Singapore at 612 ms, Mumbai at 664 ms, Montreal at 732 ms and Amsterdam at 734 ms. Every city shares one CloudFront address and the wait is shortest from Asia, which suggests that is where the model runs. Day to day, medians moved inside a band of 390 to 480 ms in every city, apart from Tokyo's one slow day, and 15 September was the slowest day in all six. In the controlled run the slower responses sat 110 to 270 ms above that run's median, the widest gap in Singapore (590 ms against 319 ms).

Test locationTypical response timeSlower response timeRequests
Amsterdam734 ms698 ms20
San Francisco761 ms712 ms20
Montreal732 ms653 ms20
Singapore612 ms590 ms20
Tokyo573 ms621 ms20
Mumbai664 ms516 ms20

Typical: half of the requests finished within this time (technical: p50 of time to first visible token). Slower: 95% of requests finished within this time (technical: p95). Statistics

Request
Authenticated streaming POST to api.deepseek.com/chat/completions
Measured from
Amsterdam · San Francisco · Montreal · Singapore · Tokyo · Mumbai
Requests
Daily measurements from 8 September 2026 to 21 September 2026. 6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai. · 8 September 2026 to 21 September 2026
Model
deepseek-v4-flash · thinking disabled · temperature 0 · max_tokens 256
Prompt
"Say 'ok' and nothing else." Identical across every city and every LLM API we benchmark
Comparison figures
ChatGPT (gpt-5.6-sol), Claude (claude-opus-5) and Grok (grok-4.6), all at low reasoning effort, from the same daily collection between 8 and 21 September 2026. Those three think briefly before the first token; this request tells DeepSeek not to
Timings taken
Parsed from the token stream: stream open, first visible token, completion, plus DNS, connect and TLS
Report coverage
6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai.
Daily requests in each location
Amsterdam: 70 · San Francisco: 70 · Montreal: 70 · Singapore: 70 · Tokyo: 70 · Mumbai: 70
Typical response time
Median of the eligible daily requests in each included location.
Slower response time
Median of recent controlled full-run p95 values in each included location.

DeepSeek API

https://api.deepseek.com/chat/completions

Independent measurement by LatencyRadar. Not affiliated with DeepSeek. · How we measure (v1.1)

How fast is your API compared to DeepSeek's?

Run a free speed test from multiple cities and find out where your users are waiting. No setup, no account required.

Test my API

No account required · Takes about 30 seconds.

More benchmarks

All benchmarks →