Benchmark report
Is the DeepSeek API slow?
DeepSeek V4 Flash time to first token from 6 cities, measured daily
Uneven around the world
First visible token typically arrived in 573 ms from Tokyo and 761 ms from San Francisco, with every city under 0.8 seconds. Asian cities wait less than North American and European ones, which suggests the model is served from Asia.
- Fastest
- 573 ms
- Tokyo
- Slowest
- 761 ms
- San Francisco
- Successful requests
- 100%
- 6 locations
Response time around the world
Tokyo
573 msSingapore
612 msMumbai
664 msMontreal
732 msAmsterdam
734 msSan Francisco
761 ms
The stream opens before the first token does
In the controlled run of 7 September DeepSeek opened the stream 130 to 245 ms before the first visible token arrived: 113 ms against 319 ms in Singapore, 275 ms against 500 ms in Amsterdam. OpenAI, Anthropic and Mistral hold the response until the first token is ready, so a client that times the first byte will read DeepSeek as faster than it is. Connecting took 2 to 4 ms and the secure handshake 4 to 5 ms; the wait before the stream opened was 119 ms in Singapore, 172 in Mumbai, 177 in Tokyo, 258 in Amsterdam, 316 in San Francisco and 319 in Montreal. Every city reached the same CloudFront address through a local point of presence, so that wait is the trip from the edge to DeepSeek, and it is shortest from Asia.
Click the location to see each stage.
Singapore388 ms
- Finding the server
- 2 ms
- Reaching the server
- 3 ms
- Setting up security
- 5 ms
- Waiting for the server
- 119 ms
- Receiving the response: most of the time
- 261 ms
- Total
- 388 ms
- Response opened
- 113 ms
- First visible word
- 319 ms
- Answer complete
- 388 ms
Finding the server is measured once per location, so it sits outside these bars. Open a city to see it.
Bars show the typical request, so the parts add up to its total. Why
Compare all locationsHide the comparison
Click a city to see its stage-by-stage breakdown.
Amsterdam579 ms
- Finding the server
- 1 ms
- Reaching the server
- 4 ms
- Setting up security
- 4 ms
- Waiting for the server
- 258 ms
- Receiving the response: most of the time
- 313 ms
- Total
- 579 ms
- Response opened
- 275 ms
- First visible word
- 500 ms
- Answer complete
- 579 ms
San Francisco592 ms
- Finding the server
- 11 ms
- Reaching the server
- 4 ms
- Setting up security
- 5 ms
- Waiting for the server: most of the time
- 316 ms
- Receiving the response
- 267 ms
- Total
- 592 ms
- Response opened
- 288 ms
- First visible word
- 524 ms
- Answer complete
- 592 ms
Montreal611 ms
- Finding the server
- 5 ms
- Reaching the server
- 3 ms
- Setting up security
- 4 ms
- Waiting for the server: most of the time
- 319 ms
- Receiving the response
- 285 ms
- Total
- 611 ms
- Response opened
- 327 ms
- First visible word
- 546 ms
- Answer complete
- 611 ms
Singapore388 ms
- Finding the server
- 2 ms
- Reaching the server
- 3 ms
- Setting up security
- 5 ms
- Waiting for the server
- 119 ms
- Receiving the response: most of the time
- 261 ms
- Total
- 388 ms
- Response opened
- 113 ms
- First visible word
- 319 ms
- Answer complete
- 388 ms
Tokyo486 ms
- Finding the server
- 1 ms
- Reaching the server
- 3 ms
- Setting up security
- 4 ms
- Waiting for the server
- 177 ms
- Receiving the response: most of the time
- 302 ms
- Total
- 486 ms
- Response opened
- 178 ms
- First visible word
- 423 ms
- Answer complete
- 486 ms
Mumbai404 ms
- Finding the server
- 8 ms
- Reaching the server
- 2 ms
- Setting up security
- 5 ms
- Waiting for the server
- 172 ms
- Receiving the response: most of the time
- 225 ms
- Total
- 404 ms
- Response opened
- 173 ms
- First visible word
- 307 ms
- Answer complete
- 404 ms
The phases come from the controlled run of 7 September 2026. Typical times pool daily requests; slower times use controlled full runs. 6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai.
Finding the server is measured once per location, so it sits outside these bars. Open a city to see it.
Bars show the typical request, so the parts add up to its total. Why
DeepSeek still feels slow?
- Check DeepSeek’s status page
- Incidents and degraded performance are posted there. A slow first token during one is not your region.
- Possibly slower at midday in China
- Our daily runs land between 03:30 and 06:50 UTC, which is late morning to early afternoon in China. The controlled run on Sunday 7 September at 21:33 UTC came back 307 to 546 ms per city, about 200 ms faster than the daily typicals. We have one evening run, so treat it as a hint. 15 September was the slowest day everywhere; Tokyo's median that day was 1.4 seconds against 0.4 to 0.8 seconds on the other thirteen.
- What this does not measure
- DeepSeek with thinking on, the chat app, tokens per second, long prompts, or other DeepSeek models. The ChatGPT, Claude and Grok numbers on this page use a different request: the same prompt at low reasoning effort, which adds a short thinking phase this request turns off.
Response time over the last 16 days
A dot marks a day well above that location’s usual level.
Show the numbers
| Day | Amsterdam | San Francisco | Montreal | Singapore | Tokyo | Mumbai |
|---|---|---|---|---|---|---|
| 5 September 2026 | 776 ms | 714 ms | 853 ms | 396 ms | 604 ms | — |
| 6 September 2026 | 561 ms | 493 ms | 543 ms | 296 ms | 331 ms | — |
| 7 September 2026 | 690 ms | 623 ms | 612 ms | 360 ms | 621 ms | 396 ms |
| 8 September 2026 | 747 ms | 614 ms | 459 ms | 470 ms | 516 ms | 560 ms |
| 9 September 2026 | 443 ms | 508 ms | 644 ms | 389 ms | 424 ms | 413 ms |
| 10 September 2026 | 779 ms | 888 ms | 905 ms | 687 ms | 782 ms | 765 ms |
| 11 September 2026 | 745 ms | 774 ms | 626 ms | 536 ms | 596 ms | 557 ms |
| 12 September 2026 | 714 ms | 631 ms | 791 ms | 658 ms | 395 ms | 613 ms |
| 13 September 2026 | 789 ms | 583 ms | 825 ms | 590 ms | 688 ms | 664 ms |
| 14 September 2026 | 823 ms | 837 ms | 923 ms | 701 ms | 780 ms | 737 ms |
| 15 September 2026 | 891 ms | 912 ms | 942 ms | 781 ms | 1406 ms ▲ | 867 ms |
| 16 September 2026 | 734 ms | 732 ms | 932 ms | 650 ms | 605 ms | 761 ms |
| 17 September 2026 | 670 ms | 810 ms | 662 ms | 492 ms | 419 ms | 525 ms |
| 18 September 2026 | 771 ms | 798 ms | 692 ms | 631 ms | 550 ms | 800 ms |
| 19 September 2026 | 670 ms | 690 ms | 726 ms | 459 ms | 526 ms | 589 ms |
| 20 September 2026 | 662 ms | 734 ms | 704 ms | 454 ms | 552 ms | 575 ms |
| 21 September 2026 | 735 ms | 876 ms | 802 ms | 665 ms | 569 ms | 656 ms |
About this measurement
POST api.deepseek.com/chat/completions · 70 requests per location over 14 daily runs · 8 September 2026 to 21 September 2026
What we tested · AI response. Time to first visible token from deepseek-v4-flash on a one-word prompt: streaming, thinking disabled, temperature 0, 256 max tokens, paid key.
How LatencyRadar measures response time →
Technical details
Doesn’t measure: DeepSeek with thinking on, the chat app, tokens per second or long prompts. Not the same request as the low-effort direct providers.
One DeepSeek surface: time to first visible token from deepseek-v4-flash at api.deepseek.com/chat/completions, a one-word prompt, streaming, thinking disabled, temperature 0, 256 max tokens, paid key. Thinking is off because DeepSeek V4 Flash thinks at high effort by default, and the request is kept identical to the OpenRouter DeepSeek V4 Flash routes on the provider matrix, so you can read direct and routed side by side.
Not measured: DeepSeek with thinking on, the chat app, tokens per second, long prompts, or other DeepSeek models. The ChatGPT, Claude and Grok figures on this page are the same prompt on the same days at low reasoning effort. Their daily medians ran 0.6 to 2.4 seconds (ChatGPT), 1.2 to 2.3 seconds (Claude) and 1.1 to 3.7 seconds (Grok) while DeepSeek's ran 0.4 to 0.9 seconds outside one slow day. Part of that gap is thinking this request turns off, so the two sets of numbers describe different requests and are not a ranking.
Typical first-token times over 8 to 21 September 2026 are 573 ms in Tokyo to 761 ms in San Francisco, with Singapore at 612 ms, Mumbai at 664 ms, Montreal at 732 ms and Amsterdam at 734 ms. Every city shares one CloudFront address and the wait is shortest from Asia, which suggests that is where the model runs. Day to day, medians moved inside a band of 390 to 480 ms in every city, apart from Tokyo's one slow day, and 15 September was the slowest day in all six. In the controlled run the slower responses sat 110 to 270 ms above that run's median, the widest gap in Singapore (590 ms against 319 ms).
| Test location | Typical response time | Slower response time | Requests |
|---|---|---|---|
| Amsterdam | 734 ms | 698 ms | 20 |
| San Francisco | 761 ms | 712 ms | 20 |
| Montreal | 732 ms | 653 ms | 20 |
| Singapore | 612 ms | 590 ms | 20 |
| Tokyo | 573 ms | 621 ms | 20 |
| Mumbai | 664 ms | 516 ms | 20 |
Typical: half of the requests finished within this time (technical: p50 of time to first visible token). Slower: 95% of requests finished within this time (technical: p95). Statistics
- Request
- Authenticated streaming POST to api.deepseek.com/chat/completions
- Measured from
- Amsterdam · San Francisco · Montreal · Singapore · Tokyo · Mumbai
- Requests
- Daily measurements from 8 September 2026 to 21 September 2026. 6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai. · 8 September 2026 to 21 September 2026
- Model
- deepseek-v4-flash · thinking disabled · temperature 0 · max_tokens 256
- Prompt
- "Say 'ok' and nothing else." Identical across every city and every LLM API we benchmark
- Comparison figures
- ChatGPT (gpt-5.6-sol), Claude (claude-opus-5) and Grok (grok-4.6), all at low reasoning effort, from the same daily collection between 8 and 21 September 2026. Those three think briefly before the first token; this request tells DeepSeek not to
- Timings taken
- Parsed from the token stream: stream open, first visible token, completion, plus DNS, connect and TLS
- Report coverage
- 6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai.
- Daily requests in each location
- Amsterdam: 70 · San Francisco: 70 · Montreal: 70 · Singapore: 70 · Tokyo: 70 · Mumbai: 70
- Typical response time
- Median of the eligible daily requests in each included location.
- Slower response time
- Median of recent controlled full-run p95 values in each included location.
DeepSeek API
https://api.deepseek.com/chat/completions
Independent measurement by LatencyRadar. Not affiliated with DeepSeek. · How we measure (v1.1)
How fast is your API compared to DeepSeek's?
Run a free speed test from multiple cities and find out where your users are waiting. No setup, no account required.
No account required · Takes about 30 seconds.