Provider Matrix

Same model, different provider: which OpenRouter route is fastest?

OpenRouter lets you run the same model ID on different companies' infrastructure with a one-line routing change. We pinned two open-weights models to five providers each, with fallbacks off and the served provider verified on every request, and measured time to first token (TTFT) from five cities on three separate dates. That answers the question an OpenRouter user is left with once the model is chosen: which provider should serve it, from where my users are, and does the answer hold from one week to the next?

What we tested · AI responses

Time until the first visible token, with each model pinned to a named provider through OpenRouter. Every row is a separate route.

Doesn’t measure: answer quality, long responses, automatic routing, or a provider’s API called directly.

Quick take

Across three runs on 22 August, 30 August and 4 September 2026, provider choice still moved the typical first token by 2 to 3× for the same model and prompt. For DeepSeek V4 Flash, Fireworks was the fastest route in 13 of 15 city-runs, but only in San Francisco was its lead larger than its own swing between runs, so that is the only city where we name a winner. Everywhere else the answer is no clear winner yet: Fireworks led, but by less than it moved from one date to the next, and in Montreal DigitalOcean was faster on two of the three dates. GLM 5.2 cannot be ranked. On 22 August three of its five providers answered only a handful of our requests, so only DigitalOcean and Novita have three clean runs, and a ranking needs four.

DeepSeek V4 Flash

Model deepseek/deepseek-v4-flash-0731, with an identical request to every provider, timed to the first visible token, on 22 August, 30 August and 4 September 2026. Each cell is the median of those runs; hover for the swing between them.

Typical first token in each city, median across the runs.
ProviderNetherlandsAmsterdamUnited StatesSan FranciscoCanadaMontrealSingaporeSingaporeJapanTokyoIndiaMumbai
Fireworks
pinned: fireworks
Winner in 1 of 5 cities
548 ms
421 ms
616 ms
703 ms
750 ms
No run saved
DigitalOcean
pinned: digitalocean
641 ms
594 ms
440 ms
1001 ms
975 ms
No run saved
Together
pinned: together
755 ms
594 ms
634 ms
1132 ms
945 ms
No run saved
SiliconFlow
pinned: siliconflow/fp8
1157 ms
1269 ms
1144 ms
1390 ms
1396 ms
No run saved
Novita
pinned: novita/fp8
1343 ms
1385 ms
1352 ms
1590 ms
1426 ms
No run saved

Quantization per endpoint: Fireworks, DigitalOcean, Together: not listed by OpenRouter · SiliconFlow, Novita: fp8.

Verdict across 3 dated runs, by location
Test locationVerdictTypicalLead vs own swingRanked
NetherlandsAmsterdamNo clear winner yet
Fireworks led on all 3 dates, but by 93 ms while its own time moved 201 ms between dates.
+93 ms vs 201 ms5
United StatesSan FranciscoFireworks421 ms+173 ms vs 94 ms5
CanadaMontrealNo clear winner yet
DigitalOcean led most often, but the fastest provider changed between dates.
5
SingaporeSingaporeNo clear winner yet
Fireworks led on all 3 dates, but by 298 ms while its own time moved 489 ms between dates.
+298 ms vs 489 ms5
JapanTokyoNo clear winner yet
Fireworks led on all 3 dates, but by 195 ms while its own time moved 418 ms between dates.
+195 ms vs 418 ms5
IndiaMumbaiNot enough comparable providers0

GLM 5.2

Model z-ai/glm-5.2, with an identical request to every provider, timed to the first visible token, on 22 August, 30 August and 4 September 2026. Each cell is the median of those runs; hover for the swing between them.

Typical first token in each city, median across the runs.
ProviderNetherlandsAmsterdamUnited StatesSan FranciscoCanadaMontrealSingaporeSingaporeJapanTokyoIndiaMumbai
DigitalOcean
pinned: digitalocean
711 ms
544 ms
532 ms
893 ms
910 ms
No run saved
Novita
pinned: novita/fp8
1347 ms
1216 ms
1264 ms
1451 ms
1143 ms
No run saved
Fireworks
pinned: fireworks
510 ms*
629 ms*
514 ms*
878 ms*
722 ms*
No run saved
Together
pinned: together
341 ms*
444 ms*
478 ms*
694 ms*
661 ms*
No run saved
SiliconFlow
pinned: siliconflow/fp8
992 ms*
958 ms*
959 ms*
1142 ms*
1006 ms*
No run saved

Quantization per endpoint: DigitalOcean, Fireworks, Together: not listed by OpenRouter · Novita, SiliconFlow: fp8.

Verdict across 3 dated runs, by location
Test locationVerdictTypicalLead vs own swingRanked
NetherlandsAmsterdamNot enough comparable providers2
United StatesSan FranciscoNot enough comparable providers2
CanadaMontrealNot enough comparable providers2
SingaporeSingaporeNot enough comparable providers2
JapanTokyoNot enough comparable providers2
IndiaMumbaiNot enough comparable providers0

A winner above had to be fastest in every one of the dated runs and lead by more than its own run-to-run swing; a city that does not meet that says so. Even then it is what we measured on those dates, not a standing recommendation, because provider load moves by the hour. The rule we hold ourselves to before recommending a route.

Conclusions

  • DeepSeek V4 Flash: Fireworks led in 13 of 15 city-runs with a typical first token of 275–764 ms, and its lead over the next provider was 93–298 ms. Its own typical time moved 94–489 ms between dates in the same city, which is why San Francisco is the only city that passes the winner rule (lead 173 ms, swing 94 ms).
  • The fastest route has the roughest tail, and it repeated. Fireworks' slowest DeepSeek responses reached 6.7–10.9 s in at least three cities on two of the three dates; DigitalOcean's never exceeded 2.1 s on any date. These are 20 requests per city per run, so read it as a pattern seen twice, not a rate.
  • Montreal is the exception to the Fireworks story: DigitalOcean was fastest there on 30 August and 4 September (379 and 440 ms) after trailing on 22 August, so the ordering changed and no winner is named.
  • Same model, same prompt: the slowest ranked DeepSeek provider still ran 2–3× the fastest (SiliconFlow 1.1–1.5 s against Fireworks 0.3–0.8 s).
  • GLM 5.2: only DigitalOcean and Novita answered our full burst on all three dates. DigitalOcean was quicker in every city on those runs (median 0.5–0.9 s against Novita's 1.1–1.5 s), but two providers cannot make a ranking. Fireworks, Together and SiliconFlow answered 1–4 of 20 requests on 22 August, which we treat as our own account's throttling, and 20 of 20 on both later dates; one more clean run puts them in the table.
  • Together's GLM route produced the slowest responses we have seen on this page: 95th percentiles of 24.8 s (San Francisco, 22 August) and 27.5 s (Montreal, 4 September) against typical times under 0.7 s.

What this page does and doesn't claim

This page measures serving infrastructure and network path for a short completion. It does not measure model quality, throughput on long generations, or price. Providers may serve the same model ID at different quantizations, which affects both speed and output; we disclose what OpenRouter's catalog lists per endpoint, including "unknown" where it says nothing. Two providers from the original seven are no longer shown: BaseTen refused every request from our account with HTTP 429 from 28 August onward, and Cloudflare was throttled or timed out in at least one city on every run after 22 August, so neither has three comparable runs. The IP addresses, certificates and edge locations we observe belong to OpenRouter's gateway, not to any upstream provider.

Reading these tables

*
Shown but not ranked: verified in some runs but not all, or too few requests answered.
Unavailable
The pinned provider errored or streamed nothing in any run.
Unverified
We could not confirm who served the response.
No run saved
No measurement was stored for this provider.

Why a cell goes unranked, and what we verify before publishing one

What these OpenRouter results mean

The collection catalog has fourteen OpenRouter routes, in two groups of seven: GLM 5.2 (z-ai/glm-5.2) and DeepSeek V4 Flash (deepseek/deepseek-v4-flash-0731), each pinned to Fireworks, Together, BaseTen (fp8), Cloudflare, DigitalOcean, Novita (fp8) and SiliconFlow (fp8) with fallbacks off. Every request is a streaming one-word prompt with reasoning off and 256 max tokens, and the provider OpenRouter reports serving it is checked against the pin; a mismatch fails the request. The tables publish the five providers with comparable saved runs. Each row measures one route: the path to OpenRouter plus the upstream provider's time to first visible token.

What a route does not show: where the upstream provider's machines are. Our IP, TLS and edge evidence describes the path to OpenRouter's gateway; the provider name is only what OpenRouter reports. It also does not show the provider's own API called directly, or OpenRouter's automatic routing with fallbacks on, which is what most users run.

One observation from the dated runs: the same model moved by two to three times in typical first token depending on the pinned provider, and a provider's own value moved between dates by about as much as its lead over the next one. That is why the page names a winner in one city only.

If OpenRouter is slow for you

Check OpenRouter’s status page
Gateway or upstream
The route is measured to openrouter.ai. A slow route may be the gateway, the upstream, or the hop between them, and this data does not separate them.
Fallbacks
With fallbacks on (OpenRouter's default) another provider may answer. These numbers describe pinned routes with fallbacks off.
Provider status pages
Each upstream has its own status page. Check the pinned provider's as well as OpenRouter's.

How this test was run

3,000 requests · 20 in each of 5 cities, for each of 5 providers, on both models, on 3 dates · 22 August, 30 August and 4 September 2026

Technical details
Request
Streaming POST to openrouter.ai/api/v1/chat/completions, one provider pinned per request
Measured from
Amsterdam · San Francisco · Montreal · Singapore · Tokyo
Publication runs
22 August, 30 August and 4 September 2026. The first through the public API, the later two by the reference collector; identical body, pinning and verification.
Models
z-ai/glm-5.2 and deepseek/deepseek-v4-flash-0731, 5 providers each
Body
Identical everywhere: prompt "Say 'ok' and nothing else.", max_tokens 256, temperature 0, reasoning disabled
Pinning
provider.order with allow_fallbacks: false, so no silent failover
Verification
The gateway-reported served provider must match the pinned one, checked on every single request
Counting rule
A provider counts in a city when at least 15 of 20 requests returned a visible token from the pinned provider, on every date. A city names a winner only with four such providers, the same one fastest on every date, and a lead larger than that provider's own swing between dates.
Warm-up
One request per provider per city, discarded before the 20 measured ones
Concurrency
One request at a time per city; 2 s between requests on the later runs
Quantization
As listed per endpoint by OpenRouter on 22 August 2026
Timings taken
Parsed from the token stream: the headline number is the first visible token

GLM 5.2 via OpenRouter

GLM 5.2 via OpenRouter pinned to Fireworks, fallbacks off, provider verified per request: time to first visible token on a one-word prompt.

Doesn’t measure: Fireworks's own API or where its machines are: the path is measured to OpenRouter's gateway; the upstream is only what OpenRouter reports.

https://openrouter.ai/api/v1/chat/completions

GLM 5.2 via OpenRouter

GLM 5.2 via OpenRouter pinned to Together, fallbacks off, provider verified per request: time to first visible token on a one-word prompt.

Doesn’t measure: Together's own API or where its machines are: the path is measured to OpenRouter's gateway; the upstream is only what OpenRouter reports.

https://openrouter.ai/api/v1/chat/completions

GLM 5.2 via OpenRouter

GLM 5.2 via OpenRouter pinned to DigitalOcean, fallbacks off, provider verified per request: time to first visible token on a one-word prompt.

Doesn’t measure: DigitalOcean's own API or where its machines are: the path is measured to OpenRouter's gateway; the upstream is only what OpenRouter reports.

https://openrouter.ai/api/v1/chat/completions

GLM 5.2 via OpenRouter

GLM 5.2 via OpenRouter pinned to Novita (fp8), fallbacks off, provider verified per request: time to first visible token on a one-word prompt.

Doesn’t measure: Novita's own API or where its machines are: the path is measured to OpenRouter's gateway; the upstream is only what OpenRouter reports.

https://openrouter.ai/api/v1/chat/completions

GLM 5.2 via OpenRouter

GLM 5.2 via OpenRouter pinned to SiliconFlow (fp8), fallbacks off, provider verified per request: time to first visible token on a one-word prompt.

Doesn’t measure: SiliconFlow's own API or where its machines are: the path is measured to OpenRouter's gateway; the upstream is only what OpenRouter reports.

https://openrouter.ai/api/v1/chat/completions

DeepSeek V4 Flash via OpenRouter

DeepSeek V4 Flash via OpenRouter pinned to Fireworks, fallbacks off, provider verified per request: time to first visible token on a one-word prompt.

Doesn’t measure: Fireworks's own API or where its machines are: the path is measured to OpenRouter's gateway; the upstream is only what OpenRouter reports.

https://openrouter.ai/api/v1/chat/completions

DeepSeek V4 Flash via OpenRouter

DeepSeek V4 Flash via OpenRouter pinned to Together, fallbacks off, provider verified per request: time to first visible token on a one-word prompt.

Doesn’t measure: Together's own API or where its machines are: the path is measured to OpenRouter's gateway; the upstream is only what OpenRouter reports.

https://openrouter.ai/api/v1/chat/completions

DeepSeek V4 Flash via OpenRouter

DeepSeek V4 Flash via OpenRouter pinned to DigitalOcean, fallbacks off, provider verified per request: time to first visible token on a one-word prompt.

Doesn’t measure: DigitalOcean's own API or where its machines are: the path is measured to OpenRouter's gateway; the upstream is only what OpenRouter reports.

https://openrouter.ai/api/v1/chat/completions

DeepSeek V4 Flash via OpenRouter

DeepSeek V4 Flash via OpenRouter pinned to Novita (fp8), fallbacks off, provider verified per request: time to first visible token on a one-word prompt.

Doesn’t measure: Novita's own API or where its machines are: the path is measured to OpenRouter's gateway; the upstream is only what OpenRouter reports.

https://openrouter.ai/api/v1/chat/completions

DeepSeek V4 Flash via OpenRouter

DeepSeek V4 Flash via OpenRouter pinned to SiliconFlow (fp8), fallbacks off, provider verified per request: time to first visible token on a one-word prompt.

Doesn’t measure: SiliconFlow's own API or where its machines are: the path is measured to OpenRouter's gateway; the upstream is only what OpenRouter reports.

https://openrouter.ai/api/v1/chat/completions

Independent measurement by LatencyRadar. Not affiliated with OpenRouter or any provider; no credits, sponsorship or commercial relationship. · How we measure (v1.0)

Get the monthly AI latency report, by test location

One email a month: how fast the major AI providers respond, measured globally.

Check your inbox for the confirmation email. Nothing is sent until you confirm. Unsubscribe any time.

How fast is your API around the world?

Run a free latency test from multiple cities and find out where your users are waiting. No setup, no account required.

Test my endpoint

Takes about 30 seconds.

More public benchmarks →