Benchmark report

Hugging Face API thumbnail

Is the Hugging Face API slow?

Response times from 6 cities

Mixed by location

Response times ranged from 32 ms in Montreal to 247 ms in Singapore. Where you are matters.

Fastest
32 ms
Montreal
Slowest
247 ms
Singapore
Successful requests
100%
6 locations

Response time around the world

Amsterdam 103ms typicalSan Francisco 114ms typicalMontreal 32ms typicalSingapore 247ms typicalTokyo 172ms typicalMumbai 209ms typical
  • CanadaMontreal
    32 ms
  • NetherlandsAmsterdam
    103 ms
  • United StatesSan Francisco
    114 ms
  • JapanTokyo
    172 ms
  • IndiaMumbai
    209 ms
  • SingaporeSingapore
    247 ms

Why is Singapore eight times slower than Montreal?

Every city reached a nearby CloudFront site (Los Angeles for San Francisco), and connecting plus the secure handshake took under 30 ms everywhere in the controlled run of 27 September. The rest is waiting: CloudFront does not cache this response, so each request travels on to Hugging Face and back. That wait was 24 ms from Montreal, 84 to 95 ms from San Francisco and Amsterdam, 163 ms from Tokyo, 203 ms from Mumbai and 247 ms from Singapore. It grows with distance from Montreal, which suggests the API is answered from eastern North America.

Click the location to see each stage.

ConnectTLSServer waitDownload
SingaporeSingapore
256 ms
Finding the server
2 ms
Reaching the server
3 ms
Setting up security
6 ms
Waiting for the server: most of the time
247 ms
Receiving the response
0 ms
Total
256 ms

Finding the server is measured once per location, so it sits outside these bars. Open a city to see it.

Bars show the typical request, so the parts add up to its total. Why

Compare all locations

Click a city to see its stage-by-stage breakdown.

ConnectTLSServer waitDownload
NetherlandsAmsterdam
103 ms
Finding the server
10 ms
Reaching the server
3 ms
Setting up security
5 ms
Waiting for the server: most of the time
95 ms
Receiving the response
0 ms
Total
103 ms
United StatesSan Francisco
112 ms
Finding the server
12 ms
Reaching the server
13 ms
Setting up security
15 ms
Waiting for the server: most of the time
84 ms
Receiving the response
0 ms
Total
112 ms
CanadaMontreal
30 ms
Finding the server
7 ms
Reaching the server
2 ms
Setting up security
4 ms
Waiting for the server: most of the time
24 ms
Receiving the response
0 ms
Total
30 ms
SingaporeSingapore
256 ms
Finding the server
2 ms
Reaching the server
3 ms
Setting up security
6 ms
Waiting for the server: most of the time
247 ms
Receiving the response
0 ms
Total
256 ms
JapanTokyo
171 ms
Finding the server
6 ms
Reaching the server
3 ms
Setting up security
5 ms
Waiting for the server: most of the time
163 ms
Receiving the response
0 ms
Total
171 ms
IndiaMumbai
210 ms
Finding the server
2 ms
Reaching the server
2 ms
Setting up security
5 ms
Waiting for the server: most of the time
203 ms
Receiving the response
0 ms
Total
210 ms

The phases come from the controlled run of 27 September 2026. Typical and slower times pool the daily requests. 6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai.

Finding the server is measured once per location, so it sits outside these bars. Open a city to see it.

Bars show the typical request, so the parts add up to its total. Why

Hugging Face still feels slow?

Check Hugging Face’s status page
Hub and API incidents are posted there. A slow response during one is not your network.
This is not a model download benchmark
We call the models listing with a limit of one, 505 bytes of JSON. Downloading weights or datasets, Inference Endpoints, Spaces and authenticated calls to private repositories are not measured.
Distance to the origin sets the time
Nothing is cached, so a nearby CloudFront site does not make the answer faster; being close to where the API runs does. Each city's daily median moved by less than 30 ms over the 14 days, so the gap between cities is geography, not bad days.

Response time over the last 18 days

Daily median of 5 requests per location, 9 September 2026 to 27 September 2026.
0250500 ms9 Sept27 Sept
Show the numbers
DayAmsterdamSan FranciscoMontrealSingaporeTokyoMumbai
9 September 2026104 ms112 ms55 ms240 ms171 ms220 ms
10 September 2026104 ms110 ms58 ms237 ms169 ms225 ms
11 September 2026103 ms113 ms55 ms240 ms169 ms214 ms
12 September 2026101 ms112 ms54 ms234 ms169 ms211 ms
13 September 2026102 ms112 ms53 ms259 ms179 ms209 ms
14 September 2026103 ms114 ms54 ms244 ms170 ms206 ms
15 September 2026105 ms109 ms31 ms245 ms173 ms223 ms
16 September 2026103 ms115 ms33 ms240 ms172 ms207 ms
17 September 2026101 ms117 ms30 ms248 ms170 ms224 ms
18 September 2026103 ms117 ms30 ms246 ms170 ms206 ms
19 September 2026103 ms112 ms33 ms237 ms169 ms205 ms
20 September 2026106 ms115 ms34 ms241 ms172 ms212 ms
21 September 2026103 ms112 ms32 ms255 ms174 ms223 ms
22 September 2026104 ms113 ms31 ms250 ms172 ms210 ms
23 September 2026103 ms112 ms30 ms245 ms172 ms209 ms
24 September 2026104 ms112 ms30 ms264 ms170 ms220 ms
25 September 2026103 ms117 ms33 ms249 ms173 ms219 ms
26 September 2026104 ms112 ms32 ms245 ms173 ms221 ms
27 September 2026102 ms115 ms31 ms261 ms172 ms218 ms

About this measurement

GET huggingface.co/api/models · 70 requests per location over 14 daily runs · 14 September 2026 to 27 September 2026

What we tested · API response. GET huggingface.co/api/models?limit=1: 505 bytes of JSON through CloudFront, uncached. The API a client calls before it downloads weights.

How LatencyRadar measures response time →

Technical details

Doesn’t measure: Model or dataset downloads, Inference Endpoints, Spaces, or an authenticated call to a private repository.

One Hugging Face surface: GET huggingface.co/api/models?limit=1, the public Hub API that lists models, answered with 505 bytes of JSON through CloudFront without caching. Libraries and tools call endpoints like this before they download anything, so it stands for the lookup step of fetching a model and for the route from each city to where the API runs.

Not measured: model and dataset downloads, Inference Endpoints, Spaces, or authenticated calls to private repositories.

Between 14 and 27 September 2026 the typical time was 32 ms in Montreal, 103 ms in Amsterdam, 114 ms in San Francisco, 172 ms in Tokyo, 209 ms in Mumbai and 247 ms in Singapore. Montreal was the fastest city and Singapore the slowest on all 14 days, and the slower times sit close to the typical ones everywhere (at most 268 ms, in Singapore). The API is steady; where you call it from is what changes the number.

Test locationTypical response timeSlower response timeRequests
Amsterdam103 ms117 ms20
San Francisco114 ms129 ms20
Montreal32 ms54 ms20
Singapore247 ms268 ms20
Tokyo172 ms177 ms20
Mumbai209 ms226 ms20

Typical: half of the requests finished within this time (technical: p50). Slower: 95% of requests finished within this time (technical: p95). Statistics

Request
GET huggingface.co/api/models
Measured from
Amsterdam · San Francisco · Montreal · Singapore · Tokyo · Mumbai
Requests
Daily measurements from 14 September 2026 to 27 September 2026. 6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai. · 14 September 2026 to 27 September 2026
Timings taken
DNS, connect, TLS, waiting for the server, download
Report coverage
6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai.
Daily requests in each location
Amsterdam: 70 · San Francisco: 70 · Montreal: 70 · Singapore: 70 · Tokyo: 70 · Mumbai: 70
Typical response time
Median of the eligible daily requests in each included location.
Slower response time
The 95th percentile of the same eligible daily requests in each included location.

Hugging Face API

https://huggingface.co/api/models?limit=1

Independent measurement by LatencyRadar. Not affiliated with Hugging Face API. · How we measure (v1.2)

How fast is your API around the world?

Run a free speed test from multiple cities and find out where your users are waiting. No setup, no account required.

Test my API

No account required · Takes about 30 seconds.

More benchmarks

All benchmarks →