Benchmark report
Is the Hugging Face API slow?
Response times from 6 cities
Mixed by location
Response times ranged from 32 ms in Montreal to 247 ms in Singapore. Where you are matters.
- Fastest
- 32 ms
- Montreal
- Slowest
- 247 ms
- Singapore
- Successful requests
- 100%
- 6 locations
Response time around the world
Montreal
32 msAmsterdam
103 msSan Francisco
114 msTokyo
172 msMumbai
209 msSingapore
247 ms
Why is Singapore eight times slower than Montreal?
Every city reached a nearby CloudFront site (Los Angeles for San Francisco), and connecting plus the secure handshake took under 30 ms everywhere in the controlled run of 27 September. The rest is waiting: CloudFront does not cache this response, so each request travels on to Hugging Face and back. That wait was 24 ms from Montreal, 84 to 95 ms from San Francisco and Amsterdam, 163 ms from Tokyo, 203 ms from Mumbai and 247 ms from Singapore. It grows with distance from Montreal, which suggests the API is answered from eastern North America.
Click the location to see each stage.
Singapore256 ms
- Finding the server
- 2 ms
- Reaching the server
- 3 ms
- Setting up security
- 6 ms
- Waiting for the server: most of the time
- 247 ms
- Receiving the response
- 0 ms
- Total
- 256 ms
Finding the server is measured once per location, so it sits outside these bars. Open a city to see it.
Bars show the typical request, so the parts add up to its total. Why
Compare all locationsHide the comparison
Click a city to see its stage-by-stage breakdown.
Amsterdam103 ms
- Finding the server
- 10 ms
- Reaching the server
- 3 ms
- Setting up security
- 5 ms
- Waiting for the server: most of the time
- 95 ms
- Receiving the response
- 0 ms
- Total
- 103 ms
San Francisco112 ms
- Finding the server
- 12 ms
- Reaching the server
- 13 ms
- Setting up security
- 15 ms
- Waiting for the server: most of the time
- 84 ms
- Receiving the response
- 0 ms
- Total
- 112 ms
Montreal30 ms
- Finding the server
- 7 ms
- Reaching the server
- 2 ms
- Setting up security
- 4 ms
- Waiting for the server: most of the time
- 24 ms
- Receiving the response
- 0 ms
- Total
- 30 ms
Singapore256 ms
- Finding the server
- 2 ms
- Reaching the server
- 3 ms
- Setting up security
- 6 ms
- Waiting for the server: most of the time
- 247 ms
- Receiving the response
- 0 ms
- Total
- 256 ms
Tokyo171 ms
- Finding the server
- 6 ms
- Reaching the server
- 3 ms
- Setting up security
- 5 ms
- Waiting for the server: most of the time
- 163 ms
- Receiving the response
- 0 ms
- Total
- 171 ms
Mumbai210 ms
- Finding the server
- 2 ms
- Reaching the server
- 2 ms
- Setting up security
- 5 ms
- Waiting for the server: most of the time
- 203 ms
- Receiving the response
- 0 ms
- Total
- 210 ms
The phases come from the controlled run of 27 September 2026. Typical and slower times pool the daily requests. 6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai.
Finding the server is measured once per location, so it sits outside these bars. Open a city to see it.
Bars show the typical request, so the parts add up to its total. Why
Hugging Face still feels slow?
- Check Hugging Face’s status page
- Hub and API incidents are posted there. A slow response during one is not your network.
- This is not a model download benchmark
- We call the models listing with a limit of one, 505 bytes of JSON. Downloading weights or datasets, Inference Endpoints, Spaces and authenticated calls to private repositories are not measured.
- Distance to the origin sets the time
- Nothing is cached, so a nearby CloudFront site does not make the answer faster; being close to where the API runs does. Each city's daily median moved by less than 30 ms over the 14 days, so the gap between cities is geography, not bad days.
Response time over the last 18 days
Show the numbers
| Day | Amsterdam | San Francisco | Montreal | Singapore | Tokyo | Mumbai |
|---|---|---|---|---|---|---|
| 9 September 2026 | 104 ms | 112 ms | 55 ms | 240 ms | 171 ms | 220 ms |
| 10 September 2026 | 104 ms | 110 ms | 58 ms | 237 ms | 169 ms | 225 ms |
| 11 September 2026 | 103 ms | 113 ms | 55 ms | 240 ms | 169 ms | 214 ms |
| 12 September 2026 | 101 ms | 112 ms | 54 ms | 234 ms | 169 ms | 211 ms |
| 13 September 2026 | 102 ms | 112 ms | 53 ms | 259 ms | 179 ms | 209 ms |
| 14 September 2026 | 103 ms | 114 ms | 54 ms | 244 ms | 170 ms | 206 ms |
| 15 September 2026 | 105 ms | 109 ms | 31 ms | 245 ms | 173 ms | 223 ms |
| 16 September 2026 | 103 ms | 115 ms | 33 ms | 240 ms | 172 ms | 207 ms |
| 17 September 2026 | 101 ms | 117 ms | 30 ms | 248 ms | 170 ms | 224 ms |
| 18 September 2026 | 103 ms | 117 ms | 30 ms | 246 ms | 170 ms | 206 ms |
| 19 September 2026 | 103 ms | 112 ms | 33 ms | 237 ms | 169 ms | 205 ms |
| 20 September 2026 | 106 ms | 115 ms | 34 ms | 241 ms | 172 ms | 212 ms |
| 21 September 2026 | 103 ms | 112 ms | 32 ms | 255 ms | 174 ms | 223 ms |
| 22 September 2026 | 104 ms | 113 ms | 31 ms | 250 ms | 172 ms | 210 ms |
| 23 September 2026 | 103 ms | 112 ms | 30 ms | 245 ms | 172 ms | 209 ms |
| 24 September 2026 | 104 ms | 112 ms | 30 ms | 264 ms | 170 ms | 220 ms |
| 25 September 2026 | 103 ms | 117 ms | 33 ms | 249 ms | 173 ms | 219 ms |
| 26 September 2026 | 104 ms | 112 ms | 32 ms | 245 ms | 173 ms | 221 ms |
| 27 September 2026 | 102 ms | 115 ms | 31 ms | 261 ms | 172 ms | 218 ms |
About this measurement
GET huggingface.co/api/models · 70 requests per location over 14 daily runs · 14 September 2026 to 27 September 2026
What we tested · API response. GET huggingface.co/api/models?limit=1: 505 bytes of JSON through CloudFront, uncached. The API a client calls before it downloads weights.
How LatencyRadar measures response time →
Technical details
Doesn’t measure: Model or dataset downloads, Inference Endpoints, Spaces, or an authenticated call to a private repository.
One Hugging Face surface: GET huggingface.co/api/models?limit=1, the public Hub API that lists models, answered with 505 bytes of JSON through CloudFront without caching. Libraries and tools call endpoints like this before they download anything, so it stands for the lookup step of fetching a model and for the route from each city to where the API runs.
Not measured: model and dataset downloads, Inference Endpoints, Spaces, or authenticated calls to private repositories.
Between 14 and 27 September 2026 the typical time was 32 ms in Montreal, 103 ms in Amsterdam, 114 ms in San Francisco, 172 ms in Tokyo, 209 ms in Mumbai and 247 ms in Singapore. Montreal was the fastest city and Singapore the slowest on all 14 days, and the slower times sit close to the typical ones everywhere (at most 268 ms, in Singapore). The API is steady; where you call it from is what changes the number.
| Test location | Typical response time | Slower response time | Requests |
|---|---|---|---|
| Amsterdam | 103 ms | 117 ms | 20 |
| San Francisco | 114 ms | 129 ms | 20 |
| Montreal | 32 ms | 54 ms | 20 |
| Singapore | 247 ms | 268 ms | 20 |
| Tokyo | 172 ms | 177 ms | 20 |
| Mumbai | 209 ms | 226 ms | 20 |
Typical: half of the requests finished within this time (technical: p50). Slower: 95% of requests finished within this time (technical: p95). Statistics
- Request
- GET huggingface.co/api/models
- Measured from
- Amsterdam · San Francisco · Montreal · Singapore · Tokyo · Mumbai
- Requests
- Daily measurements from 14 September 2026 to 27 September 2026. 6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai. · 14 September 2026 to 27 September 2026
- Timings taken
- DNS, connect, TLS, waiting for the server, download
- Report coverage
- 6 of 6 test locations included: Amsterdam, San Francisco, Montreal, Singapore, Tokyo, Mumbai.
- Daily requests in each location
- Amsterdam: 70 · San Francisco: 70 · Montreal: 70 · Singapore: 70 · Tokyo: 70 · Mumbai: 70
- Typical response time
- Median of the eligible daily requests in each included location.
- Slower response time
- The 95th percentile of the same eligible daily requests in each included location.
Hugging Face API
https://huggingface.co/api/models?limit=1
Independent measurement by LatencyRadar. Not affiliated with Hugging Face API. · How we measure (v1.2)
How fast is your API around the world?
Run a free speed test from multiple cities and find out where your users are waiting. No setup, no account required.
No account required · Takes about 30 seconds.