All topicsTopic 23 / 24
Intermediate 1 minute
Rate Limiting
Rate limiting caps requests from a user, token, IP, or service over time. It protects capacity, reduces accidental overload, and makes abuse harder. Common algorithms include fixed windows, sliding windows, and token buckets. A useful limit has a clear response, such as HTTP 429, and considers legitimate bursts, distributed clients, and whether limits are shared across instances.
Key idea
Limit work before overload, and tell callers when to try again.
See it in one picture
Follow the arrowsClient
Gate
Budget
requests leftAPI
Real-world example
A public search API may allow 60 requests per minute per API key. When a client exceeds that budget, the gateway returns 429 and a retry hint instead of letting every request degrade the service.
Quick check