Ram
All topicsTopic 23 / 24
Intermediate 1 minute

Rate Limiting

Rate limiting caps requests from a user, token, IP, or service over time. It protects capacity, reduces accidental overload, and makes abuse harder. Common algorithms include fixed windows, sliding windows, and token buckets. A useful limit has a clear response, such as HTTP 429, and considers legitimate bursts, distributed clients, and whether limits are shared across instances.

Key idea

Limit work before overload, and tell callers when to try again.

See it in one picture

Follow the arrows

Real-world example

A public search API may allow 60 requests per minute per API key. When a client exceeds that budget, the gateway returns 429 and a retry hint instead of letting every request degrade the service.

Quick check

Can you spot it?

0/2 answered
Question 1

1. What does rate limiting control?

Question 2

2. What status commonly signals a limit?