1. What Is Rate Limiting and Why Does It Exist?
API rate limiting is a technique that controls the number of requests a client can make to an API within a given time window. If you have ever received an HTTP 429 Too Many Requests response, you have encountered rate limiting firsthand.
Rate limiting exists for several practical reasons:
- Server protection: Without limits, a single client could overwhelm a server with millions of requests, degrading performance for everyone.
- Fair usage: Rate limits ensure that no single consumer monopolizes shared resources.
- Abuse prevention: Limits help defend against brute-force attacks, credential stuffing, and scraping operations that violate terms of service.
- Cost management: Cloud infrastructure costs scale with traffic. Rate limiting keeps usage predictable and within budget.
- Service stability: Downstream dependencies (databases, third-party APIs) also have capacity limits. Rate limiting acts as a pressure valve.
Rate limiting is not about restricting developers -- it is about making the API reliable for all developers. An unprotected API eventually fails for everyone.
2. Common Rate Limiting Strategies
There are four widely used algorithms for rate limiting. Each makes different tradeoffs between simplicity, fairness, and burst tolerance.
Token Bucket
The token bucket algorithm is the most common approach. Imagine a bucket that holds a fixed number of tokens. Tokens are added at a constant rate (e.g., 10 per second). Each request consumes one token. If the bucket is empty, the request is rejected.
Key property: Token bucket allows bursts. If a client has been idle, the bucket fills up, and they can make a burst of requests up to the bucket capacity. This makes it well-suited for APIs where occasional traffic spikes are normal.
class TokenBucket:
capacity = 100 # max tokens
refill_rate = 10 # tokens per second
tokens = 100 # current tokens
last_refill = now()
function allow_request():
elapsed = now() - last_refill
tokens = min(capacity, tokens + elapsed * refill_rate)
last_refill = now()
if tokens >= 1:
tokens -= 1
return true
return false
Leaky Bucket
The leaky bucket processes requests at a fixed rate, regardless of how many arrive. Incoming requests are queued, and the queue is processed at a constant rate. If the queue is full, new requests are dropped.
Key property: The output rate is perfectly smooth. This is ideal when you need to protect a downstream service that cannot handle bursts -- for example, a payment processor with strict throughput limits.
Fixed Window Counter
The simplest approach: divide time into fixed windows (e.g., every minute) and count requests per window. If the count exceeds the limit, reject the request until the next window starts.
Key weakness: The boundary problem. A client can send the maximum number of requests at the end of one window and the maximum again at the start of the next, effectively doubling their rate for a brief period.
Sliding Window Log
Sliding window keeps a timestamp log of every request. To check the rate, count the timestamps within the last N seconds. This eliminates the boundary problem of fixed windows but requires more memory since you store every request timestamp.
Sliding Window Counter (Hybrid)
A practical compromise: combine the previous window's count (weighted by how much of it overlaps with the current window) with the current window's count. This approximates the accuracy of a sliding log with the memory efficiency of a fixed counter.
| Algorithm | Burst Tolerance | Memory | Accuracy | Best For |
|---|---|---|---|---|
| Token Bucket | High | Low | Good | General-purpose APIs |
| Leaky Bucket | None | Low | Good | Smooth output required |
| Fixed Window | Boundary issue | Very Low | Moderate | Simple internal services |
| Sliding Window Log | None | High | Exact | Strict compliance needs |
| Sliding Window Counter | Minimal | Low | Good | Production APIs at scale |
3. HTTP Rate Limit Headers Explained
Well-designed APIs communicate rate limit status through HTTP response headers. As a consumer, these headers tell you exactly where you stand without having to guess.
Standard Headers
HTTP/1.1 200 OK
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 742
X-RateLimit-Reset: 1708520400
Content-Type: application/json
X-RateLimit-Limit-- The maximum number of requests allowed in the current window.X-RateLimit-Remaining-- How many requests you have left before hitting the limit.X-RateLimit-Reset-- Unix timestamp (seconds) indicating when the window resets.Retry-After-- Sent with429responses. Tells you how many seconds to wait, or provides an HTTP date for when to retry.
The IETF is working on standardized rate limit headers via RFC 9110 and draft-ietf-httpapi-ratelimit-headers. The new names are RateLimit-Limit, RateLimit-Remaining, and RateLimit-Reset (without the X- prefix). If you are building a new API, consider supporting both formats during the transition period.
Reading Headers in JavaScript
const response = await fetch('https://api.example.com/data');
const limit = response.headers.get('X-RateLimit-Limit');
const remaining = response.headers.get('X-RateLimit-Remaining');
const resetTimestamp = response.headers.get('X-RateLimit-Reset');
console.log(`${remaining}/${limit} requests remaining`);
console.log(`Resets at: ${new Date(resetTimestamp * 1000).toISOString()}`);
if (parseInt(remaining) < 10) {
console.warn('Approaching rate limit -- consider slowing down');
}
You can test API responses and inspect these headers directly with the QTool API Tester -- no setup required, works right in your browser.
4. Handling 429 Too Many Requests
The 429 Too Many Requests status code means you have exceeded the allowed rate. How you respond to it determines whether your application degrades gracefully or crashes in a loop.
The Naive Approach (Do Not Do This)
// DO NOT DO THIS -- immediate retry without delay
async function fetchData(url) {
const response = await fetch(url);
if (response.status === 429) {
return fetchData(url); // infinite tight loop!
}
return response.json();
}
This creates a tight retry loop that hammers the server even harder, making the problem worse for everyone.
The Right Approach: Respect Retry-After
async function fetchWithRetry(url, maxRetries = 3) {
for (let attempt = 0; attempt <= maxRetries; attempt++) {
const response = await fetch(url);
if (response.status !== 429) {
return response;
}
if (attempt === maxRetries) {
throw new Error(`Rate limited after ${maxRetries} retries`);
}
// Check Retry-After header
const retryAfter = response.headers.get('Retry-After');
let waitMs;
if (retryAfter) {
// Retry-After can be seconds or an HTTP date
const seconds = parseInt(retryAfter);
if (!isNaN(seconds)) {
waitMs = seconds * 1000;
} else {
waitMs = new Date(retryAfter) - Date.now();
}
} else {
// Fallback: exponential backoff
waitMs = Math.pow(2, attempt) * 1000;
}
console.log(`Rate limited. Waiting ${waitMs}ms before retry...`);
await new Promise(resolve => setTimeout(resolve, waitMs));
}
}
5. Exponential Backoff with Jitter
Exponential backoff increases the wait time between retries: 1 second, 2 seconds, 4 seconds, 8 seconds, and so on. But pure exponential backoff has a critical problem: if 1,000 clients all get rate limited at the same time, they will all retry at the same intervals, creating synchronized spikes.
Jitter solves this by adding a random component to each delay, spreading retries across time and preventing the "thundering herd" effect.
Full Jitter (Recommended)
function getBackoffDelay(attempt, baseDelay = 1000, maxDelay = 30000) {
// Full jitter: random between 0 and exponential cap
const exponentialDelay = Math.min(maxDelay, baseDelay * Math.pow(2, attempt));
return Math.random() * exponentialDelay;
}
// Equal jitter: half exponential + half random
function getEqualJitterDelay(attempt, baseDelay = 1000, maxDelay = 30000) {
const exponentialDelay = Math.min(maxDelay, baseDelay * Math.pow(2, attempt));
const halfDelay = exponentialDelay / 2;
return halfDelay + Math.random() * halfDelay;
}
// Decorrelated jitter (AWS recommendation)
let previousDelay = 1000;
function getDecorrelatedDelay(baseDelay = 1000, maxDelay = 30000) {
const delay = Math.min(maxDelay, Math.random() * previousDelay * 3);
previousDelay = Math.max(baseDelay, delay);
return delay;
}
AWS published an excellent analysis showing that full jitter produces the best results in most scenarios. It minimizes total completion time across all clients while keeping server load smooth. Decorrelated jitter is a close second and simpler to implement in stateful contexts.
Production-Ready Retry Wrapper
class RateLimitedClient {
constructor(baseUrl, options = {}) {
this.baseUrl = baseUrl;
this.maxRetries = options.maxRetries || 5;
this.baseDelay = options.baseDelay || 1000;
this.maxDelay = options.maxDelay || 60000;
}
async request(path, options = {}) {
const url = `${this.baseUrl}${path}`;
for (let attempt = 0; attempt <= this.maxRetries; attempt++) {
const response = await fetch(url, options);
// Success or client error (not rate limited)
if (response.status !== 429 && response.status !== 503) {
return response;
}
if (attempt === this.maxRetries) {
throw new Error(
`Request to ${path} failed after ${this.maxRetries} retries`
);
}
const waitMs = this.calculateDelay(response, attempt);
console.warn(
`[Attempt ${attempt + 1}/${this.maxRetries}] ` +
`Rate limited. Waiting ${Math.round(waitMs)}ms...`
);
await this.sleep(waitMs);
}
}
calculateDelay(response, attempt) {
// Prefer server-provided Retry-After
const retryAfter = response.headers.get('Retry-After');
if (retryAfter) {
const seconds = parseInt(retryAfter);
if (!isNaN(seconds)) return seconds * 1000;
const date = new Date(retryAfter);
if (!isNaN(date)) return Math.max(0, date - Date.now());
}
// Full jitter exponential backoff
const cap = Math.min(
this.maxDelay,
this.baseDelay * Math.pow(2, attempt)
);
return Math.random() * cap;
}
sleep(ms) {
return new Promise(resolve => setTimeout(resolve, ms));
}
}
// Usage
const api = new RateLimitedClient('https://api.example.com');
const data = await api.request('/v1/users?page=1');
const json = await data.json();
6. Implementing Rate Limiting Server-Side
If you are building an API, you need to implement rate limiting on the server side. The most common approach for distributed systems is to use Redis as a shared counter.
Express.js with Redis (Sliding Window)
const Redis = require('ioredis');
const redis = new Redis();
function rateLimiter(options = {}) {
const {
windowMs = 60 * 1000, // 1 minute
maxRequests = 100,
keyGenerator = (req) => req.ip,
message = 'Too many requests, please try again later.'
} = options;
return async (req, res, next) => {
const key = `ratelimit:${keyGenerator(req)}`;
const now = Date.now();
const windowStart = now - windowMs;
// Remove expired entries, add current, count total
const pipeline = redis.pipeline();
pipeline.zremrangebyscore(key, 0, windowStart);
pipeline.zadd(key, now, `${now}-${Math.random()}`);
pipeline.zcard(key);
pipeline.pexpire(key, windowMs);
const results = await pipeline.exec();
const requestCount = results[2][1];
// Set rate limit headers
res.set('X-RateLimit-Limit', maxRequests);
res.set('X-RateLimit-Remaining',
Math.max(0, maxRequests - requestCount));
res.set('X-RateLimit-Reset',
Math.ceil((now + windowMs) / 1000));
if (requestCount > maxRequests) {
res.set('Retry-After',
Math.ceil(windowMs / 1000));
return res.status(429).json({
error: {
code: 'RATE_LIMIT_EXCEEDED',
message,
retryAfter: Math.ceil(windowMs / 1000)
}
});
}
next();
};
}
// Apply to routes
app.use('/api/', rateLimiter({
windowMs: 60 * 1000,
maxRequests: 100
}));
// Stricter limit for auth endpoints
app.use('/api/auth/', rateLimiter({
windowMs: 15 * 60 * 1000,
maxRequests: 10,
message: 'Too many login attempts.'
}));
You can validate your rate limit header responses and JSON error bodies using the QTool JSON Editor to format and inspect the output.
7. Real-World API Rate Limits
Understanding how major APIs set their limits gives you a sense of industry norms and helps you design your own limits appropriately.
| API | Rate Limit | Window | Notes |
|---|---|---|---|
| GitHub REST | 5,000 requests | 1 hour | Authenticated. 60/hour unauthenticated. |
| Twitter/X v2 | 300-900 requests | 15 minutes | Varies by endpoint and tier. |
| Stripe | 100 requests | 1 second | Per-second with burst allowance. |
| OpenAI | Varies by tier | Per minute | Both RPM and TPM (tokens per minute). |
| Shopify | 40 requests | Per second | Leaky bucket at 2 req/second restore. |
Notice the variety: GitHub uses a generous hourly window, Stripe uses per-second limits with bursting, and OpenAI limits both request count and token usage simultaneously. Your choice should reflect your API's usage patterns and infrastructure constraints.
8. Best Practices Checklist
Whether you are consuming or building rate-limited APIs, follow these guidelines:
As an API Consumer
- Always read rate limit headers before making the next request. Proactive throttling beats reactive retrying.
- Implement exponential backoff with jitter for 429 handling. Never use fixed-delay retries.
- Cache responses when possible. The fastest request is the one you do not make.
- Use bulk endpoints instead of making many individual requests.
- Set a maximum retry count (3-5 attempts). Infinite retries waste resources and mask bugs.
- Log rate limit events to identify patterns and optimize your request patterns.
As an API Provider
- Always return rate limit headers on every response, not just 429s.
- Include
Retry-Afteron 429 responses. Do not make clients guess. - Use different limits for different endpoints. Auth endpoints need stricter limits than read-only endpoints.
- Consider tiered limits based on authentication level or subscription plan.
- Document your rate limits clearly in your API documentation.
- Use Redis or a dedicated rate limiting service for distributed environments. In-memory counters do not work across multiple server instances.
- Return a descriptive JSON error body with the 429 status, not just an empty response. Format and validate it with the JSON Editor.
The best rate limiting strategy is invisible to well-behaved clients. They never hit the limit because the headers and documentation guide them to stay within bounds.
9. Frequently Asked Questions
What is API rate limiting and why does it exist?
API rate limiting is a technique that controls the number of requests a client can make to an API within a given time window. It exists to protect servers from being overwhelmed, ensure fair usage across all consumers, prevent abuse and denial-of-service attacks, and manage infrastructure costs by keeping traffic within capacity.
What is the difference between token bucket and leaky bucket algorithms?
The token bucket algorithm adds tokens at a fixed rate and allows bursts up to the bucket capacity, making it flexible for bursty traffic. The leaky bucket algorithm processes requests at a constant rate regardless of input, smoothing out traffic spikes. Token bucket is better when you want to allow occasional bursts, while leaky bucket is better when you need a perfectly steady output rate.
How should you handle HTTP 429 Too Many Requests errors?
When you receive a 429 error, first check the Retry-After header for the server's suggested wait time. If no Retry-After header is present, implement exponential backoff with jitter: start with a base delay (e.g., 1 second), double it on each retry, add random jitter to prevent thundering herd problems, and set a maximum retry count (typically 3-5 attempts) to avoid infinite loops.
What HTTP headers are used for API rate limiting?
The most common rate limiting headers are X-RateLimit-Limit (maximum requests allowed per window), X-RateLimit-Remaining (requests left in the current window), X-RateLimit-Reset (Unix timestamp when the window resets), and Retry-After (seconds to wait before retrying, sent with 429 responses). The IETF is standardizing these under RateLimit-Limit, RateLimit-Remaining, and RateLimit-Reset.
What is exponential backoff with jitter and why is it important?
Exponential backoff with jitter is a retry strategy where wait times increase exponentially (1s, 2s, 4s, 8s) with a random component added to each delay. The jitter prevents the thundering herd problem, where many clients that were rate-limited at the same time all retry simultaneously, causing another spike. Without jitter, synchronized retries can overwhelm a server repeatedly.