Design a Rate Limiter: A System Design Interview Walkthrough
"Design a rate limiter" shows up constantly, both as its own question and as a follow-up tacked onto something bigger ("now add rate limiting to the API you just designed"). It's popular with interviewers because it looks small but has a surprising amount of real depth: several genuinely different algorithms, a distributed-systems problem hiding in plain sight once you have more than one server, and concrete product decisions about what happens to a request that gets blocked.
1. Scope it before you pick an algorithm
"Rate limiter" is underspecified until you nail down what's being limited and why. Ask, and state your assumptions:
- What's the limit keyed on — per user, per IP, per API key, per endpoint? Different keys often need different limits running at the same time.
- Is this protecting the backend from overload, enforcing a paid-tier quota, or preventing abuse (e.g. login brute-forcing)? The reason changes how strict and how precise the limiter needs to be.
- What should happen on a rejected request — hard reject with 429, silent drop, or queue and delay? This is a product decision as much as an engineering one.
- Single server or a fleet behind a load balancer? This is the question that decides whether the interesting part of the interview is the algorithm or the distributed coordination.
2. The algorithms, and the trade-off each one makes
This is the part interviewers are usually most interested in — not whether you've memorized the names, but whether you can explain what each one gets wrong at the edges.
- Fixed window counter — count requests in the current time bucket (e.g. this calendar minute), reset at the boundary. Simple and cheap, but bursty at the edges: a client can send its full limit at 11:59:59 and again at 12:00:00, doubling its effective rate for a moment.
- Sliding window log — store a timestamp per request and count how many fall in the trailing window. Accurate, no boundary burst, but memory grows with request volume, which is expensive at high QPS.
- Sliding window counter — approximate the sliding log by weighting the previous and current fixed windows. Fixes most of the boundary-burst problem at a fraction of the memory of a full log — the usual production compromise.
- Token bucket — a bucket refills with tokens at a fixed rate up to a cap; each request consumes a token, and an empty bucket rejects. Naturally allows short bursts up to the bucket size while still enforcing a long-run average rate, which is why it's the default choice for most public APIs (it's what Stripe and AWS use).
- Leaky bucket — requests join a queue that drains at a fixed rate; a full queue rejects. Smooths bursts into a constant output rate, at the cost of added latency for queued requests, which token bucket doesn't impose.
Token bucket is the safe default to lead with in an interview — it's simple to reason about, allows reasonable bursts, and is genuinely what most production systems use — but naming the fixed-window boundary problem and why sliding window counter fixes it is what signals you actually understand the trade-off, not just the vocabulary.
3. The real problem: making it work across many servers
A rate limiter that lives in each app server's memory is trivial and also wrong the moment you have more than one server behind a load balancer — a client can get the full limit on every server independently, multiplying their effective quota by your fleet size. This is usually the pivot point of the interview.
- Centralize the counter in a shared store — Redis is the standard choice, since `INCR` with a TTL gives you an atomic fixed-window counter in one round trip, and it's fast enough to sit on the hot path of every request.
- Watch for race conditions on non-atomic patterns: reading a counter, checking it in application code, then writing it back is a check-then-act race under concurrent requests. Prefer an atomic primitive (`INCR`, or a small Lua script for token bucket logic) over read-modify-write.
- Decide where enforcement actually happens: at the API gateway / edge (one place, protects every backend service, but coarser-grained) versus in each service (fine-grained, but now every service needs the same correct logic and the same shared store).
- At very high scale, even Redis becomes a bottleneck or single point of failure — mention sharding the limiter by key, or accepting slightly relaxed accuracy (each node keeps a local approximation and syncs periodically) as the next trade-off once a single shared store stops being enough.
4. What to do with a rejected request
Don't skip this — it's a small section but it's easy, concrete signal that you're thinking about the caller's experience, not just the mechanism:
- Return 429 Too Many Requests, not a generic error, so well-behaved clients can distinguish "back off" from "broken."
- Set a Retry-After header so clients know when to try again instead of retrying immediately and making the overload worse.
- Expose remaining-quota headers (X-RateLimit-Remaining, X-RateLimit-Reset) so clients can self-throttle before they ever hit the limit.
5. What a strong finish looks like
In the last few minutes, be ready to reason about the limiter's own failure modes, unprompted:
- What happens if the shared Redis instance goes down — does every request get rejected (fail closed, safe but hurts availability) or allowed through uncounted (fail open, available but momentarily unprotected)? Most production systems fail open for a rate limiter and fail closed for auth, and being able to say why is good signal.
- How do you support different limits for different tiers (free vs. paid) on the same endpoint without duplicating the whole system?
- How would you rate-limit something bursty-by-nature, like a batch upload, without the fixed-average logic punishing normal usage?
Rate limiter isn't on BuildTheSystem's live question list yet — it's on the roadmap alongside a handful of other infrastructure-flavored questions — but the same live, typed-canvas format applies to the questions that are up now: talk through your design out loud while an AI interviewer watches what you actually draw and pushes on the choices you make, instead of grading a write-up after the fact.
Rehearse this out loud, not just on paper
BuildTheSystem runs a live mock interview on a typed-component whiteboard — the interviewer listens, waits while you draw, and reacts the moment you connect something questionable.
Start a mock interviewRead next
7 System Design Interview Mistakes That Get Candidates Rejected
The recurring mistakes that sink otherwise-solid system design interview performances, from skipping requirements to silent whiteboarding.
How to Design a URL Shortener (Bitly): The Complete Guide
Learn how to design a URL shortener step by step — short code generation, redirect-path caching, and data storage trade-offs — then rehearse the exact same question out loud in a live mock interview.