Back to the blog
8 min read

Design a Rate Limiter: A System Design Interview Walkthrough

walkthroughrate limiter

"Design a rate limiter" shows up constantly, both as its own question and as a follow-up tacked onto something bigger ("now add rate limiting to the API you just designed"). It's popular with interviewers because it looks small but has a surprising amount of real depth: several genuinely different algorithms, a distributed-systems problem hiding in plain sight once you have more than one server, and concrete product decisions about what happens to a request that gets blocked.

1. Scope it before you pick an algorithm

"Rate limiter" is underspecified until you nail down what's being limited and why. Ask, and state your assumptions:

  • What's the limit keyed on — per user, per IP, per API key, per endpoint? Different keys often need different limits running at the same time.
  • Is this protecting the backend from overload, enforcing a paid-tier quota, or preventing abuse (e.g. login brute-forcing)? The reason changes how strict and how precise the limiter needs to be.
  • What should happen on a rejected request — hard reject with 429, silent drop, or queue and delay? This is a product decision as much as an engineering one.
  • Single server or a fleet behind a load balancer? This is the question that decides whether the interesting part of the interview is the algorithm or the distributed coordination.

2. The algorithms, and the trade-off each one makes

This is the part interviewers are usually most interested in — not whether you've memorized the names, but whether you can explain what each one gets wrong at the edges.

  • Fixed window counter — count requests in the current time bucket (e.g. this calendar minute), reset at the boundary. Simple and cheap, but bursty at the edges: a client can send its full limit at 11:59:59 and again at 12:00:00, doubling its effective rate for a moment.
  • Sliding window log — store a timestamp per request and count how many fall in the trailing window. Accurate, no boundary burst, but memory grows with request volume, which is expensive at high QPS.
  • Sliding window counter — approximate the sliding log by weighting the previous and current fixed windows. Fixes most of the boundary-burst problem at a fraction of the memory of a full log — the usual production compromise.
  • Token bucket — a bucket refills with tokens at a fixed rate up to a cap; each request consumes a token, and an empty bucket rejects. Naturally allows short bursts up to the bucket size while still enforcing a long-run average rate, which is why it's the default choice for most public APIs (it's what Stripe and AWS use).
  • Leaky bucket — requests join a queue that drains at a fixed rate; a full queue rejects. Smooths bursts into a constant output rate, at the cost of added latency for queued requests, which token bucket doesn't impose.

Token bucket is the safe default to lead with in an interview — it's simple to reason about, allows reasonable bursts, and is genuinely what most production systems use — but naming the fixed-window boundary problem and why sliding window counter fixes it is what signals you actually understand the trade-off, not just the vocabulary.

3. The real problem: making it work across many servers

A rate limiter that lives in each app server's memory is trivial and also wrong the moment you have more than one server behind a load balancer — a client can get the full limit on every server independently, multiplying their effective quota by your fleet size. This is usually the pivot point of the interview.

  • Centralize the counter in a shared store — Redis is the standard choice, since `INCR` with a TTL gives you an atomic fixed-window counter in one round trip, and it's fast enough to sit on the hot path of every request.
  • Watch for race conditions on non-atomic patterns: reading a counter, checking it in application code, then writing it back is a check-then-act race under concurrent requests. Prefer an atomic primitive (`INCR`, or a small Lua script for token bucket logic) over read-modify-write.
  • Decide where enforcement actually happens: at the API gateway / edge (one place, protects every backend service, but coarser-grained) versus in each service (fine-grained, but now every service needs the same correct logic and the same shared store).
  • At very high scale, even Redis becomes a bottleneck or single point of failure — mention sharding the limiter by key, or accepting slightly relaxed accuracy (each node keeps a local approximation and syncs periodically) as the next trade-off once a single shared store stops being enough.

4. What to do with a rejected request

Don't skip this — it's a small section but it's easy, concrete signal that you're thinking about the caller's experience, not just the mechanism:

  • Return 429 Too Many Requests, not a generic error, so well-behaved clients can distinguish "back off" from "broken."
  • Set a Retry-After header so clients know when to try again instead of retrying immediately and making the overload worse.
  • Expose remaining-quota headers (X-RateLimit-Remaining, X-RateLimit-Reset) so clients can self-throttle before they ever hit the limit.

5. What a strong finish looks like

In the last few minutes, be ready to reason about the limiter's own failure modes, unprompted:

  • What happens if the shared Redis instance goes down — does every request get rejected (fail closed, safe but hurts availability) or allowed through uncounted (fail open, available but momentarily unprotected)? Most production systems fail open for a rate limiter and fail closed for auth, and being able to say why is good signal.
  • How do you support different limits for different tiers (free vs. paid) on the same endpoint without duplicating the whole system?
  • How would you rate-limit something bursty-by-nature, like a batch upload, without the fixed-average logic punishing normal usage?

Rate limiter isn't on BuildTheSystem's live question list yet — it's on the roadmap alongside a handful of other infrastructure-flavored questions — but the same live, typed-canvas format applies to the questions that are up now: talk through your design out loud while an AI interviewer watches what you actually draw and pushes on the choices you make, instead of grading a write-up after the fact.

Rehearse this out loud, not just on paper

BuildTheSystem runs a live mock interview on a typed-component whiteboard — the interviewer listens, waits while you draw, and reacts the moment you connect something questionable.

Start a mock interview