100 requests per minute, and 200 got through

A fixed window limiter set to 100 requests a minute can let 200 through in two seconds. I follow the interview question from that first sketch all the way to two servers racing for one counter.

A fixed window limiter set to 100 requests per minute

The limiter in the animation above is set to 100 requests per minute, and it just let 200 through in under two seconds. It isn't broken. It's the simplest rate limiter there is, and a lot of real code works exactly like it: count requests per clock minute, wipe the count when the minute changes. A client that sends right before the change and right after it gets twice the limit through.

This is also a classic interview question. You're asked to design a rate limiter, you sketch a counter, and the interviewer asks what happens at 0:59. I'll follow that conversation here. Each section starts with the next question, and the animations show the actual algorithms, so you can mess with them and see where they break.

A counter and a clock

The interviewer asks

“Design a rate limiter for our public API. 100 requests per minute per API key.”

The first answer most people give is a good place to start. Keep one counter per API key per minute. When a request comes in, look up the counter for the current minute. If it's under the limit, add one and let the request through. If it's at the limit, answer with HTTP 429 Too Many Requests [1].

function allow(key, now) {
  const minute = Math.floor(now / 60_000);
  const count = counters.get(`${key}:${minute}`) ?? 0;
  if (count >= LIMIT) return false;
  counters.set(`${key}:${minute}`, count + 1);
  return true;
}

That's a fixed window. You need one integer per client, and old counters just expire. The animation below uses a limit of 10 per minute so you can count the slips. Press Send a request as fast as you like, or switch the traffic to “right at 1:00”.

Fixed window, limit 10 per minute

Watch the yellow band. With steady traffic it never goes above 10. With traffic bunched around the boundary it shows 20: ten just before 1:00 and ten just after. Figma's engineering blog shows the same thing with a limit of 5 per minute [2]. Five requests at 11:00:59, five more at 11:01:00, and all ten go through.

How bad that is depends on what the limit protects. A 2x spike for a second or two is fine if you just want to stop a script from hammering you all day. It's a real problem if the thing behind the limiter can't take more than 100 a minute.

Remember every request

The interviewer asks

“OK, the boundary leaks. Can you make it exact?”

Stop thinking in clock minutes. Instead of “requests this minute”, count “requests in the last 60 seconds”, measured back from right now. For that you need the time of every request you let through. That list is the sliding log.

function allow(key, now) {
  const log = logs.get(key) ?? [];
  while (log.length && log[0] <= now - 60_000) log.shift();
  if (log.length >= LIMIT) return false;
  log.push(now);
  logs.set(key, log);
  return true;
}

The next animation gets the same bunched traffic. The bracket is the last 60 seconds, and a timestamp gets crossed off when it falls out of it.

Sliding log, limit 10 per minute

It never lets more than 10 through in any 60 seconds. What it costs you is memory: each client now needs a timestamp for every allowed request instead of one integer. At 100 per minute that's up to 100 timestamps per client, 8 bytes each (a 64-bit timestamp in milliseconds), so a million active clients need 800 MB just for timestamps. Figma did this math for a limit of 500 requests a day and got 5 million stored values for 10,000 users [2].

Two counters and a guess

The interviewer asks

“Can you get close to that without storing every timestamp?”

Keep the per-minute counters from the first version, but when a request arrives, look at two of them: this minute's and last minute's. Then pretend last minute's requests were spread evenly across it. If we're 15 seconds into this minute, the last 60 seconds still include the final 45 seconds of the previous one, so you count 75% of it.

const elapsed = now % 60_000;
const estimate = previous * (60_000 - elapsed) / 60_000 + current;
if (estimate + 1 > LIMIT) return false;

Cloudflare described this in 2017 with a worked example [3]: a limit of 50 per minute, 42 requests in the previous minute and 18 so far in the current one, 15 seconds in: 42 × 0.75 + 18 = 49.5. Figma wrote about something similar the same year [2]. For an hourly limit they keep 60 one-minute counters and add them up, which means more counters to store but less guessing.

In the next one, the clock moves through minute 2, and you pick how minute 1's traffic really came in. The dashed slips under the line are what the estimate assumes, and the solid ones are what happened.

Sliding window counter, limit 10 per minute

If minute 1's traffic came early, the guess still counts requests that are already older than 60 seconds, so it's too strict. If it came late, the guess spreads them out, counts too few, and lets a few extra through. So you store two integers per client instead of a whole log, and the count is sometimes a bit off.

Cloudflare measured how far off on real traffic: 400 million requests from 270,000 sources. The estimate missed the real rate by 6% on average, but that only changes the answer for clients sitting close to the limit, so just 0.003% of requests were wrongly allowed or wrongly limited [3]. Their counters lived in a Memcached fork with simple commands like GET, SET and INCR, and that's all this approach needs.

A bucket of tokens

The interviewer asks

“Some clients send bursts. A page load fires 20 API calls at once. Can you allow that without raising the average rate?”

Every limiter so far counts the requests from the last minute. A token bucket doesn't count requests at all. Each client gets a bucket that holds up to N tokens, and tokens drip in at a fixed rate. A request takes one token, and if there's none left it gets a 429. A full bucket lets a burst of N through at once, and after that the client gets exactly the refill rate.

The drip doesn't need a timer or a background job. You store two numbers per client, the token count and the time you last looked, and top the bucket up when the next request shows up:

function allow(b, now) {
  const refill = (now - b.last) / 1000 * RATE;
  b.tokens = Math.min(CAPACITY, b.tokens + refill);
  b.last = now;
  if (b.tokens < 1) return false;
  b.tokens -= 1;
  return true;
}

In the animation, the coin forming under the tap is that top-up. Try a bucket of 1 with bursty traffic, then a bucket of 12.

Token bucket

So there are two settings instead of one: the bucket size decides how big a burst you accept, and the rate decides the average. Stripe wrote in 2017 that they use a token bucket for their API, with the limiters backed by Redis [4]. AWS API Gateway documents its throttling the same way: the rate is tokens added per second, and the burst is the size of the bucket [5].

It doesn't fix 0:59 on its own, though. A bucket of 100 that refills at 100 per minute can still let about 200 through in 60 seconds: the 100 it saved up, plus 100 more dripping in. If a hard ceiling per minute matters, keep the bucket small compared to the rate.

A bucket with a hole

The interviewer asks

“The payment provider behind this API accepts 2 calls per second. Bursts or not. Now what?”

The leaky bucket is the other famous bucket, and the name gets used for two different things. In the first one, the requests themselves go into the bucket and leak out of a hole at the bottom at a steady rate. The bucket is a queue. A burst doesn't get rejected, it waits its turn, and you only answer 429 when the queue is full.

The animation below sends the same bursts through both, with the same size and rate: the token bucket on top, the queue below.

Token bucket and leaky bucket, same size and rate

The token bucket passes the burst straight to the server. The queue turns it into an even drip, but requests have to wait: here the last one in each burst waits almost 3 seconds. That's what you want in front of something that really can't take a burst, like that payment provider. NGINX's limit_req works this way when you set burst and leave out nodelay: extra requests are delayed and only rejected once they go past the burst size [6]. It answers with 503 by default, so if you want a 429 you have to set limit_req_status.

The second meaning counts instead of queueing. Shopify's API docs describe it with marbles: every request drops a marble into the bucket, the bucket leaks at a fixed rate, and a full bucket means 429 [7]. Nothing waits in this one. Turn it upside down and you have a token bucket, because marbles in the bucket are just tokens missing from it. For the same size and rate, the two make the same decision on every request.

Many servers, one limit

The interviewer asks

“You have 20 API servers behind a load balancer. Where does the counter live?”

A counter in each server's memory doesn't work. A client whose requests get spread over 20 servers can get up to 20 times the limit, so 2,000 a minute instead of 100. The counter has to move to a shared store, often Redis, and that opens a new way to get it wrong. Here it is step by step:

Two servers, one counter

Read, decide, write is three steps, and another server can slip in between any two of them. With thousands of requests per second on the same key, it will. The fix is to let the store read and write in one operation. For a counter that's INCR, which adds one and returns the new value, and you decide based on what it returns [8]. For a token bucket, where the update is more than adding one, the usual answer is a small Lua script. Redis runs a script from start to finish without running anything else in between [9].

Even the Redis docs point out a race in one of their own rate limiter examples: a client that crashes between INCR and EXPIRE leaves a counter with no expiry, which the docs call a leak. Their fix is a Lua script too [8].

Saying no politely

The interviewer asks

“What do you send back when you say no?”

Status 429 Too Many Requests, from RFC 6585 [1]. The same section says the response may include a Retry-After header, which tells the client how long to wait, and RFC 9110 defines it as a whole number of seconds or an HTTP date [10]. A client that respects it stops hammering you. RFC 6585 also says caches must not store a 429.

GitHub goes further and tells you where you stand on every response, not only when you're blocked [11]. Authenticated users get 5,000 requests an hour, and the headers say how many are left and when the count resets, in UTC epoch seconds:

HTTP/1.1 200 OK
x-ratelimit-limit: 5000
x-ratelimit-remaining: 4987
x-ratelimit-used: 13
x-ratelimit-reset: 1790001200

There's an IETF draft to standardize this with two headers, RateLimit and RateLimit-Policy [12]. It's still a draft, version 11 as of May 2026. Plenty of blog posts show the older RateLimit-Limit and RateLimit-Remaining names, which were replaced in 2023.

Side by side

AlgorithmPer clientBurstsExact?Who wrote about it
Fixed window1 counterup to 2x the limit at the boundaryper clock minuteRedis docs, INCR pattern
Sliding log1 timestamp per requestnever over the limityesFigma (2017), considered it
Sliding window counter2 countersclose to the limitestimateCloudflare (2017)
Token bucket2 numbersup to the bucket size, on purposethe average, not every 60 sStripe (2017), AWS API Gateway
Leaky bucket, queuethe waiting requestssmoothed, requests waityes, on the way outNGINX limit_req with burst
Leaky bucket, meter2 numberssame as token bucketthe average, not every 60 sShopify

If I had to pick one without knowing more, I'd take the token bucket with a small bucket. It stores two numbers per client, and you decide up front how big a burst is OK.

Sources

  1. Nottingham, M. and R. Fielding, “Additional HTTP Status Codes”, RFC 6585, April 2012. Section 4.
  2. Mahdi, N., “An alternative approach to rate limiting”, Figma blog, April 2017.
  3. Desgats, J., “How we built rate limiting capable of scaling to millions of domains”, Cloudflare blog, June 2017.
  4. Tarjan, P., “Scaling your API with rate limiters”, Stripe blog, March 2017.
  5. Amazon Web Services, “Throttle requests to your REST APIs”, Amazon API Gateway Developer Guide.
  6. NGINX, “Module ngx_http_limit_req_module”, nginx.org documentation.
  7. Shopify, “Shopify API limits”, Shopify developer documentation.
  8. Redis, “INCR”, command documentation. See the rate limiter patterns.
  9. Redis, “Scripting with Lua”, Redis documentation.
  10. Fielding, R., Ed., Nottingham, M., Ed., and J. Reschke, Ed., “HTTP Semantics”, STD 97, RFC 9110, June 2022. Section 10.2.3.
  11. GitHub, “Rate limits for the REST API”, GitHub Docs.
  12. IETF HTTPAPI Working Group, “RateLimit header fields for HTTP”, Internet-Draft draft-ietf-httpapi-ratelimit-headers-11, May 2026.

Get the next article by email

I send one email when a new article is out, and nothing else.

Keep reading

Cache stampede

Soon

A popular key expires and ten thousand requests hit the database in the same second.

System Design

Design a URL shortener

Soon

The classic warm-up question, and the follow-ups that make it hard.

Interview questionsSystem Design