Drawbly

Technical explanations · Drawbly

API rate limiting: why can six requests pass at once?

By Drawbly ·

An API can allow “three requests per ten seconds” and still accept six requests almost together. The answer depends on how the limit counts. This worked example compares a fixed window with a token bucket using the same stated rate.

Six fictional requests arrive from 9.8 to 10.2 seconds. A fixed window accepts three before the ten-second boundary and three after it. A three-token bucket accepts the first three but rejects the next three because it has refilled less than one token.
Same six arrivals, two counting rules. Green means accepted; red means rate limited. The clock and limit are fictional teaching values.

One burst, just across a boundary

Imagine a fictional report API that permits three requests per ten seconds for one client key. Six requests arrive at elapsed times 9.8, 9.9, 9.9, 10.0, 10.1 and 10.2 seconds. Assume the client has sent nothing earlier, the bucket starts full, each request costs one token, and the counter uses the same clock for all six. These assumptions make the arithmetic inspectable; actual services can have different scopes and distributed enforcement.

Fixed window: the counter resets at ten seconds

A fixed-window limiter divides time into blocks: [0,10), [10,20) and so on. The first three arrivals consume the entire first window's allowance. At 10.0 the count resets, so the next three are also accepted. The result is six accepted in 0.4 seconds, even though the configured quota reads “three per ten seconds.” The boundary effect is a property of separate window counters, not an implementation error. A fixed window is simple when a coarse quota is acceptable; it is a poor fit if that boundary burst would overload the protected work.

Redis's rate-limiting guide describes this edge effect and compares alternative algorithms. Counting per client needs an explicit key, and a shared counter needs an atomic update if requests can reach several workers; otherwise simultaneous checks can overshoot the intended limit.

Token bucket: the budget refills continuously

Now give the same client a bucket of three tokens, refilling at 0.3 token per second. A request consumes one token when at least one is available. The first three arrivals spend three tokens from the initially full bucket. By 10.2, only 0.12 token has refilled since the first arrival at 9.8; there is still no whole token to spend. All three later arrivals are rate limited. A quiet client can store up to three tokens and make a short burst, but the bucket cannot spend another full burst merely because the wall clock crossed ten seconds.

This is an idealized single-key model. AWS API Gateway documents token-bucket throttling with a steady rate and burst capacity, and warns that its throttle targets are best effort rather than guaranteed hard ceilings. A distributed service, concurrent requests or separate keys may show different observed results. State the key, bucket capacity, refill rule and whether limiting is local or shared before promising a numerical bound.

What should a client do with HTTP 429?

The rejected requests in this example could receive 429 Too Many Requests. RFC 6585 says a 429 response may include a Retry-After header; it does not prescribe how the server identifies a user or counts requests. RFC 9110 allows Retry-After as either an HTTP date or a nonnegative number of delay seconds. Honor a supplied delay. Without one, slow retries with backoff and jitter so many clients do not retry together. A 429 is not permission to resend immediately in a loop.

Retries raise a separate correctness question: if a timed-out write actually succeeded, a retry can repeat the side effect. The idempotency-key example shows how to address that duplication. A rate limiter controls arrival pressure; an idempotency key controls repeated work. An API may need both.

Which rule should you draw?

Use a fixed-window counter when its simplicity and a boundary burst are acceptable. Use a token bucket when you want to permit a limited burst after quiet time while controlling the sustained arrival rate. Neither is a complete capacity plan: identify the expensive operation, concurrency limit, shared state, and response behavior too. A sliding-window or queue design may fit a stricter rolling bound, but it is outside this small example.

Open the editable Drawbly diagram and replace the fictional clock and counts with your own API's key and rule. Keep the six arrival marks aligned between the two rows; that alignment makes the difference easy to review. For a related request flow, see the webhook versus polling example.

Technical references

AWS API Gateway throttling, Redis rate-limiting algorithms, RFC 6585 §4 and RFC 9110 §10.2.3. The report API and six requests are fictional.