Cloud Native Patternsintermediate8 min

Rate Limiting

Smooth your own outbound calls so you stay under a downstream service's limits instead of slamming into them.

Imagine pouring a large jug of water into a funnel. Dump it all at once and it overflows everywhere; pour it at a controlled pace and every drop makes it through. When your application calls an external API — a payment provider, a mapping service, a partner feed — you're pouring requests into someone else's funnel, and they've decided exactly how fast it can drain.

The Rate Limiting pattern is how you do the controlled pour. Rather than firing requests the instant you have them and hoping the provider keeps up, you deliberately pace your own outbound traffic to stay within the limit the downstream service allows.

The problem

Almost every external service caps how often you may call it — say, 100 requests per second, or 10,000 per day. Exceed it and you don't just get a polite slowdown: you get rejected requests, 429 Too Many Requests errors, temporary bans, or even billing penalties. Your work fails not because anything was wrong with it, but because you sent it too fast.

This is easy to confuse with throttling, but the direction is opposite. Throttling is defensive on the inbound side — it protects your service from being overwhelmed by callers. Rate limiting is considerate on the outbound side — it protects a downstream service from being overwhelmed by you. Naively retrying rejected calls only makes it worse, hammering an already-saturated limiter and triggering ever-longer penalties.

Below, a nightly sync fires 20 calls at a partner that accepts 5 a second, and retries every 429 a second later. Before the retries start, predict how many requests it will take to get all 20 through.

How it works

The classic mechanism is a token bucket that lives in your own client. The bucket holds tokens and refills at exactly the rate the downstream service permits. Every outbound call must spend a token. If one is available, the call goes immediately; if the bucket is empty, the call waits in a queue until a token arrives, instead of being fired off to fail.

The bucket's two numbers do different jobs. Its refill rate sets the long-run pace: 5 tokens a second means 5 calls a second, however much work piles up. Its size caps the burst: after a quiet spell the bucket is full, so that many calls can go at once, but never more. A spike of 1,000 jobs drains out smoothly at the allowed pace instead of being rejected en masse.

Step through the same 20 calls below, this time behind a bucket that holds 5 tokens and gains 5 more every second. Predict when the last call leaves, then flip to No limiter to compare the two at each step. (Real limiters usually drip tokens in one at a time, here one every 200 ms, which spreads the calls out even more evenly. The arithmetic is the same.)

Notice what pacing did not cost you: both versions finished at the same moment, because the partner's quota set the speed all along. The bucket just stopped you from paying for that speed limit in errors. Even a well-paced client should still expect the odd 429 — clocks drift, and other clients may share your quota — so honor the Retry-After header and put the call back in the queue rather than retrying it on the spot.

Tip

Coordinate the bucket across instances. A per-process limiter is fine for one worker, but ten workers each pacing to the full limit will collectively blow past it tenfold. When you scale out, the token bucket usually needs to live in shared state (a cache like Redis) so the whole fleet shares one budget.

Check yourself

You run 4 copies of a sync worker. Each has its own token bucket, set to the partner's limit of 5 calls a second. What does the partner see at peak?

When to use it

Use rate limiting whenever you call a service that publishes a quota and you'd rather pace yourself than be cut off. It pairs naturally with queue-load-leveling: a queue absorbs the bursty work, and the rate limiter drains it at a sustainable speed. It also makes your retry logic far gentler — instead of retrying immediately into a wall, retries wait for the next available token, so you stop amplifying the very congestion you're trying to recover from.

Where it's overkill: low-volume calls that never approach any limit, or fire-and-forget traffic where the occasional rejection genuinely doesn't matter. But the moment you're doing bulk work against a metered API — sending notifications, syncing records, scraping a feed — pacing your outbound flow is the difference between steady throughput and a stream of rejections.

Watch out

A limiter smooths bursts; it can't create quota. If your app produces 8 jobs a second against a 5-a-second limit, the queue grows by 3 every second, forever, and work ends up waiting for hours. Put a bound on the queue, watch how old its oldest item is, and when it backs up, slow the producer, drop low-value work, or ask the provider for a bigger quota.

Check yourself

Your nightly import gets a stream of 429s from a partner API, even though you already retry with exponential backoff. What fixes the problem at its source?

Key takeaways

  • Rate limiting governs the calls *you* make outward, keeping you inside a downstream service's published quota.
  • It's the mirror image of throttling: throttling protects your own service from callers; rate limiting protects you from overwhelming someone else's.
  • A common mechanism is a token bucket — you spend a token per call and refill at the allowed rate, smoothing bursts into a steady stream.
  • Excess work waits in a buffer or queue rather than being fired off and rejected, so you avoid wasted calls and error storms.
  • It turns unpredictable bursts into a predictable, compliant flow — and avoids the penalty of tripping a provider's limiter.

Keep going