Imagine pouring a large jug of water into a funnel. Dump it all at once and it overflows everywhere; pour it at a controlled pace and every drop makes it through. When your application calls an external API — a payment provider, a mapping service, a partner feed — you're pouring requests into someone else's funnel, and they've decided exactly how fast it can drain.
The Rate Limiting pattern is how you do the controlled pour. Rather than firing requests the instant you have them and hoping the provider keeps up, you deliberately pace your own outbound traffic to stay within the limit the downstream service allows.
The problem
Almost every external service caps how often you may call it — say, 100 requests per second, or 10,000 per day. Exceed it and you don't just get a polite slowdown: you get rejected requests, 429 Too Many Requests errors, temporary bans, or even billing penalties. Your work fails not because anything was wrong with it, but because you sent it too fast.
This is easy to confuse with throttling, but the direction is opposite. Throttling is defensive on the inbound side — it protects your service from being overwhelmed by callers. Rate limiting is considerate on the outbound side — it protects a downstream service from being overwhelmed by you. Naively retrying rejected calls only makes it worse, hammering an already-saturated limiter and triggering ever-longer penalties.
Below, a nightly sync fires 20 calls at a partner that accepts 5 a second, and retries every 429 a second later. Before the retries start, predict how many requests it will take to get all 20 through.
How it works
The classic mechanism is a token bucket that lives in your own client. The bucket holds tokens and refills at exactly the rate the downstream service permits. Every outbound call must spend a token. If one is available, the call goes immediately; if the bucket is empty, the call waits in a queue until a token arrives, instead of being fired off to fail.
The bucket's two numbers do different jobs. Its refill rate sets the long-run pace: 5 tokens a second means 5 calls a second, however much work piles up. Its size caps the burst: after a quiet spell the bucket is full, so that many calls can go at once, but never more. A spike of 1,000 jobs drains out smoothly at the allowed pace instead of being rejected en masse.
Step through the same 20 calls below, this time behind a bucket that holds 5 tokens and gains 5 more every second. Predict when the last call leaves, then flip to No limiter to compare the two at each step. (Real limiters usually drip tokens in one at a time, here one every 200 ms, which spreads the calls out even more evenly. The arithmetic is the same.)
Notice what pacing did not cost you: both versions finished at the same moment, because the partner's quota set the speed all along. The bucket just stopped you from paying for that speed limit in errors. Even a well-paced client should still expect the odd 429 — clocks drift, and other clients may share your quota — so honor the Retry-After header and put the call back in the queue rather than retrying it on the spot.
Coordinate the bucket across instances. A per-process limiter is fine for one worker, but ten workers each pacing to the full limit will collectively blow past it tenfold. When you scale out, the token bucket usually needs to live in shared state (a cache like Redis) so the whole fleet shares one budget.
You run 4 copies of a sync worker. Each has its own token bucket, set to the partner's limit of 5 calls a second. What does the partner see at peak?
When to use it
Use rate limiting whenever you call a service that publishes a quota and you'd rather pace yourself than be cut off. It pairs naturally with queue-load-leveling: a queue absorbs the bursty work, and the rate limiter drains it at a sustainable speed. It also makes your retry logic far gentler — instead of retrying immediately into a wall, retries wait for the next available token, so you stop amplifying the very congestion you're trying to recover from.
Where it's overkill: low-volume calls that never approach any limit, or fire-and-forget traffic where the occasional rejection genuinely doesn't matter. But the moment you're doing bulk work against a metered API — sending notifications, syncing records, scraping a feed — pacing your outbound flow is the difference between steady throughput and a stream of rejections.
A limiter smooths bursts; it can't create quota. If your app produces 8 jobs a second against a 5-a-second limit, the queue grows by 3 every second, forever, and work ends up waiting for hours. Put a bound on the queue, watch how old its oldest item is, and when it backs up, slow the producer, drop low-value work, or ask the provider for a bigger quota.
Your nightly import gets a stream of 429s from a partner API, even though you already retry with exponential backoff. What fixes the problem at its source?