You order a custom cake. The bakery doesn't make you stand at the counter for three hours while they bake it — they take your order, hand you a ticket number, and tell you to check back later. When you return and your number's up, you collect the cake.
The Asynchronous Request-Reply pattern is that bakery ticket, for APIs. Some operations are just too slow to finish inside a single HTTP request. Rather than holding the connection open and hoping it doesn't time out, you accept the request immediately, hand back a ticket, and let the caller check back for the result.
The problem
HTTP is built for quick, synchronous exchanges: send a request, get a response, done. But plenty of real work — generating a report, transcoding a video, running a big calculation — takes seconds or minutes. Try to do it inside one request and everything fights you: browsers and load balancers time out after 30–60 seconds, the connection ties up a server thread the whole time, and a network blip means the client has no idea whether the work finished or not.
The nastiest part is that a timeout only closes the connection. The server usually keeps working, because nothing tells it to stop. Step through one report that takes about 90 seconds. First predict what the timeout does to the work; then, when the client is left holding a bare 504, decide what it should do and see where that leads.
Either way, the client is stuck because it has no handle on the job: no way to ask "how is the report I asked for doing?" You can't just make the slow thing fast. So instead of pretending the call is quick, you need a way to break the link between asking for the work and waiting for it to complete.
How it works
The flow has three steps. First, the client sends the request and the API accepts it immediately — it validates the input, drops a job onto a queue, and responds with 202 Accepted plus a status URL (the ticket) in the Location header. The connection closes in milliseconds.
Second, a background worker picks the job off the queue and does the heavy lifting at its own pace, completely decoupled from the original caller. Third, the client polls the status URL: it gets 200 OK with "still working" until the job finishes, at which point the status endpoint points to the finished result (often a 303 See Other redirect to the result resource).
Step through the same report, asynchronous this time. Predict what the very first call returns before you look. Then flip to Synchronous at any step to see the old approach at the same moment.
Make accepting a job idempotent. A client that times out before getting its 202 will retry — and you don't want to start the same expensive job twice. Key the request so a retry returns the existing job's status URL instead of creating a duplicate.
Your API returns 202 Accepted with a status URL, yet some users get the same report generated twice. What's the most likely cause?
Designing the status endpoint
The status URL is the contract between client and server, so a few details matter:
- Tell clients how often to ask. Send a
Retry-Afterheader with the 202 and with each "still running" reply. Without it, eager clients poll every few milliseconds and the status endpoint takes more load than the real work. - Report failure, not just progress. The status must be able to say
failed, with a reason, so the client can stop polling and tell the user. - Keep status and result separate. The status resource describes the job; the result lives at its own URL. When the job finishes, the status endpoint redirects there, and the result can be cached or downloaded like any other resource.
- Clean up. Finished results and old job records should expire after a while, and a
DELETEon the status URL is a natural way to let clients cancel a job they no longer need.
A status that can only say "running" is a trap. If a worker crashes halfway through and nothing records it, the job's status stays "running" forever — and every client polls forever. Give each job a deadline, record failures where the status endpoint can see them, and only return 202 once the job is safely on the queue, so the ticket never points at work that was silently lost.
When to use it
Use this pattern whenever an operation is too slow to fit comfortably in a single request but the client still needs the eventual result — report generation, media processing, bulk imports, anything compute-heavy. It keeps your front-end connections short and snappy, and lets the background work scale independently behind a queue with competing consumers.
Skip it when the work is genuinely fast; the extra machinery and the polling protocol aren't worth it for a 50-millisecond query. And if the client can't poll — or you'd rather push the result the moment it's ready — consider webhooks, WebSockets, or server-sent events instead. Async request-reply is the simplest fit when the client is happy to come back and ask, "Is it done yet?"
A worker crashed halfway through job 42. A day later a client is still polling /status/42 and hearing "running". What was missing?