Cloud Native Patternsintermediate8 min

Asynchronous Request-Reply

When the work takes too long for one HTTP call, accept the request fast, do the work in the background, and let the caller poll for the result.

You order a custom cake. The bakery doesn't make you stand at the counter for three hours while they bake it — they take your order, hand you a ticket number, and tell you to check back later. When you return and your number's up, you collect the cake.

The Asynchronous Request-Reply pattern is that bakery ticket, for APIs. Some operations are just too slow to finish inside a single HTTP request. Rather than holding the connection open and hoping it doesn't time out, you accept the request immediately, hand back a ticket, and let the caller check back for the result.

The problem

HTTP is built for quick, synchronous exchanges: send a request, get a response, done. But plenty of real work — generating a report, transcoding a video, running a big calculation — takes seconds or minutes. Try to do it inside one request and everything fights you: browsers and load balancers time out after 30–60 seconds, the connection ties up a server thread the whole time, and a network blip means the client has no idea whether the work finished or not.

The nastiest part is that a timeout only closes the connection. The server usually keeps working, because nothing tells it to stop. Step through one report that takes about 90 seconds. First predict what the timeout does to the work; then, when the client is left holding a bare 504, decide what it should do and see where that leads.

Either way, the client is stuck because it has no handle on the job: no way to ask "how is the report I asked for doing?" You can't just make the slow thing fast. So instead of pretending the call is quick, you need a way to break the link between asking for the work and waiting for it to complete.

How it works

The flow has three steps. First, the client sends the request and the API accepts it immediately — it validates the input, drops a job onto a queue, and responds with 202 Accepted plus a status URL (the ticket) in the Location header. The connection closes in milliseconds.

Second, a background worker picks the job off the queue and does the heavy lifting at its own pace, completely decoupled from the original caller. Third, the client polls the status URL: it gets 200 OK with "still working" until the job finishes, at which point the status endpoint points to the finished result (often a 303 See Other redirect to the result resource).

Step through the same report, asynchronous this time. Predict what the very first call returns before you look. Then flip to Synchronous at any step to see the old approach at the same moment.

Tip

Make accepting a job idempotent. A client that times out before getting its 202 will retry — and you don't want to start the same expensive job twice. Key the request so a retry returns the existing job's status URL instead of creating a duplicate.

Check yourself

Your API returns 202 Accepted with a status URL, yet some users get the same report generated twice. What's the most likely cause?

Designing the status endpoint

The status URL is the contract between client and server, so a few details matter:

  • Tell clients how often to ask. Send a Retry-After header with the 202 and with each "still running" reply. Without it, eager clients poll every few milliseconds and the status endpoint takes more load than the real work.
  • Report failure, not just progress. The status must be able to say failed, with a reason, so the client can stop polling and tell the user.
  • Keep status and result separate. The status resource describes the job; the result lives at its own URL. When the job finishes, the status endpoint redirects there, and the result can be cached or downloaded like any other resource.
  • Clean up. Finished results and old job records should expire after a while, and a DELETE on the status URL is a natural way to let clients cancel a job they no longer need.
Watch out

A status that can only say "running" is a trap. If a worker crashes halfway through and nothing records it, the job's status stays "running" forever — and every client polls forever. Give each job a deadline, record failures where the status endpoint can see them, and only return 202 once the job is safely on the queue, so the ticket never points at work that was silently lost.

When to use it

Use this pattern whenever an operation is too slow to fit comfortably in a single request but the client still needs the eventual result — report generation, media processing, bulk imports, anything compute-heavy. It keeps your front-end connections short and snappy, and lets the background work scale independently behind a queue with competing consumers.

Skip it when the work is genuinely fast; the extra machinery and the polling protocol aren't worth it for a 50-millisecond query. And if the client can't poll — or you'd rather push the result the moment it's ready — consider webhooks, WebSockets, or server-sent events instead. Async request-reply is the simplest fit when the client is happy to come back and ask, "Is it done yet?"

Check yourself

A worker crashed halfway through job 42. A day later a client is still polling /status/42 and hearing "running". What was missing?

Key takeaways

  • Async request-reply decouples a slow operation from the HTTP request that triggers it.
  • The API accepts the request, returns 202 Accepted with a status URL, and processes the work in the background.
  • The client polls the status endpoint until the result is ready, then fetches it.
  • It keeps front-end connections short and lets the heavy work scale independently behind a queue.
  • The trade-off is more moving parts and a polling protocol the client must follow.

Keep going