Cloud Native Patternsintermediate8 min

Compute Resource Consolidation

Pack several small, related workloads onto shared compute instead of paying for a separate machine per task.

Imagine renting a whole delivery van for every single parcel — one van for a letter, another van for a small box, a third for a postcard. Each van costs the same fixed amount whether it's full or nearly empty, and most of them roll out almost empty. The obvious fix is to load many parcels into one van. That's the idea behind Compute Resource Consolidation.

The problem

In the cloud it's tempting to give every small task its own compute instance — one VM or container for the image thumbnailer, another for the nightly report job, another for the email sender. It feels clean, but each instance carries fixed overhead: the OS, the runtime, monitoring agents, and a minimum size you pay for whether or not the task is busy. Since most of these tasks are bursty or light, you end up with a fleet of machines sitting at 5% utilization while the bill reflects 100% provisioning. You're paying for capacity nobody is using, and managing a sprawl of instances to boot.

Step through the bill for ten such services below. Each VM costs, say, $70 a month. Predict how much of the $700 actually pays for work before you look.

How it works

Instead of one instance per task, consolidate compatible tasks onto a shared set of compute units. A single instance (or small pool) hosts several cooperating tasks together, so the fixed overhead of that machine is amortized across all of them and average utilization climbs. The image worker, the report job, and the email sender share the same node, each taking a slice of its CPU and memory. You provision for the combined demand rather than summing each task's worst case in isolation — and because small tasks are rarely all busy at the same moment, that's usually far less hardware.

Step through it below: predict what each node's CPU meter reads once the ten services share two nodes. Then a bad deploy sends one service into a busy loop, and you'll see the price of sharing. Once you've seen it, switch on per-service limits and replay the same bad deploy.

Tip

Group tasks that get along. The cost win comes from sharing, and sharing works best for tasks with compatible scaling, lifecycle, and trust: they can be deployed, scaled, and patched together, and none of them needs a security boundary from the others. Keep anything that needs strong isolation — as the bulkhead pattern would dictate — on its own instance.

Check yourself

You packed 12 services onto 3 nodes with no resource limits. One service starts leaking memory, and soon services all over its node are being killed and restarted. What would have contained it?

Sharing safely

A shared node needs rules, the way a shared flat needs a rota. Three are standard, and container platforms and orchestrators support all of them:

  • Limits cap what one task can take: no more than this much CPU, no more than this much memory. A runaway task hits its own ceiling and slows down, or is restarted, instead of starving its neighbors.
  • Reservations (often called requests) guarantee the minimum a task needs, so the scheduler only places it on a node with that much room to spare.
  • Quotas cap a whole team or namespace, so one group can't fill the shared pool with its own workloads.

Then watch the node as a whole. Consolidation turns ten quiet machines into a couple of busy ones, so a node that fails now takes several tasks with it. Run at least two, keep some headroom for bursts, and alert on sustained high CPU or memory before your users notice.

Watch out

Don't size by the average. "Ten tasks at 5% each fit in half a machine" is only true if they're busy at different times. Ten report jobs that each average 5% over the week, but do all of it flat out between 9:00 and 17:00 on Monday, need ten machines' worth of CPU for those eight hours, all at once. Look at when each task is busy, consolidate the ones whose peaks don't line up, and size each node for the combined peak.

When to use it

Consolidation pays off when you have many small, light, or bursty tasks whose individual instances would mostly sit idle, and where those tasks have similar scaling and lifecycle needs and a comparable trust level. It complements elastic scaling and competing consumers by making each shared node do real work, and it's a manual cousin of serverless, which consolidates for you behind the scenes.

Don't consolidate tasks that genuinely conflict: those with very different scaling curves, hard isolation or security boundaries, or peaks that land at the same time and would force you to provision for the worst of all worlds. When in doubt, isolate what must be isolated and pool the rest.

Check yourself

Ten report jobs each average 5% CPU over the week on their own VMs, but all of that work happens on Monday between 9:00 and 17:00, when each one runs flat out. You plan to put all ten on one VM, since 10 × 5% = 50%. What's wrong with the plan?

Key takeaways

  • Many small tasks each on their own instance waste money: every machine carries fixed overhead and most sit nearly idle.
  • Compute Resource Consolidation packs multiple cooperating tasks onto a shared set of compute units to raise utilization.
  • It lowers cost and management overhead by amortizing the fixed cost of each instance across many tasks.
  • The trade-off is reduced isolation: tasks contend for CPU and memory, and one noisy neighbor can starve the rest of its node — per-task limits and quotas are what make sharing safe.
  • Size for the combined peak, not the average: consolidate tasks whose busy times don't line up, and isolate the ones that genuinely conflict.

Keep going