Cloud Native Patternsadvanced8 min

Leader Election

When several identical instances run side by side, pick exactly one to own the jobs that must happen only once.

Picture a relay team where every runner is equally fast and equally ready. The race only works if exactly one of them is holding the baton at any moment. If two grab it, chaos; if nobody does, the team stalls. Someone has to be the runner right now — and if they trip, a teammate must snatch the baton instantly.

Leader Election gives a fleet of identical service instances that same single-baton rule. They're interchangeable, but for certain jobs exactly one of them must be in charge, and the role must pass on cleanly if that one goes down.

The problem

Running many copies of a service is how we scale and stay available — and we work hard to keep those copies stateless and interchangeable so any of them can handle any request. But some tasks break if more than one instance does them. Think of a nightly billing job, advancing a shared workflow, or assigning work from a queue: if every instance runs it, you get duplicated effort, double charges, or corrupted state.

Try the two obvious rules below. Pick one at midnight, predict what happens on the second night, then flip to the other rule.

Neither rule works. Letting everyone run the job fails every night, and hard-coding "instance A does it" fails the first night A crashes, gets redeployed, or scales away — quietly, because nothing complains about work that didn't happen. You need the fleet to agree, on its own and continuously, on which single instance currently owns the special work.

How it works

The instances compete to acquire one shared, exclusive token — usually a lease or distributed lock backed by a store that can guarantee only one holder at a time (a database row, a blob lease, a coordination service like ZooKeeper or etcd). Whichever instance grabs it becomes the leader and takes responsibility for the single-owner work; the rest see it's taken and wait as standbys.

The catch is failure. The leader must keep renewing its lease on a heartbeat. If it crashes or hangs, it stops renewing, the lease expires, and the standbys race to claim the now-free token — electing a fresh leader automatically, with no human in the loop. Each time the lease changes hands, the store hands out a higher term number (also called an epoch or fencing token). It costs nothing while things go well, and it's what lets the system reject an old leader that wakes up late.

Step through it on a timeline below, one lane per node. When the leader goes quiet, predict what the standbys do, then flip to Stays healthy to see what was keeping them quiet.

Tip

Don't roll your own consensus. Correct distributed election is notoriously subtle. Use a battle-tested lease or coordination primitive — a blob lease, etcd, ZooKeeper, or a database row claimed with a conditional update — rather than inventing one.

Check yourself

Your leader renews a 15-second lease every 5 seconds, and it crashes right after a renewal. Roughly how long until a standby can take over?

Split brain and fencing

Leases have a nasty edge case. A leader can stall without dying — a long garbage-collection pause, a frozen VM, a network blip — and when it resumes, it has no idea that time passed. As far as it knows it still holds the lease, so it carries on working while the standby that replaced it is working too. Two leaders at once is called split brain, and it's exactly the double-running you set out to prevent.

The fix is the term. Every action the leader takes carries its term, and the store it writes to remembers the highest term it has seen. Once the new leader has written with term 2, anything stamped term 1 is rejected, so the old leader finds out it was replaced the moment it tries to act. Checking the clock before each write isn't enough on its own: the pause can strike between the check and the write.

Watch out

A paused leader doesn't know it was paused. Followers can't tell a dead leader from a slow one, and a slow leader can't tell that its lease ran out. Make the lease longer than your worst normal pause, have the store reject stale terms, and keep the leader's work idempotent so the rare overlap can't double-charge anyone.

Check yourself

A leader froze in a long garbage-collection pause, and a standby took over as term 8. The old leader wakes up and sends a write stamped term 7. What should happen?

When to use it

Reach for leader election when a task in a multi-instance system must be performed by exactly one instance at a time, yet must survive the loss of whichever instance currently holds the role — coordinating a workflow, running a singleton scheduler, or managing shared resources.

Don't use it for work that can run in parallel. If you've got a pile of independent messages to process, competing consumers is faster and simpler — many instances pulling from the same queue, each handling different items. Leader election adds a coordination bottleneck and a single point of activity, so reserve it strictly for the jobs that genuinely demand a single owner.

Key takeaways

  • Leader election designates one instance among many to coordinate work that must not run in parallel, and passes the role on when that instance fails.
  • Election usually relies on a shared lease that only one instance can hold at a time; the others wait as standbys.
  • The leader must keep renewing its lease. If it stops, the lease expires and a standby takes over, so failover takes about one lease length.
  • Each new leader gets a higher term. Stamp the leader's writes with it, so a paused old leader is rejected instead of causing split brain.
  • It's the right tool only for genuinely single-owner work — for parallelizable jobs, prefer competing consumers.

Keep going