War story

The Elasticsearch deadlock that parked every index: ILM warm phases on a single node

, 1 min read

This cluster is a single Elasticsearch 9.1.5 node, and it holds compliance audit evidence for a US fintech. A stuck lifecycle there is not cosmetic. Twice it stopped moving indices through their lifecycle, and I treated those as two separate incidents. They weren't.

What broke

First, the cluster climbed to its shard ceiling — 999 of 1,000 — and indices stopped moving through their lifecycle. I spotted it from the host's disk space.

A month later it came back in a different form: a warm-phase deadlock, with indices waiting in check-migration indefinitely.

What I tried

I treated the shard-ceiling incident and the warm-phase deadlock as two separate problems. They were the same root cause in two different forms.

What worked

On a single node, any ILM policy with a warm or cold phase must set replicas to zero in that phase. Otherwise check-migration waits for replica copies that can never be allocated.

In the policy, that's one action in the phase:

"warm": {
  "actions": {
    "allocate": { "number_of_replicas": 0 }
  }
}

Fixing the policy doesn't move indices that are already stuck. I moved those to their next phase manually.

After the fix: shards 810 → 433, ILM errors at zero, disk steady at 71%. I verified the warm-phase deadlock closed on 12 August 2026.

Worth knowing

  • YELLOW with around 40 unassigned shards is now the expected steady state. Fleet creates new indices with one replica, and the warm phase strips that replica at its first transition.
  • The first incident was caught from disk space, not from Elasticsearch itself. Afterwards I set up an alert rule so it can't go unnoticed again.

Start with a $500 audit.

See exactly where you stand. Actionable findings in a week.

The full report and debrief call. Delivery in 5–7 business days.