Contents

Architecture & System Design › Cloud Design Patterns

Scheduler Agent Supervisor

Coordinating a multi-step job and recovering steps that fail.

Also known as: scheduler agent supervisor, scheduler-agent-supervisor, distributed task pattern

Scheduler-agent-supervisor coordinates distributed work in three roles: the scheduler (orders tasks — when, in what sequence), agents (execute tasks on workers), and the supervisor (monitors progress, retries failures, reassigns stuck work). Cron-at-scale formalised: reliable distributed execution with visibility and recovery built in.

scheduler: run A, then B+C parallel, then D (workflow definition)
agents: execute on fleet (report heartbeats + results)
supervisor: stuck? failed? → retry/reassign/escalate

It suits scheduled pipelines, multi-step workflows and fleet operations (rollouts, scans, migrations) — anywhere fire-and-forget jobs need ordering, monitoring and guaranteed completion. Workflow engines productise the pattern; hand-rolled versions serve narrower needs.

The classic mistakes:

  • Scheduler as SPOF. One scheduler instance dying halts all coordination. Elect/replicate schedulers; agents tolerate scheduler gaps gracefully.
  • Agents without heartbeats. Silent agent death strands tasks as “running” forever. Heartbeat every task; supervisor times out and reassigns.
  • Non-idempotent tasks. Retried tasks double-applying (duplicate emails, double charges) corrupt under the exact failures supervision exists for. Idempotence prerequisite, not optional.
  • Supervisor split-brain. Two supervisors both reassigning duplicate everything. Single supervisor (elected) or partitioned supervision with fencing.
  • Opaque progress. Tasks as black boxes (“running” for hours, no detail) defeat monitoring and debugging. Structured progress reporting per task.
  • Retry storms. Blanket aggressive retries on systemic failure amplify precisely when systems strain. Bounded retries with backoff, dead-lettering, escalation.
  • Schedule drift. Clock skew and scheduler gaps shifting cron semantics (missed or doubled runs). Idempotent windows and catch-up policies defined explicitly.

When to use it: ordered, monitored, recoverable distributed work — pipelines, rollouts, fleet ops. Scheduler decides, agents do, supervisor guarantees: reliability through separated concerns.