Backend Development › Queues & Async Processing
Workflow Engine
Durable orchestration of long multi-step processes, like Temporal.
Also known as: workflow engine, workflow orchestration, durable workflow
A workflow engine runs long-running, multi-step processes — “onboard a customer”, “process a refund”, “fulfil an order” — as an explicit, durable state machine. Each step can call a service, wait for an event, or branch, and the engine persists where the process is, retries failed steps, enforces timeouts, and resumes after a crash.
[start] → charge card → reserve stock → (wait for shipment event)
→ send confirmation → [done]
any step fails → compensate earlier steps
It brings structure to workflows that would otherwise be spread across queues, timers and ad-hoc database flags. The engine handles the hard parts: durability (the workflow survives restarts), retries with backoff, timeouts, and visibility into in-flight instances.
It’s closely related to the saga pattern (workflows with compensating actions) and often used to implement one. Examples include orchestration services and “durable execution” libraries.
The classic mistakes:
- Putting business logic in the wrong place. The engine orchestrates; the actual work should live in services or activities it calls. Mixing heavy logic into workflow definitions makes them hard to test and change.
- Non-idempotent steps. Workflow steps get retried; each must be safe to run more than once (see idempotence). The engine retries, so your steps must tolerate it.
- Ignoring determinism requirements. Some engines require the workflow definition to be deterministic (it replays history); calling
now()or random directly in the definition breaks replay. Use the engine’s facilities for time/randomness. - Forgetting compensation. A failure midway needs to undo prior steps (refund, release stock); without compensations you leave inconsistent state (see saga).
- Unbounded/long waits. A workflow that can wait indefinitely for an external event needs timeouts, or instances pile up forever.
- Operational blind spots. In-flight workflows, stuck instances and failures need monitoring and a way to retry or cancel manually.
- Reaching for it too soon. A simple two-step process may not need a workflow engine; a job queue or a database flag suffices. The engine earns its place with genuinely multi-step, long-lived, fault-prone processes.
When to use it: for complex, long-running, multi-service processes where durability, retries, timeouts and compensation are essential — clearly better than hand-rolling state across queues and flags. For short, simple asynchronous work, a job queue is lighter. It’s the orchestration layer that makes sagas and multi-step distributed transactions manageable in practice.