Architecture & System Design › Performance & Scalability
Capacity Planning
Predicting and providing the resources needed for future load.
Also known as: capacity management, resource planning, forecasting capacity, sizing infrastructure
Capacity planning is ensuring that the system will have enough resources (compute, memory, storage, network, database capacity, licenses) for the load it will face, before the load arrives, without wasting money on resources it won’t use. Under-provisioning causes outages and slowness. Over-provisioning wastes budget.
The basic loop
- Know current usage and the load drivers. Requests per second, data growth, active users, batch sizes, and the resources they consume (CPU, memory, I/O, connections).
- Find the relationship between load and resources: “each 1,000 requests per second needs about 4 cores”, “each million users adds 200 GB”.
- Forecast load: organic growth, seasonality (holidays, month-end, sales), planned launches and marketing events, new customers. Include peaks, not just averages (percentiles).
- Compute needed capacity, including headroom: a target utilization well below 100% so spikes and failures are absorbed (many teams plan for roughly 50 to 70% peak utilization, but it depends on the workload).
- Account for failure: if one node or zone dies, can the rest carry the load (N+1, N+2 redundancy) (redundancy)?
- Account for lead time: how long does it take to add capacity (instant autoscaling, or weeks to order hardware and approve a budget)?
- Validate with load tests (load testing) and watch real metrics.
- Repeat regularly, and compare forecasts with reality.
Measuring the right things
- Utilization and saturation, per resource: CPU, memory, disk I/O, network, database connections, queue depth (USE method, metrics).
- The bottleneck. Capacity is set by the first resource to run out, often a database or a downstream dependency, not the web servers.
- Efficiency per unit of work, to watch for regressions: a release that doubles CPU per request halves your capacity.
Practical considerations
- Autoscaling helps but has limits: scale-up delay, quotas, cold starts, the database that doesn’t scale with it. Don’t treat it as a substitute for planning.
- Cloud makes it flexible, not free. Know the cost per unit and set budgets (FinOps).
- Plan for non-linear behavior: systems degrade sharply as they approach saturation (queues grow, latency explodes), so linear extrapolation fails.
- Storage grows monotonically unless you manage retention and archiving (data retention).
- Plan for events: launches, campaigns, seasonal peaks, with a known date and a test beforehand.
- Quotas and limits: cloud service quotas, API limits, connection limits and licenses can bite before CPU does.
- Document assumptions and review them as the product changes.
Mindset
Capacity planning is reducing surprises. A rough estimate with clear assumptions and a regular check beats a precise model nobody updates. Combine quick estimation (back-of-the-envelope) with measurement, and keep headroom for the unexpected.