Infrastructure & Operations › Observability
Metrics
Numeric measurements over time.
Also known as: application metrics, system metrics, monitoring metrics, time series metrics, Prometheus metrics
Metrics are numeric measurements recorded over time: requests per second, error rate, response time, queue depth, CPU usage, active users. They’re compact and cheap to store and query, which makes them ideal for dashboards, trends and alerts.
Compare with logs (detailed text events) and traces (one request’s path): metrics answer “how is the system behaving overall, and is it changing?”, and logs and traces help find out why (observability, logs, metrics and traces).
The basic types
| Type | Meaning | Example |
|---|---|---|
| Counter | Only goes up (resets on restart) | http_requests_total, errors_total |
| Gauge | A value that goes up and down | queue_depth, memory_bytes, active_connections |
| Histogram | Distribution of values in buckets, to compute percentiles | request_duration_seconds |
You usually care about a counter’s rate (requests per second), not its raw total (metric types).
from prometheus_client import Counter, Histogram
REQUESTS = Counter("http_requests_total", "Requests", ["method", "status"])
LATENCY = Histogram("http_request_duration_seconds", "Request latency")
@LATENCY.time()
def handle(request):
...
REQUESTS.labels(method="GET", status="200").inc()
(Prometheus is a widely used metrics system, often paired with Grafana for dashboards: Prometheus and Grafana.)
What to measure
Start with proven checklists:
- Golden signals: latency, traffic, errors, saturation (golden signals).
- RED for services: Rate, Errors, Duration (RED method).
- USE for resources: Utilization, Saturation, Errors (USE method).
- Business metrics too: orders per minute, sign-ups, payments failed. Often the first sign of trouble.
Pitfalls
- High cardinality labels: each distinct combination of label values creates a separate time series. Putting
user_idor a full URL with IDs in a label can create millions of series and crash or bankrupt your monitoring (cardinality). Use labels with small, bounded sets of values. - Averages hide the tail: use histograms and look at percentiles (percentiles).
- Measuring everything: lots of unused metrics add cost and clutter. Measure what you’d act on.
- Naming and units: consistent names, with the unit in the name (
_seconds,_bytes). - A metric nobody looks at or alerts on is wasted. Connect them to dashboards and alerts.