Contents

Infrastructure & Operations › Observability

Metrics

Numeric measurements over time.

Also known as: application metrics, system metrics, monitoring metrics, time series metrics, Prometheus metrics

Metrics are numeric measurements recorded over time: requests per second, error rate, response time, queue depth, CPU usage, active users. They’re compact and cheap to store and query, which makes them ideal for dashboards, trends and alerts.

Compare with logs (detailed text events) and traces (one request’s path): metrics answer “how is the system behaving overall, and is it changing?”, and logs and traces help find out why (observability, logs, metrics and traces).

The basic types

TypeMeaningExample
CounterOnly goes up (resets on restart)http_requests_total, errors_total
GaugeA value that goes up and downqueue_depth, memory_bytes, active_connections
HistogramDistribution of values in buckets, to compute percentilesrequest_duration_seconds

You usually care about a counter’s rate (requests per second), not its raw total (metric types).

from prometheus_client import Counter, Histogram

REQUESTS = Counter("http_requests_total", "Requests", ["method", "status"])
LATENCY = Histogram("http_request_duration_seconds", "Request latency")

@LATENCY.time()
def handle(request):
    ...
    REQUESTS.labels(method="GET", status="200").inc()

(Prometheus is a widely used metrics system, often paired with Grafana for dashboards: Prometheus and Grafana.)

What to measure

Start with proven checklists:

  • Golden signals: latency, traffic, errors, saturation (golden signals).
  • RED for services: Rate, Errors, Duration (RED method).
  • USE for resources: Utilization, Saturation, Errors (USE method).
  • Business metrics too: orders per minute, sign-ups, payments failed. Often the first sign of trouble.

Pitfalls

  • High cardinality labels: each distinct combination of label values creates a separate time series. Putting user_id or a full URL with IDs in a label can create millions of series and crash or bankrupt your monitoring (cardinality). Use labels with small, bounded sets of values.
  • Averages hide the tail: use histograms and look at percentiles (percentiles).
  • Measuring everything: lots of unused metrics add cost and clutter. Measure what you’d act on.
  • Naming and units: consistent names, with the unit in the name (_seconds, _bytes).
  • A metric nobody looks at or alerts on is wasted. Connect them to dashboards and alerts.