Contents

Engineering Craft › Debugging · also in Performance & Scalability

Flame Graph

A visualization of where a program spends its time.

Also known as: flame graph, flamegraph, profiling visualization

A flame graph is a picture of a profiler’s output. The profiler samples the program’s call stacks many times a second; the flame graph stacks those samples so you can see, at a glance, where the time goes. It’s the standard way to read a CPU profile.

How to read it:

  • The x-axis is the call stack, and a frame’s width is the share of samples in which that function was on the stack. Width = time.
  • The y-axis is call depth: the function at the bottom called the one above it, and so on up to whatever was running.
  • A wide frame is a hot spot — something taking a lot of time. A wide frame high up, with a narrow parent, is a function that’s expensive in itself; a wide parent with many children means the time is spread below it.
[                 main (whole program)                 ]
[  parse (40%)  ][  compute (55%)  ]
      [ lex ][ read ]   [ hotfn ][ gc ][ other ]

(An inverted or “icicle” graph flips it; the meaning is the same.)

The classic mistakes:

  • Reading height as importance. Tall doesn’t mean slow. A deeply nested but narrow stack is cheap; the wide frames are where the time is.
  • Confusing sampled and instrumented profiles. Sampling profilers add little overhead and are safe to run under load; instrumenting every call changes timing and can mislead. Know which you have.
  • Only looking at CPU. A program can be slow because it’s waiting — on I/O, locks, or the network — which a CPU profile won’t show. Off-CPU analysis (or traces/spans) is needed for that.
  • Optimising the wrong frame. The widest leaf is not always the best fix; sometimes an inefficient caller producing huge inputs is the real cause. Read the parents, not just the hot leaf.
  • Trusting it under unrealistic load. A profile of an idle process tells you nothing about production. Profile with representative work.

A flame graph turns “it feels slow” into “these two functions dominate”. Combined with observability for the system view and latency percentiles for user impact, it points you at a specific function to optimise — which is exactly what makes profiling worth the effort.