Contents

Data Engineering › Batch & Distributed Processing

Notebooks (Jupyter)

Interactive documents mixing code, output and notes.

Also known as: Jupyter, Jupyter notebook, JupyterLab, Colab

A notebook is an interactive document made of cells: some contain code, some text, and the output (tables, charts) appears right below the code that produced it. Jupyter is the best-known one; Google Colab and Databricks notebooks are similar.

# cell 1
import pandas as pd
df = pd.read_csv("orders.csv")

# cell 2
df.groupby("country")["total"].sum().sort_values(ascending=False).head()
# output table appears here

You run cells one at a time and inspect results as you go, which makes notebooks ideal for exploring a new dataset, trying ideas, drawing quick charts, and sharing an analysis with notes.

Where they cause trouble

  • Hidden state. Cells can run out of order, and a variable from a deleted cell is still in memory. The notebook looks correct but can’t be reproduced from top to bottom. Habit: restart the kernel and run all cells before trusting or sharing it.
  • Hard to review in Git. The file stores outputs and metadata as JSON, so diffs are noisy. Clear outputs before committing, or use tools that pair a notebook with a plain text version.
  • Hard to test and reuse. Logic buried in cells can’t easily be imported or unit tested.
  • Secrets. Don’t paste credentials into cells; outputs and files get shared.
  • Large outputs (huge tables) bloat the file.

Notebooks and production pipelines

Notebooks are great for exploration but a poor basis for a scheduled production job. A common path is to explore in a notebook, then move the settled logic into normal Python modules or SQL models with tests, and have the orchestrator run those. Some platforms do run notebooks on a schedule; if you do, apply the same care about versioning and errors.