Data Engineering › Batch & Distributed Processing
pandas
Python's standard DataFrame library for data analysis.
Also known as: pandas library, pandas DataFrame, pd
pandas is Python’s standard library for working with tabular data. Its central object is the DataFrame: a table with labelled columns, held in memory. It lets you load, filter, join, group and reshape data in a few lines.
import pandas as pd
orders = pd.read_csv("orders.csv") # load
big = orders[orders["total"] > 100] # filter rows
by_country = (
orders.groupby("country")["total"]
.sum()
.sort_values(ascending=False)
)
orders["created_at"] = pd.to_datetime(orders["created_at"])
merged = orders.merge(customers, on="customer_id", how="left") # join
merged.to_parquet("orders_enriched.parquet") # save
It is a common choice for exploration in a notebook, quick data cleaning, and small to medium pipeline steps.
The main limit: memory
pandas loads the whole table into RAM, and operations can create copies, so a file of a few gigabytes may need many times that in memory. When data outgrows one machine’s memory, options are loading only the needed columns, processing in chunks, switching to a single-node engine built for larger-than-memory work, or a distributed tool like Spark.
Classic mistakes
- Looping over rows with
foror.applywhere a vectorized column operation would be much faster. - Chained assignment confusion. Changing a slice may or may not change the original DataFrame; use
.loc[...]and.copy()deliberately. Behavior has been changing between versions, so check the docs for yours. - Silent type problems. A column of numbers with one bad value becomes text. Check
df.dtypes. - Missing values. Missing numbers show up as
NaN, and comparisons with them behave unexpectedly. See null. - Using it as a database. For joins and aggregations over large tables, push the work to SQL in the warehouse.