Contents

Data Engineering › Data Engineering Foundations

Volume, Velocity, Variety

The three dimensions that make data "big", and which one is actually your problem.

Also known as: 3 Vs, three Vs of big data, big data 3 Vs

Volume, velocity and variety are the three classic dimensions that make data “big”:

QuestionExamples of “hard”
VolumeHow much data?Terabytes to petabytes, billions of rows
VelocityHow fast does it arrive and need processing?Thousands of events per second; answers needed in seconds
VarietyHow many shapes and sources?Tables, JSON, logs, images, text, many systems

They matter because each one pushes you toward a different solution.

Which one is actually your problem?

This is the useful question. Most teams have one problem, or none. The classic mistake is adopting a big, complex distributed system “because we do big data” when the data would fit on one machine.

A rough check: a modern laptop or single server can process datasets of many gigabytes, sometimes more, with tools like pandas or a single-node engine, far faster than a cluster for the same job. Reach for distributed tools when you have measured that you’ve outgrown one machine.

Variety is the easiest to underestimate: ten source systems with ten definitions of “customer” can be harder than a billion clean rows.

Some people add more Vs (veracity, value). The first three are the common core; the extra ones are less standardized.