Data Engineering › Storage, Formats & Lakehouse
JSON Lines
One JSON object per line, easy to stream and append.
Also known as: JSONL, NDJSON, newline-delimited JSON
JSON Lines (also called JSONL or NDJSON, for newline-delimited JSON) is a file format with one complete JSON object per line.
{"event": "signup", "user_id": 1, "ts": "2026-03-01T09:00:00Z"}
{"event": "login", "user_id": 1, "ts": "2026-03-01T09:05:12Z"}
{"event": "purchase", "user_id": 1, "amount": 19.99}
Compare with a normal JSON file holding an array: [ {...}, {...}, {...} ]. That file must be read as a whole before you know it is valid, and appending means editing the end of the file.
Why data teams like it
- Streams well. Read it one line at a time without loading it all into memory.
- Easy to append. New records are just new lines.
- Fault tolerant. A bad line affects only that record, not the whole file.
- Splittable. A big file can be cut at line boundaries and processed in parallel.
- Flexible. Each line can have different fields, which suits logs and events.
import json
with open("events.jsonl") as f:
for line in f:
record = json.loads(line)
print(record["event"])
Things to watch
- Newlines inside values must be escaped (
\n); each record has to stay on one physical line. Standard JSON serializers do this for you. - No schema. Types can drift between lines (
"amount": "19.99"vs19.99), and fields come and go. Validate on ingestion. See schema drift. - Bigger than binary formats. Field names repeat on each line and it compresses less efficiently than columnar formats. For analytics at scale, convert to Parquet. See row vs columnar formats.
- Compressing the files (e.g. gzip) saves space, at the cost of splitting.
It is a good “landing” format for raw event data.