Contents

Data Engineering › Data Engineering Foundations

Structured, Semi-Structured and Unstructured Data

Tables, JSON-like records, and free-form text, images or audio.

Also known as: structured data, semi-structured data, unstructured data, data structure types

Data comes in three broad shapes, based on how much structure it has.

TypeShapeExamples
StructuredFixed schema: rows and columns with defined typesSQL tables, spreadsheets, CSV with a stable header
Semi-structuredSelf-describing, with fields and nesting, but flexibleJSON, XML, Avro, log lines in a pattern, event payloads
UnstructuredNo predefined organizationFree text, emails, PDFs, images, audio, video
Structured:       id | name | total          (every row has the same columns)
Semi-structured:  {"id": 1, "items": [{"sku": "A1", "qty": 2}], "coupon": null}
Unstructured:     "Hi, my order never arrived and I'd like a refund..."

Why the difference matters

  • Structured data is easy to query, validate and join, because the schema is known up front (schema on write: you enforce the shape when you store it).
  • Semi-structured data is flexible. Different records can have different fields, and fields can be nested or repeated. It fits event payloads and API responses, but needs more care: fields may be missing or change type, and you often work out the schema when you read it (schema on read). Some warehouses have special types for querying JSON.
  • Unstructured data has to be processed before it can be analyzed: extracting text, running models over images or audio, creating embeddings. It’s usually stored as files in object storage with metadata on the side.

What to do about it

  • Land raw semi-structured data as it arrived, then parse it into structured tables in a transformation step. Keeping the raw copy lets you re-parse when you find a mistake.
  • Expect schemas to drift in semi-structured sources: new fields, renamed fields, changed types (schema drift).
  • Choose formats that carry types and schemas for storage between pipeline stages (columnar formats and Avro) over loose text where you can.
  • Most real systems mix all three: a support ticket has structured fields (id, status), semi-structured metadata and an unstructured message.

Most of the analytical value in a business is still in structured data, but unstructured content, such as text and documents, is increasingly a source for analysis and for AI applications. See data volume, velocity and variety.