Database Operations
Running databases in production: access, backups, replication and upgrades.
Backend Engineer track
Junior
Write correct code, ship small changes safely, ask good questions.
- BackupsCopies of data for restoring, and why untested backups don't count.
Mid-level
Own a feature end to end without hand-holding.
Core: start here
- Database ReplicationCopying data to other servers for availability and read scaling.
- Read ReplicaA copy of the database for serving reads.
7 more mid-level concepts
- Bulk Loading (COPY)Loading large amounts of data much faster than row-by-row inserts.
- Database Authentication and TLSHow clients prove who they are to the database, and encrypting that connection.
- Database Connection PoolReusing database connections, with tools like PgBouncer at scale.
- Database Version UpgradesMoving to a new major version with little downtime.
- Logical vs Physical BackupsSQL dumps vs copies of the data files.
- Restore TestingActually restoring backups regularly to prove they work.
- Roles and Privileges (GRANT, REVOKE)Controlling who can read and change which data.
Senior
Own a system, its failure modes, and its trade-offs.
Core: start here
- ShardingSplitting data across databases by key.
13 more senior concepts
- Database ExtensionsAdding capabilities to a database, like PostGIS or pgvector.
- Database High AvailabilityAutomatic failover for databases, with tools like Patroni.
- Key-Range vs Hash PartitioningSplitting data by ranges of keys vs by hashed keys.
- Last Write WinsResolving conflicting writes by timestamp, and the data it silently loses.
- Leaderless ReplicationAny node accepts writes, with quorums to reconcile, as in Cassandra.
- PartitioningDividing data into parts, within one machine or across many.
- Point-in-Time RecoveryRestoring a database to any moment using a base backup plus the log.
- Query Statistics (pg_stat_statements)Finding which queries use the most time in aggregate.
- Row-Level SecurityThe database itself filtering rows per user or tenant.
- Streaming vs Logical ReplicationCopying the raw log vs copying row changes, and what each lets you do.
- Synchronous vs Asynchronous ReplicationWaiting for replicas to confirm vs not, trading safety for latency.
- Table PartitioningSplitting one huge table into smaller physical pieces, e.g. by month.
- Using the Database as a QueueJob queues in SQL with SKIP LOCKED, and when that's enough.
Staff
Shape how many teams build, across systems.
Nothing here yet.
Principal
Set technical direction for the organization.
Nothing here yet.
Data Engineer track
Junior
Build and fix pipelines from clear specs; write correct SQL.
Core: start here
- PartitioningDividing data into parts, within one machine or across many.
2 more junior concepts
- BackupsCopies of data for restoring, and why untested backups don't count.
- Bulk Loading (COPY)Loading large amounts of data much faster than row-by-row inserts.
Mid-level
Own pipelines and models end to end, including their quality.
- Database Authentication and TLSHow clients prove who they are to the database, and encrypting that connection.
- Database ReplicationCopying data to other servers for availability and read scaling.
- Database Version UpgradesMoving to a new major version with little downtime.
- Key-Range vs Hash PartitioningSplitting data by ranges of keys vs by hashed keys.
- Logical vs Physical BackupsSQL dumps vs copies of the data files.
- Read ReplicaA copy of the database for serving reads.
- Restore TestingActually restoring backups regularly to prove they work.
- Roles and Privileges (GRANT, REVOKE)Controlling who can read and change which data.
- Row-Level SecurityThe database itself filtering rows per user or tenant.
- ShardingSplitting data across databases by key.
- Table PartitioningSplitting one huge table into smaller physical pieces, e.g. by month.
Senior
Design the platform's storage, processing and modeling choices.
- Database ExtensionsAdding capabilities to a database, like PostGIS or pgvector.
- Database High AvailabilityAutomatic failover for databases, with tools like Patroni.
- Last Write WinsResolving conflicting writes by timestamp, and the data it silently loses.
- Leaderless ReplicationAny node accepts writes, with quorums to reconcile, as in Cassandra.
- Point-in-Time RecoveryRestoring a database to any moment using a base backup plus the log.
- Query Statistics (pg_stat_statements)Finding which queries use the most time in aggregate.
- Streaming vs Logical ReplicationCopying the raw log vs copying row changes, and what each lets you do.
- Synchronous vs Asynchronous ReplicationWaiting for replicas to confirm vs not, trading safety for latency.
- Using the Database as a QueueJob queues in SQL with SKIP LOCKED, and when that's enough.
Staff
Shape how the whole organization produces and uses data.
Nothing here yet.
Principal
Set data strategy and architecture across the company.
Nothing here yet.
Frontend Engineer track
Junior
Build UI that works, ship small changes safely, ask good questions.
- BackupsCopies of data for restoring, and why untested backups don't count.
Mid-level
Own a feature end to end without hand-holding.
Nothing here yet.
Senior
Own an app's architecture, performance, and failure modes.
Nothing here yet.
Staff
Shape how many teams build, across apps.
Nothing here yet.
Principal
Set technical direction for the organization.
Nothing here yet.