LakeBase Didn't Reinvent Postgres. It Moved the fsync.

databases postgres architecture mechanism cloud

Databricks shipped LakeBase Postgres: real Postgres, wire compatible, standard drivers, ORMs, psql, pgvector and PostGIS included. Generally available on AWS since January 22, 2026, Azure GA in March, adoption reportedly growing at twice the rate of their warehousing product. It is not a reimplementation; the query engine is unmodified Postgres. The thing that changed is the thing underneath the commit. A transaction no longer becomes durable when a local SSD flushes. It becomes durable when a quorum of storage nodes agrees it happened. The marketing headline is database branching, “Git for your database.” The headline is the side effect.

What LakeBase is

The product facts, kept short. Databricks acquired Neon in May 2025; LakeBase is the productized result. Serverless Postgres with scale-to-zero, Postgres 18 with pgvector today, an SLA since July 2026, SOC 2 Type 2, PCI-DSS and HITRUST landed this month, storage quota default of 32 TB, Lakehouse Sync pushing operational data into Delta tables without pipelines. PostGIS, extensions, the tooling you already use. None of that is the interesting part.

The mechanism

The architecture splits Postgres into two layers connected by a WAL stream.

Compute is stateless Postgres. It parses, plans, executes, and holds only transient state: shared buffers in RAM plus a local NVMe cache. It owns no durable data. That is why it can restart, autoscale, or scale to zero without a recovery story; there is nothing on it to recover.

The WAL is pulled out of the machine into a group of safekeepers. A transaction commits when a quorum acknowledges the log record via Paxos consensus, not when a local fsync returns. This is the sentence everything else falls out of: the commit is a distributed-systems event.

Pageservers hold the data files, rebuilt from the WAL. They materialize page versions on demand at any requested LSN, acting as a write-through cache above object storage. Object storage (S3 on AWS) holds the immutable page history and sits off the hot read path; only pageservers read from it. The read path is buffer pool, local NVMe, pageserver, object storage, stopping at the first cache hit.

You traded the fsync latency of a local SSD for the consensus latency of a replicated log. Databricks argues this is a wash for anyone already running synchronous replication: one network round trip replacing another rather than adding one. Fair, but note what changed. The durability of your commit now depends on the health of a distributed quorum, and during a partition the system must choose between blocking writes and risking a durability gap. A monolith never had to make that choice.

Why branching and scale-to-zero are consequences, not features

Here is the part worth internalising. Four bullet points on the LakeBase landing page are one architectural choice wearing four hats.

A branch is a metadata operation pointing at an LSN. No data is copied. A 10 GB database and a 2 TB database branch in the same number of seconds, and you pay only for the pages the branch modifies; touch 1 GB in a 100 GB database and the branch costs about 1 GB extra.

Scale-to-zero is the same sentence read from the other side. Compute owns nothing durable, so it can evaporate. Point-in-time recovery is compute attaching to a past LSN, not restoring data. A read replica is a second compute attaching to existing storage, no data movement. Four features, one cause.

The tradeoff ledger

Externalized durability is a trade, and the ledger has two columns.

PropertyMonolithic PostgresLakeBase
Commit acklocal fsyncquorum consensus over the network
Cold page readbase image on diskreplay WAL deltas forward to the LSN
Database branchpg_dump, minutes to hoursmetadata operation, seconds
Idle computebilled around the clock$0
Superuser, tablespaces, logical replicationavailablenot available

The cold-page row deserves respect. Resolving a page’s current version can mean replaying a long chain of WAL deltas over a base image, and if compaction and image generation are mistuned, you get tail latency that does not show up in a demo and does show up at 2 a.m. Databricks’s own engineering confirms this was real: they disabled Full Page Writes and pushed image generation into the pageserver, reporting 94% WAL traffic reduction and up to 5× write throughput. You do not spend that engineering effort on a problem that does not exist.

The rest of the ledger, from third-party analyses. Branching has no merge semantics: a branch is a snapshot at an LSN and never resyncs with its parent, so the Git mental model breaks exactly where Git is most useful, at merge time. Long-lived branches pin page history and quietly inflate storage; immutable storage is cheap to write and expensive to forget. Connections are ephemeral in ways that surprise session state: a 24-hour idle timeout, a 3-day maximum lifetime, a roughly 500 ms cold resume, and advisory locks, temp tables, and prepared statements dying with the connection. High-availability configurations cannot scale to zero, so production compute bills around the clock. And governance is aligned, not merged: Unity Catalog governs the lakehouse query surfaces while Postgres roles govern direct connections, two authorization paths to keep consistent.

None of these are fatal. All of them are the price.

Where this generalises

The transferable test: when a database vendor lists features, group them by the architectural choice that implies them, then price the choice, not the list.

LakeBase’s list, branching, scale-to-zero, PITR, instant replicas, sync tables into Delta, one storage foundation serving OLTP and analytics, all falls out of “durability lives in a distributed log over object storage.” The costs, consensus in the commit path, cold-page tail management, GC discipline, session ephemerality, fall out of the same sentence. You cannot take the features and decline the costs; they are the same decision.

This applies to every externalized-durability system you will evaluate from here on: Aurora’s separated storage layer, Neon itself, every disaggregated OLTP engine on the roadmap. Each one puts the same trade on the table in a different wrapper, and the evaluation is always the same two questions. Which features are consequences of the storage decision? What does the storage decision cost at 2 a.m. rather than in the demo?

The trap

The trap is to read the feature list as the product. Branching is not the product. A quorum-committed WAL in object storage is the product, and branching is one of the things that architecture can do that a monolith cannot.

The second trap is the demo-versus-production gap. The vendor-reported pgbench numbers, around 1,731 TPS for LakeBase against roughly 1,508 for Aurora PostgreSQL on the same 4M-row workload, are steady-state numbers. Steady state is where disaggregated architectures look best; the tails are where the ledger lives. Cold-page replay latency, quorum behavior under partition, GC pressure from a forgotten branch: none of it appears in a 90-second demo, all of it appears in a page.

The closing rule. When durability becomes a distributed-systems event, every consequence of that decision, on both sides of the ledger, looks like a feature until you know which sentence it fell out of. LakeBase’s sentence is “the commit lives in a quorum over object storage.” Read the sentence, and the bullets stop being marketing and become arithmetic.

$ cat GIT .md
· 7 min read

84 Repositories Vanished. The Fix Was mkdir.

S3 has no empty directories, and in a fully packed git repository the refs directories are exactly that: empty. A file-by-file sync carried every object intact and dropped the two directories git requires to call something a repository, so 84 of 517 came back unreadable while the database still said they had commits. The migration's health checks stayed green the whole time, because none of them ever open a repository.

git aws s3 gitlab mechanism
$ cat KUBERNETES .md
· 7 min read

The Constraint Was Satisfied. The Zone Was Empty Anyway.

A topology spread constraint is only ever a statement about the population its selector matches, and an operator-set label can quietly pool sibling Deployments into one shared count. Three collectors each satisfied their hard zone spread while the union of them left a zone empty, and the standard repair converged on the same wrong answer every time.

kubernetes scheduling opentelemetry finops mechanism
$ cat CLICKHOUSE .md
· 8 min read

136 Million PUTs for 17 GiB of Data

Object storage bills per operation, and a ClickHouse part on an S3 disk is not one object but one per column. So the cost of a cold tier is a function of how many parts exist, not how many bytes they hold, and every setting that starves merges becomes a line on the bill. Two chart defaults did exactly that, and the fix that stopped it had never been committed.

clickhouse s3 finops observability mechanism