โ† Back to all articles
Systems ยท Database ยท Serverless
๐Ÿ˜
NEON ยท SERVERLESS POSTGRES

Neon Serverless Postgres Architecture:
Storage-Compute Split, Branching, and Scale-to-Zero

On Neon the query engine is not glued to a disk. Compute is an ephemeral Postgres VM. Data lives in a history-preserving object store. That split is why branching and scale-to-zero are cheap instead of marketing.

๐Ÿ“… August 11, 2026 โœ๏ธ Barnabas Waweru โฑ 14 min read ๐Ÿท Systems ยท Neon ยท Postgres ยท Serverless
ansi ยท wordmark ยท neon storage
โ–ˆโ–ˆโ–ˆโ•—   โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ–ˆโ•—   โ–ˆโ–ˆโ•—
โ–ˆโ–ˆโ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ•โ•โ•โ•โ•โ–ˆโ–ˆโ•”โ•โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ•‘
โ–ˆโ–ˆโ•”โ–ˆโ–ˆโ•— โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ–ˆโ–ˆโ•— โ–ˆโ–ˆโ•‘
โ–ˆโ–ˆโ•‘โ•šโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ•โ•โ•  โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘โ•šโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•‘
โ–ˆโ–ˆโ•‘ โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•”โ•โ–ˆโ–ˆโ•‘ โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ•‘
โ•šโ•โ•  โ•šโ•โ•โ•โ•โ•šโ•โ•โ•โ•โ•โ•โ• โ•šโ•โ•โ•โ•โ•โ• โ•šโ•โ•  โ•šโ•โ•โ•โ•

โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—
โ–ˆโ–ˆโ•”โ•โ•โ•โ•โ•โ•šโ•โ•โ–ˆโ–ˆโ•”โ•โ•โ•โ–ˆโ–ˆโ•”โ•โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•”โ•โ•โ•โ•โ• โ–ˆโ–ˆโ•”โ•โ•โ•โ•โ•
โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—   โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•”โ•โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—  
โ•šโ•โ•โ•โ•โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ•โ•โ•  
โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘   โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•”โ•โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ•‘โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•”โ•โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—
โ•šโ•โ•โ•โ•โ•โ•โ•   โ•šโ•โ•    โ•šโ•โ•โ•โ•โ•โ• โ•šโ•โ•  โ•šโ•โ•โ•šโ•โ•  โ•šโ•โ• โ•šโ•โ•โ•โ•โ•โ• โ•šโ•โ•โ•โ•โ•โ•โ•

Postgres Is Not Your Database

The Mental Model You've Been Missing

Traditional Postgres thinking: you rent a VM, Postgres owns the disk, connections arrive on port 5432, the WAL is written locally. That model is twenty years old, and it's the reason your dev database costs $40/month to sit idle, your CI pipeline has to hydrate a fresh database from scratch on every run, and you lose data the moment your disk fills up.

Neon's answer is architectural, not cosmetic. It takes the Postgres query engine and removes it from the storage layer entirely. Compute and storage become independently scalable, separately billed, separately operated resources. The query engine runs in ephemeral VMs, computes, and the data lives in a distributed, object-storage-backed system called Lakebase (Neon's internal storage engine, based on the open-source Neon Storage project).

This separation is why Neon can offer things that were architecturally impossible on traditional Postgres hosts: copy-on-write database branches, scale-to-zero with millisecond wakeup, read replicas that cost no extra storage, and point-in-time restore without a full backup.

KEY DISTINCTION

On RDS or DigitalOcean, deleting your dev instance frees your CPU, but the EBS volume is still there, still billed. On Neon, scale-to-zero suspends the compute while your data persists in shared storage. You pay for active CU-seconds, not idle machines. For dev databases that run 2 hours a day, that is a 10ร— cost reduction by default.

โˆž Branching
Copy-on-write database clones in milliseconds. Zero data copying, zero load on parent.
โšก Scale-to-Zero
Compute suspends after 5 min idle. Resumes in hundreds of milliseconds on next query.
๐Ÿ”Œ Serverless Driver
HTTP transport for edge functions; WebSocket for transactions. Drop-in for standard pg.
๐Ÿ“– Read Replicas
Independent read computes on shared storage. No data replication. Spin up in seconds.
๐Ÿ• Point-in-Time
Restore to any second within your history window. No separate backup storage required.
๐Ÿงฒ pgvector Native
HNSW and IVFFlat indexes baked in. halfvec for 3072-dim embeddings. No separate vector store.

The Engine Room: Storage-Compute Separation

How Lakebase Works Under the Hood

When you send a query to a Neon endpoint, a Compute, a lightweight Postgres process, receives it. The compute runs your SQL just like standard Postgres: it parses, plans, and executes. But here's where the architecture diverges: instead of reading pages from a local disk, it fetches them from the Neon Storage Service (Lakebase) over the network.

Lakebase is a distributed, WAL-based object store. Every write your Postgres compute makes becomes a WAL record shipped to Lakebase's Safekeeper nodes, a quorum-replicated write-ahead log layer. Once durably written, the WAL is processed by Pageserver nodes, which materialise page versions and serve them to computes on demand.

The critical insight: the Pageserver maintains a full history of every page version since your project was created (bounded by your history window). This is what enables both branching and point-in-time restore, they are not separate backup mechanisms. They are natural consequences of a history-preserving storage layer.

NEON ARCHITECTURE DIAGRAM
  CLIENT (Edge Function / Vercel / Cloudflare Worker)
       โ”‚
       โ”‚  HTTP  (neon() serverless driver, one-shot queries)
       โ”‚  WS    (Pool/Client, transactions, session state)
       โ”‚
  โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  โ”‚  NEON PROXY (SNI routing โ†’ project endpoint)         โ”‚
  โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”‚
  โ”‚  โ”‚ Connection Poolโ”‚   โ”‚  PgBouncer  (transaction  โ”‚   โ”‚
  โ”‚  โ”‚ (10k max conns)โ”‚   โ”‚  mode, -pooler hostname)  โ”‚   โ”‚
  โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ”‚
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
             โ”‚                        โ”‚
  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  โ”‚  COMPUTE  (ephemeral Postgres VM, scales to zero)     โ”‚
  โ”‚                                                         โ”‚
  โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”‚
  โ”‚  โ”‚  Standard Postgres Query Engine                  โ”‚   โ”‚
  โ”‚  โ”‚  Parser โ†’ Planner โ†’ Executor โ†’ Buffer Cache      โ”‚   โ”‚
  โ”‚  โ”‚                          โ”‚ page miss             โ”‚   โ”‚
  โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ”‚
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ โ”˜
                                โ”‚  page request (NRL protocol)
  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  โ”‚  LAKEBASE STORAGE (Neon Storage, distributed, durable)โ”‚
  โ”‚                                                         โ”‚
  โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”‚
  โ”‚  โ”‚  SAFEKEEPER nodes โ”‚   โ”‚  PAGESERVER nodes         โ”‚  โ”‚
  โ”‚  โ”‚  (WAL quorum,     โ”‚   โ”‚  (materialise page vers., โ”‚  โ”‚
  โ”‚  โ”‚   3-way repl.)    โ”‚โ”€โ”€โ–บโ”‚   serve on demand)        โ”‚  โ”‚
  โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ”‚
  โ”‚                                    โ”‚                      โ”‚
  โ”‚                          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”‚
  โ”‚                          โ”‚  S3-compatible object store โ”‚  โ”‚
  โ”‚                          โ”‚  (immutable page layers)    โ”‚  โ”‚
  โ”‚                          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ”‚
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
    

Safekeeper: The Write-Ahead Log Quorum

When your Postgres compute commits a transaction, it ships the WAL record to a set of Safekeeper nodes, typically three, geographically distributed within the region. A write is acknowledged only when a quorum has confirmed receipt. This gives Neon the same durability guarantee as synchronous streaming replication without requiring a standby Postgres process.

Once durably written, the WAL flows to the Pageserver. The Pageserver reads page deltas, applies them to the base page versions, and caches the results in a tiered store: hot pages in memory, warm pages in fast SSD, cold pages in S3-compatible object storage. This tiered materialization is what keeps cold-start latency low, the Pageserver can reconstruct the latest page version without replaying the entire WAL.

Pageserver: History-Preserving Storage

The Pageserver does not overwrite old page versions. Instead, it stores deltas, the difference between consecutive page states at specific LSN (Log Sequence Number) positions. To serve a page at LSN N, the Pageserver finds the nearest base image and applies deltas forward to reach N.

This immutable-append design is what makes point-in-time restore a first-class primitive rather than an afterthought. And it is exactly what makes branching possible: to branch at LSN N, you simply tell the Pageserver to serve page requests for the new branch using the existing page history up to N. No data is copied. The branch's writes go into a new delta chain. The parent never sees them.

The Branch Tree: Git for Your Database

What a Branch Actually Is

A Neon branch is a named pointer to an LSN in the storage layer. When you create a branch, Neon records the parent branch, the fork point LSN, and allocates a new endpoint, a compute, to serve queries against the branch. The endpoint gets its own connection string, its own scale-to-zero lifecycle, and its own compute size allocation.

The branch shares all page history below the fork LSN with its parent. Writes to the branch become new deltas in the branch's own timeline. The parent branch never sees branch writes, and the branch never sees post-fork writes to the parent unless you explicitly merge or rebase (via the Neon CLI's restore command).

Branch creation is instantaneous regardless of database size. Whether your database is 100 MB or 100 GB, the branch is available within a few seconds, almost all of that is compute spin-up, not data copying.

IMPORTANT: Branch vs Compute

A branch is a storage concept. A compute is a runtime concept. You can have a branch with no running compute (it will start on demand). You can have multiple computes on the same branch. Read replicas are exactly this: additional read-only computes on the same branch. Do not conflate "branch" with "database instance", Neon separates them deliberately.

Branch-per-Preview: The CI/CD Pattern

01
PR Opened
GitHub Action triggers neondatabase/create-branch-action@v6 with the PR number as branch name.
02
Branch Created
Neon forks from main at the current LSN. Action outputs db_url_with_pooler.
03
Preview Deploy
Vercel action receives DATABASE_URL pointing to the PR branch. Preview runs against isolated data.
04
PR Merged / Closed
neondatabase/delete-branch-action@v3 removes the branch. Storage freed; compute was already idle.
GITHUB ACTIONS, BRANCH PER PR
# .github/workflows/preview.yml
on:
  pull_request:
    types: [opened, synchronize, reopened, closed]

jobs:
  preview:
    runs-on: ubuntu-latest
    steps:
      - uses: neondatabase/create-branch-action@v6
        id: neon
        if: github.event.action != 'closed'
        with:
          project_id: ${{ secrets.NEON_PROJECT_ID }}
          api_key:    ${{ secrets.NEON_API_KEY }}
          branch_name: preview/pr-${{ github.event.pull_request.number }}

      - uses: vercel/action@v1
        if: github.event.action != 'closed'
        with:
          env: DATABASE_URL=${{ steps.neon.outputs.db_url_with_pooler }}

      - uses: neondatabase/delete-branch-action@v3
        if: github.event.action == 'closed'
        with:
          project_id: ${{ secrets.NEON_PROJECT_ID }}
          api_key:    ${{ secrets.NEON_API_KEY }}
          branch: preview/pr-${{ github.event.pull_request.number }}

Schema-Only Branches & Branch TTL

When working with sensitive production data, Neon supports schema-only branching, create a branch that carries the schema but not the row data. Useful for testing migrations on a realistic structure without exposing PII to developer machines.

Branches can also carry a TTL (expiration date). Set an expiration at branch creation and Neon will automatically delete the branch after that date, no cleanup cron required. Ideal for ephemeral feature branches and CI environments with known lifespans.

The Sleeping Giant: Scale-to-Zero Compute

Suspend and Resume Mechanics

After 5 minutes of inactivity (configurable on paid plans, fixed on Free), the Neon control plane suspends the compute. "Suspend" means the Postgres process exits, the VM is deallocated, and billing stops. The data, safely in Lakebase, persists indefinitely.

When the next query arrives via the Neon Proxy, the proxy detects the suspended state, signals the control plane to start a compute, waits for readiness, and then routes the connection. The client sees a connection delay of a few hundred milliseconds on the first query, not seconds, not a timeout. The Pageserver's page cache is already warm from before the suspend, so the reactivated compute can serve data immediately.

Note: Computes larger than 16 CU are exempt from scale-to-zero and run always-active. Very large computes need stable performance characteristics that aren't compatible with the suspend/resume cycle.

COMPUTE LIFECYCLE
  ACTIVE โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–บ IDLE (5 min no queries)
     โ”‚                                        โ”‚
     โ”‚  queries arrive                        โ”‚  control plane signals suspend
     โ”‚  billed per CU-second                  โ”‚
     โ”‚                                        โ–ผ
     โ”‚                                   SUSPENDED
     โ”‚                                   (compute deallocated,
     โ”‚                                    data in Lakebase,
     โ”‚                                    billing stopped)
     โ”‚                                        โ”‚
     โ”‚  next query arrives                    โ”‚  proxy detects suspend
     โ”‚  โ—„โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
     โ”‚
     โ”œโ”€โ–บ control plane provisions compute VM
     โ”œโ”€โ–บ Postgres starts, connects to Pageserver
     โ”œโ”€โ–บ Proxy routes connection (few hundred ms delay)
     โ””โ”€โ–บ ACTIVE, queries served normally
    
COLD START IN PRODUCTION

For production traffic, cold-start latency is real. Handle it at the application layer, not with hope. Two patterns:

  • Connection string option: append ?connect_timeout=10 to give the Postgres client enough time to wait for the compute to warm.
  • Retry at application layer: catch Prisma P1001 / PG connection errors on first attempt, retry once with 1-2s backoff. Almost always succeeds.

Alternatively, on paid plans, disable scale-to-zero for production endpoints and enable it only for dev/test branches.

Autoscaling: The Other Half of the Picture

Scale-to-zero handles the lower bound (idle โ†’ zero). Autoscaling handles the upper bound. You set a min and max CU range per compute. Neon monitors CPU and memory pressure and scales within the range without restart, this is live compute resize, not instance replacement.

This combination means your database automatically tracks your workload: zero CUs at night, 0.25 CU for light morning traffic, scaling to 4 CU during a traffic spike, then back down. You pay only for what the workload actually consumed, measured in CU-seconds.

The Connection Layer: PgBouncer and the Serverless Driver

The 10,000-Connection Problem

Standard Postgres creates a new OS process for each connection. On a 0.25 CU compute (1 GB RAM), max_connections is 104. Subtract 7 reserved for Neon's superuser: you have 97. That's fine for a monolith, disastrous for serverless functions where every request may spawn a new connection.

Neon bundles PgBouncer at every endpoint. Enable it by switching to the pooled connection string, identical to the standard one except the hostname includes -pooler:

# Standard (unpooled, for migrations, long-lived connections)
postgresql://user:[email protected]/neondb?sslmode=require

# Pooled (PgBouncer, for serverless, Vercel Functions, Edge)
postgresql://user:[email protected]/neondb?sslmode=require
    

PgBouncer runs in transaction mode: it holds a real Postgres connection only for the duration of a transaction, then returns it to the pool. Up to 10,000 client connections can share the underlying pool.

POOLED VS UNPOOLED, DECISION TABLE
Scenario Use Why
Vercel / Netlify Functions Pooled Connection-per-request; pooler prevents max_connections exhaustion
Cloudflare Workers / Edge HTTP (neon() driver) No TCP; HTTP fetch is the only viable transport at the edge
Prisma / Drizzle migrations Unpooled DDL holds session state; transaction-mode pooler drops SET/PREPARE
pg_dump / logical replication Unpooled Requires persistent session state (LISTEN/NOTIFY, replication slots)
Multi-statement transactions Pooled or Unpooled Pooled works in transaction mode; avoid SET/PREPARE within transaction
Long-lived Node.js server Unpooled + own pool Manage your own pg.Pool for session affinity and full Postgres features

The Serverless Driver: HTTP vs WebSocket Transport

@neondatabase/serverless is a drop-in replacement for the pg npm package that adds two new transports: HTTP and WebSocket. The standard pg package requires TCP, which is unavailable in most edge runtimes. The serverless driver solves this without changing your query syntax.

neon(), HTTP Transport
  • Single request per query (or batch)
  • Lowest cold-start latency
  • No transactions (single-statement only)
  • Template literal = auto-parameterization
  • Best for: Route Handlers, Server Actions
Pool / Client, WebSocket Transport
  • Full Postgres protocol over WS
  • Multi-statement transactions (BEGIN/COMMIT)
  • Session state, prepared statements
  • node-postgres compatible API
  • Best for: complex transactional logic
DRIVER PATTERNS, NEXT.JS APP ROUTER
// lib/db.ts, HTTP transport (default: use for everything without transactions)
import { neon } from '@neondatabase/serverless'
const sql = neon(process.env.DATABASE_URL!)
export { sql }

// Route Handler: one-shot query, safe, parameterized
const rows = await sql`SELECT id, email FROM users WHERE id = ${userId}`

// Multiple non-interactive queries in one round-trip (transaction: false)
const [users, posts] = await sql.transaction([
  sql`SELECT * FROM users LIMIT 10`,
  sql`SELECT * FROM posts LIMIT 10`
], { isolationLevel: 'ReadCommitted' })

// โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
// lib/db-pool.ts, WebSocket transport (transactions only)
import { Pool } from '@neondatabase/serverless'

export async function POST(req: Request, ctx: { waitUntil: (p: Promise) => void }) {
  const pool = new Pool({ connectionString: process.env.DATABASE_URL })
  const client = await pool.connect()
  try {
    await client.query('BEGIN')
    await client.query('UPDATE accounts SET bal = bal - $1 WHERE id = $2', [amount, from])
    await client.query('UPDATE accounts SET bal = bal + $1 WHERE id = $2', [amount, to])
    await client.query('COMMIT')
    return Response.json({ ok: true })
  } catch (err) {
    await client.query('ROLLBACK')
    throw err
  } finally {
    client.release()
    ctx.waitUntil(pool.end()) // close after response, non-blocking
  }
}

The Read Farm: Replicas Without Replication

Why Neon Read Replicas Are Different

Traditional read replicas involve streaming replication: WAL records flow from primary to standby, which applies them to a full copy of the data. You pay for double the storage, and replica lag is a function of network latency and write volume.

Neon read replicas work differently. Since all computes, primary and replicas, read from the same Lakebase Pageserver, there is no data to replicate. A read replica is just another compute pointed at the same storage branch. You pay only for the additional compute-hours the replica consumes.

Replica lag is architectural: the primary writes WAL to Safekeepers, which flow to the Pageserver. Replicas read from the Pageserver. The lag is the Pageserver's materialization latency, typically milliseconds, not seconds. And because replicas share the Pageserver's page cache, they benefit from pages already warmed by the primary.

READ REPLICA TOPOLOGY
  WRITE TRAFFIC                     READ TRAFFIC
       โ”‚                                 โ”‚
  โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”           โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  โ”‚ PRIMARY COMPUTE   โ”‚           โ”‚ READ REPLICA #1      โ”‚
  โ”‚ (read-write)      โ”‚           โ”‚ (read-only compute)  โ”‚
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜           โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
           โ”‚ WAL                             โ”‚ page requests
           โ–ผ                                 โ”‚
  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  โ”‚                  LAKEBASE PAGESERVER                      โ”‚
  โ”‚           (single source of truth for all pages)          โ”‚
  โ”‚                                                           โ”‚
  โ”‚  Pages at LSN N are identical whether read by the         โ”‚
  โ”‚  primary or any replica. No physical replication needed.  โ”‚
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
    
โœ“ No Extra Storage Cost
Data is not duplicated. Replicas share pages with the primary. You pay only for compute time.
โœ“ Instant Provisioning
Spin up a read replica in seconds. No data copying, no initial sync lag, no catch-up period.
โœ“ Autoscaling + Scale-to-Zero
Each replica has its own CU range and scale-to-zero setting. Analytics replica can stay always-on; dev replica scales to zero.
โœ“ Independent Compute Size
Run analytics on a 4 CU replica without affecting the 0.25 CU primary serving your app's writes.

Cost Reality Check

NEON FREE TIER, WHAT YOU ACTUALLY GET

Neon's Free tier is genuinely useful for development, but understand its limits before you hit them in production.

  • Compute: 0.25 CU (1 GB RAM); always scale-to-zero (cannot disable); 5 min idle timeout fixed
  • Storage: 512 MB included; 0.5 GiB additional = $0.10/GiB-month after that
  • Branches: max 10 branches per project on Free
  • History window: 6 hours (paid plans: 1 day default, up to 30 days on Scale)
  • Projects: 1 project on Free (10 on Launch, unlimited on Scale)
PLAN COMPARISON, AUG 2026
Feature Free Launch ($19/mo) Scale ($69/mo)
Compute included 0.25 CU (fixed) 300 CU-hr/mo 750 CU-hr/mo
Storage included 0.5 GiB 10 GiB 50 GiB
Branches (max) 10 Unlimited Unlimited
Scale-to-zero Forced (cannot disable) Optional Optional
History window 6 hours 7 days 30 days
Max compute size 0.25 CU 8 CU 56 CU

The Real Cost Drivers

Compute: You pay for active CU-seconds. A 1 CU compute running 8 hours/day at $0.16/CU-hr costs ~$39/month, less than most managed Postgres services that run idle 24/7. With scale-to-zero, a dev database that runs 2h/day costs under $5/month.

Storage: Billed per GiB-month of data stored. This includes your current data and the WAL history within your window. Long history windows increase storage cost, a database with heavy writes and a 30-day history window stores a lot of deltas. Monitor storage on the Neon console.

Branches: Branches themselves are free, they share storage with the parent. You pay for compute on branch endpoints if they're running. Delete PR branches when PRs close; an open PR branch with an idle endpoint is just storage, no compute cost.

Data transfer: Egress from Neon to your application is billed. Edge-native deployments (Vercel Edge, Cloudflare Workers) that co-locate with Neon's AWS regions benefit from lower latency and lower egress costs.

COST OPTIMIZATION CHECKLIST
  • Delete PR branches when PRs close, use the GitHub Action delete hook
  • Set branch TTL for known-lifespan test environments
  • Use scale-to-zero on dev and staging; disable for production on paid plans
  • Set autoscaling min to 0.25 CU and max to what your peak requires, don't provision for peak always
  • Use the HTTP transport (neon()) for read-heavy paths, avoids WS connection overhead
  • Deploy in the same AWS region as your Neon project to minimize data transfer costs

Key Takeaways

1
Storage-compute separation is the root of everything. Neon's architecture is not a configuration option, it's a fundamentally different model. Every feature (branching, scale-to-zero, read replicas, PITR) flows from separating the query engine from the history-preserving storage layer. If you internalize this, Neon's behavior stops being surprising.
2
Use pooled connection strings for serverless; unpooled for migrations. This is the most common source of breakage. PgBouncer in transaction mode drops session-level state. DDL migrations that hold session state will silently fail or behave incorrectly through the pooler. Always set Prisma's directUrl or Drizzle's migration config to the unpooled URL.
3
neon() for reads, Pool for transactions. The HTTP transport has the lowest overhead and is edge-compatible. Use it by default in Next.js Route Handlers and Server Actions. Only reach for the WebSocket Pool when you genuinely need multi-statement transactions, and remember to close the Pool with ctx.waitUntil(pool.end()).
4
Branch cost โ‰  storage cost. Branches share parent storage via copy-on-write. A branch is essentially free until you write to it. The storage cost is the write delta accumulated since the fork. The compute cost is only incurred while a branch endpoint is actively running. Delete idle PR branches to keep costs clean.
5
Cold starts require explicit handling, not hope. Scale-to-zero is a feature, not a defect, but cold-start latency is real. Add ?connect_timeout=10 to your connection string and implement a single retry with short backoff for Prisma P1001 / PG ETIMEDOUT errors. On paid plans, disable scale-to-zero on your production compute if sub-100ms p99 is a requirement.
6
pgvector is a first-class citizen, with a dimension caveat. Neon supports pgvector HNSW and IVFFlat indexing natively. HNSW supports up to 2,000 dimensions. For OpenAI's text-embedding-3-large (3072 dimensions), use the halfvec type, same semantics, half the storage, within the HNSW ceiling.