Skip to main content
Subscribe
Backend Infrastructure

How OpenAI Scales PostgreSQL for 800 Million ChatGPT Users

How would you build a system that 800 million people use? ChatGPT is one, and the way OpenAI scales PostgreSQL for it is almost stubbornly simple. Its core data lives on one Postgres database: one writer, nearly 50 copies. OpenAI laid the whole thing out in January 2026 (OpenAI, Scaling PostgreSQL to power 800 million ChatGPT users, 2026).

This post builds that system up from nothing. It starts with ten people on one box and adds only the parts OpenAI actually runs, one break at a time, until it arrives at the machine in that document. I built this ladder for a Machine Room episode, and this is the written version.

One rule holds throughout. The rungs (10, 100, 10,000, 1,000,000 users) are my thought experiment. Every mechanism and every figure in the fixes is OpenAI’s, quoted from their post.

The episode’s narration is synthetic (an AI voice). The script and every figure in it are the ones in this post.

Ten people, one box (the message, slowly)

Diagram: ten users on the left send one message to a single box containing the app and a Postgres database on Microsoft Azure, with the OpenAI and PostgreSQL logos beside it

At the bottom of the ladder, everything sits on one box: the app and a Postgres database, on Azure, the same cloud OpenAI uses (its production primary is an Azure PostgreSQL flexible server, per OpenAI’s post). Ten people use it. Nothing is interesting yet, which is exactly why it’s the right place to slow down.

Follow one message. You type “hi” and press send. Before anything is saved, the app has questions. Who are you? What plan are you on? What are this chat’s settings? That’s three reads. Then it stores your message, which is one write.

Three reads and one write. Hold on to that ratio, because it decides everything that follows. OpenAI describes its own workloads as “primarily read-heavy”, and the whole design leans on that fact. Reads and writes will start behaving very differently once the crowd grows, and the fixes for each will look nothing alike.

With ten people, the box answers before anyone notices. There’s no queue and no waiting. A single Postgres instance is a remarkably capable machine, and most products never outgrow this rung.

At a hundred people, things get a little warmer. The right move here is the boring one, and it’s the one OpenAI says it made first. After launch, it optimised the application and database layers, then scaled up to a bigger instance, and only after that scaled out. Fewer wasted queries, a bigger box.

That order matters. It’s tempting to reach for distributed systems the moment a graph bends upward. OpenAI didn’t. It squeezed the single machine first.

But every box has a ceiling. You can buy a bigger one, and then a bigger one again, and eventually there isn’t a bigger one to buy. The rest of this post is about what happens when you hit those ceilings, one at a time, and which part of the system takes the hit.

Ten thousand, and the doors (PgBouncer)

Diagram: a wave of 10,000 requests reaches the box, where PgBouncer (the doorman) holds a few dozen open connections in front of a Postgres instance with a 5,000-connection limit, and one message's write passes through

Azure PostgreSQL caps each instance at 5,000 connections, and OpenAI has had “incidents caused by connection storms that exhausted all available connections” (OpenAI, 2026). At ten thousand users, the first thing to break isn’t the database’s thinking. It’s the doors.

Picture a launch. Ten thousand people press send in the same minute. Every one of those sends needs a connection to Postgres before it can run a single query. Opening a connection isn’t free; OpenAI measured the average at 50 milliseconds before its fix.

Five thousand get in. The rest wait, time out, and retry. The retries arrive on top of the next wave of sends, which makes the line longer, which causes more timeouts. The database wasn’t tired. It had plenty of capacity to answer queries. It simply ran out of doors.

This was the moment I found hardest to explain in the episode, because it feels wrong. The machine is idle and failing at the same time.

The fix is a doorman. OpenAI put PgBouncer in front of Postgres as a proxy layer, running in statement or transaction pooling mode. Instead of every request opening its own connection, PgBouncer keeps a small set of real connections open to the database. When a request needs one, it borrows a connection for a single transaction and hands it straight back.

Now walk the message through. Your send arrives at PgBouncer, not at Postgres. It gets a connection that’s already open, runs its three reads and one write, and returns the connection to the pool. The next person’s send picks up the same connection a moment later. Thousands of clients share a few real doors.

The payoff, in OpenAI’s own numbers: “average connection time dropped from 50 milliseconds (ms) to 5 ms.” The 5,000 limit stops being the thing that wakes people up at night.

Two details from the post are easy to skip and shouldn’t be. The proxy, the clients, and the replicas are co-located in each region, so the borrowed connection doesn’t cross an ocean. And idle timeouts are critical, because a client that grabs a connection and sits on it defeats the whole point of pooling. Each replica runs its own Kubernetes deployment with multiple PgBouncer pods behind one Service.

A million, and the reads (replicas)

Diagram: a wave of 1,000,000 requests passes PgBouncer; the Postgres primary (the writer) streams to nearly 50 read replicas (copies) arranged in four regions, with one message travelling through

OpenAI runs “nearly 50 read replicas spread over multiple regions globally” behind a single primary (OpenAI, 2026). At a million users, the doorman is fine. The database itself is the limit: one machine answering every query.

So make copies. A read replica holds the same data as the primary and can answer the same questions. Your three reads (who you are, your plan, this chat’s settings) don’t care which machine answers them, as long as the answer is right.

A copy can’t take a write, though. If two machines both accepted writes independently, you’d have two versions of the truth and no clean way to decide which one wins. So every send gets split. The one write goes to the primary, the writer. The reads go to the copies. OpenAI offloads reads to replicas “wherever possible”; some reads stay on the primary because they sit inside a write transaction.

How do the copies stay current? The writer streams its write-ahead log (WAL) to every replica, in order. Each replica replays the log and ends up with the same rows. If WAL is new to you, how WAL and transactions keep Postgres consistent covers the fundamentals.

Spreading those copies across regions means a read gets answered near the person asking. There’s a cost, though: lag. A copy may not have your newest message for a heartbeat after you send it. In a chat app, nobody notices.

There’s a ceiling here too, and OpenAI names it plainly. The primary streams WAL to every replica, so each extra copy adds network and CPU load on the writer. “We can’t keep adding replicas indefinitely.” The next rung is cascading replication, where intermediate replicas relay WAL to further replicas so the primary doesn’t ship to every one directly (PostgreSQL docs, Cascading Replication). OpenAI says that setup would let it “scale to potentially over a hundred replicas without overwhelming the primary”, and it’s still in testing.

One more detail. OpenAI runs multiple replicas per region, with headroom, so that losing one replica doesn’t turn into a regional outage.

A hundred million, and the same question (the cache and the miss storm)

Diagram: a wave of 100,000,000 requests hits a cache holding who-you-are, your-plan and this-chat's-settings entries in front of regional read replicas, with the Postgres primary (the writer) off to the side

OpenAI’s caching layer serves “most of the read traffic” before it ever reaches Postgres (OpenAI, 2026). That matters at this scale: during the ImageGen launch, more than 100 million new users signed up within a week. Even nearly 50 copies shouldn’t answer the same question millions of times.

And it is the same question. Your plan doesn’t change between one message and the next. Neither does who you are. Most reads ask something that was answered a second ago, so a cache sits in front of the copies and keeps those answers in memory.

Walk the message through. You send. The app asks who you are: cache hit. Your plan: cache hit. This chat’s settings: cache hit. Not one of the three reads reaches a replica. That’s the point of the layer.

Now the cache’s bad day. OpenAI’s post names “widespread cache misses from a caching-layer failure” as one way its outages start. Here’s my illustration of how that looks: a cache restart or a deploy empties the entries, and in the same second every read that used to hit the cache lands on the copies instead. The misses don’t disappear. They go straight to Postgres, all at once. That’s a cache-miss storm.

The storm is worst on shared keys. If thousands of requests miss on the same entry at the same moment, all of them go to the database for the identical row. The database does the same work thousands of times, and the cache gets refilled thousands of times with the same answer.

OpenAI’s guard is a per-key lease. In its words: “only a single reader that misses on a particular key fetches the data from PostgreSQL … All other requests wait for the cache to be updated.” One request goes to the database. Everyone else waits a moment and then reads the fresh entry from the cache.

It’s a small rule with a big effect. A storm that would have sent thousands of identical queries to the copies sends one.

Notice what all three fixes so far have in common. The doorman, the copies, and the cache each make reads cheaper or rarer. None of them helps with the write.

The writer (why OpenAI doesn’t shard Postgres)

Diagram: the Postgres primary (the writer) with a headroom gauge, the cache, nearly 50 copies, and Cosmos DB as the sharded store made of many slices, with heavy writes routed to the slices and one message kept on the writer

Every write in this system still lands on one machine. OpenAI’s stated rule is to “minimize load on the primary as much as possible” (reads and writes alike) so that it has “sufficient capacity to handle write spikes” (OpenAI, 2026). The writer is the one machine nothing copies, so it gets protected.

Why is a write so expensive in Postgres? Because of how it handles concurrent access (MVCC). Updating “even a single field” copies the entire row to a new version. The old version becomes a dead tuple until vacuum cleans it up. That’s write amplification: a one-field change costs a whole row. It’s also read amplification, because queries have to scan past those dead tuples. The WAL fundamentals post linked above also covers why MVCC keeps row versions at all.

This is how a new feature breaks you. A feature that writes a lot doesn’t trip the doorman or the cache. It aims straight at the writer.

So OpenAI keeps the writer as empty as possible. It fixed bugs that caused redundant writes (the same reason idempotent writes matter for retries). It introduced lazy writes. It rate-limits backfills, which “can sometimes take over a week”. No new tables go on the current deployment; new workloads default to sharded systems. Only lightweight schema changes are allowed, with a 5-second timeout, because something like a column-type change can rewrite a whole table.

The heaviest move: shardable, write-heavy workloads were migrated to “sharded systems such as Azure Cosmos DB”. The message itself stays on the writer. There’s one truth, and hundreds of endpoints assume one database.

That leads to the honest version of the hook. OpenAI didn’t split its Postgres database. It moved the heaviest writing out of it.

Hacker News said the quiet part. In the thread on the post (348 points, 136 comments), one commenter, mannyv, wrote: “they scaled PostgreSQL by offloading a lot of it to Azure CosmosDB.” Another, bhouston, put it more drily: “Ah yes, OpenAI is sharding now.” (Hacker News). Those are opinions, not findings, but they’re fair ones.

So why not shard Postgres itself? OpenAI’s answer is cost. Sharding the existing workloads would mean “changes to hundreds of application endpoints and potentially taking months or even years.” And reads dominate anyway; the read side already scales outward. OpenAI is also explicit that it’s “not ruling out sharding PostgreSQL in the future.”

Eight hundred million, and the week it nearly broke

Diagram: rate limits at four checkpoints (app, doorman, proxy, query) in front of the Postgres primary and its standby, with the cache, nearly 50 copies and the Cosmos DB sharded store behind, an 800,000,000-user wave on the left and the user ladder from 10 to 800,000,000 along the bottom

From here on, the numbers are OpenAI’s, not mine. In twelve months, ChatGPT’s Postgres had one SEV-0 incident. It came during the ImageGen launch, “when write traffic suddenly surged by more than 10x as over 100 million new users signed up within a week” (OpenAI, 2026). The spike aimed at the one machine nothing copies.

OpenAI describes the shape every one of its outages takes. Something upstream spikes: a burst of cache misses, an expensive multi-way join, a write storm from a new feature. Latency rises. Requests time out. Clients retry, and the retries add load. Load raises latency further. OpenAI calls it a “vicious cycle”.

The guards are all about breaking that loop before it reaches the database. OpenAI doesn’t date when each one arrived, so I won’t either.

Then there’s the quiet benefit of everything earlier. Because reads moved to the copies, a primary outage is “no longer a SEV0 since reads remain available.” People can still open their chats and read their history while the writer recovers.

One SEV-0 in twelve months. And it landed on the writer, the only machine the design can’t copy its way out of.

The whole machine, one message

Diagram: the whole system for one message, from an 800,000,000-user wave through the four checkpoints, the app, PgBouncer, the Postgres primary with headroom and its standby, the cache, nearly 50 regional copies and the Cosmos DB sharded store

The finished system serves “millions of queries per second for 800 million users” with “low double-digit millisecond p99 client-side latency and five-nines availability,” on load that has “grown by more than 10x” in the past year (OpenAI, 2026; the single-primary, nearly-50-replica setup is also reported by InfoQ, 2026). Here’s one last send through all of it.

You press send. The request passes the rate limits. It reaches the app, which needs its three reads. Who you are: cache hit. Your plan: cache hit. This chat’s settings: a miss. One reader takes the lease, fetches the row from a nearby copy through PgBouncer, and fills the entry for everyone after you.

Then the write. It goes through PgBouncer to the primary, which has headroom because so much was moved off it. The primary streams the change as WAL to nearly 50 copies and to the standby. Your message didn’t touch the sharded store at all; that’s where the heavy, shardable writing lives.

It’s worth saying what “800 million users” means here. It isn’t 800 million people pressing send at once. It’s the user base, and the load that arrives from it is what the doorman, the copies, and the cache absorb.

The way I hold the whole design in my head is one sentence: scale the reads, spare the writer.

Lollipop ladder on a log scale from 10 to 800 million users. Rungs 10, 100, 10,000 and 1,000,000 are a thought experiment; 100 million and 800 million are OpenAI's numbers. Fixes: one box, bigger box, PgBouncer, nearly 50 read replicas, cache with per-key lease, spare the writer plus four rate-limit checkpoints and a hot standby.

Key Takeaways

Frequently Asked Questions

Why doesn’t OpenAI shard PostgreSQL?

Sharding the existing workloads would mean “changes to hundreds of application endpoints and potentially taking months or even years,” according to OpenAI. Reads dominate and already scale out through replicas. Shardable, write-heavy work went to sharded systems such as Azure Cosmos DB instead. OpenAI says it’s not ruling out sharding Postgres later.

What does PgBouncer do for ChatGPT?

PgBouncer sits between ChatGPT’s services and Postgres as a connection pooler. It keeps a small set of real connections open and lends one per transaction or statement. OpenAI says average connection time dropped from 50 ms to 5 ms, and thousands of clients now share a few real connections under the 5,000-per-instance Azure limit.

How many read replicas does ChatGPT’s Postgres have?

Nearly 50, spread across multiple regions, all fed by one primary. Each replica adds WAL-shipping load on the writer, so OpenAI is testing cascading replication, where intermediate replicas relay WAL onward. It says that could let it “scale to potentially over a hundred replicas.” Regions run multiple replicas with headroom for failures.

What is a cache-miss storm, and how does OpenAI guard against it?

A cache-miss storm happens when cache hit rates drop and the burst of misses hits Postgres directly, often thousands of requests for the same key. OpenAI’s guard is a per-key lease: “only a single reader that misses on a particular key fetches the data from PostgreSQL.” Everyone else waits for the refreshed entry.

Why is a write so expensive in Postgres?

Postgres uses MVCC, so updating “even a single field” copies the entire row to a new version. The old one lingers as a dead tuple. That’s write amplification, plus read amplification from scanning dead rows. It’s why OpenAI keeps the primary as empty as possible, with headroom for spikes like ImageGen’s 10x write surge.

Scale the reads, spare the writer

Every fix in this design serves the reads. The doorman, the copies, and the cache all exist so that the reads never need the writer, and the writer’s own rules (no new tables, lazy writes, slow backfills, heavy writing moved out) exist so it has room on the one week that matters. That week came with ImageGen, and the system held with one SEV-0 on the single machine nobody can copy.

It’s not a clever architecture. It’s a disciplined one. If you’d rather watch the ladder assemble itself, the animated version is in the episode at the top of this post.

Written by Nishil Bhave

Builder, maker, and tech writer at MakeToCreate.

Never miss a post

Get the latest tech insights delivered to your inbox. No spam, unsubscribe anytime.

Related Posts