The Pager is the Product: PostgreSQL in Cloud Environments
Friday, October 02 · 13:30–14:20
When a managed Postgres service takes over part of the stack, the actions your on-call engineer can take during an incident change too. This talk compares self-managed and managed PostgreSQL in the cloud from the point of view of the person holding the pager at 2 a.m. It uses four questions: what you can change with your own access, whether you can see the layer that's failing, what the application has to do after failover, and whether the rotation can actually carry out the recovery.
I'll cover what self-management asks of you, such as fencing, promotion, restore drills and not depending on one specialist. I'll also cover where a managed service's guardrails help and where they block an emergency procedure you haven't rehearsed. A fast database failover can still mean a slow recovery: pools full of dead connections, stale endpoints and unbounded retries often decide your real recovery time. The talk also covers what a second region adds to either model, and how to agree on downtime and data-loss tolerance before you need them.
I'll finish with a decision framework: how to choose the operating model your rotation can support, how to explain that choice to leadership and to on-call, which signals should make you revisit it, and how to rehearse a cutover when you move between models.