Killing the Cron Job: Rescuing a Production Odoo Workload on PostgreSQL
Wednesday, September 30 · 11:30–12:20
When we walked in, every query running longer than three minutes was being killed by a cron job. That cron job was the only thing keeping the system standing, and it was hiding everything that was actually broken. Core financial reports were taking over thirty minutes to run, most of them never finishing. CPU was pinned, IOPS were spiking, and bulk delete operations on operational tables were taking more than two hours.
This talk walks through the effort on a high-volume Odoo deployment, with the actual numbers from the engagement. Phase one was about stopping the bleeding: right-sizing the resources and tuning parameters against the observed workload. Reports that previously failed to complete in over thirty minutes now finish in under two minutes. Phase two was about finding what was structurally broken: a multi-day measurement window using pg_gather against real production traffic, the discovery of nine high-volume transactional tables with missing foreign key indexes, and a targeted index strategy that the workload immediately validated. Bulk deletes that previously took more than 120 minutes now complete in approximately 62 minutes. Temporary file generation dropped from 280 GB per day to 192.9 GB per day. Read IOPS spikes during business hours fell from 10–15 per day to 5–7 per day.
The talk's central argument is uncomfortable but useful: every safety mechanism that "fixes" a performance problem without addressing root cause is a tarp over the hole in the floor. The teams that get out of this trap learn to measure twice. Once urgently, to stop the bleeding so the business can keep running. Once patiently, with a real observation window, to find what is actually broken.