Backups, upgrades and operations
Odoo high availability and database failover
Keep Odoo up when a server dies: stateless app nodes behind a load balancer, a shared filestore, a PostgreSQL streaming replica and a safe failover.
CICDoo Engineering Updated 8 min read
The short answer
Odoo high availability has two halves. The Odoo side is easy to duplicate: run the same code on several servers against one database, share the filestore and sessions, and put a load balancer with health checks in front. The database is the hard half: keep a PostgreSQL streaming replica on another server, through a replication slot with a cap on retained WAL, and promote it when the primary is lost. Promotion is a one-way door, so fence the old primary first and decide whether a human or a tool like Patroni pulls the trigger.
On this page
What "high availability" means for Odoo
An Odoo deployment fails in two different places, and each needs its own answer:
- The Odoo process or its server. A crashed worker, a full disk, a kernel panic, a host that a provider reboots. Odoo itself keeps no state in memory that matters between requests, so the answer is more copies of it.
- The PostgreSQL database. Every record lives there, and there is exactly one writable copy. The answer is a second, continuously updated copy that can take over.
Most outages people call "Odoo is down" are the first kind. The second kind is rarer and far more expensive, because a lost database server without a standby means restoring from the last backup and losing everything since.
Before building anything, write down two numbers: how long you can be down (recovery time objective) and how much recent data you can afford to lose (recovery point objective). Nightly backups alone give you a recovery point of up to a day. A streaming replica brings it down to seconds.
Running several Odoo app nodes
Odoo app nodes are interchangeable as long as everything they share lives outside them.
One database, one dbfilter
Every node points at the same PostgreSQL server and database, with identical odoo.conf database settings:
db_host = 10.0.0.10
db_port = 5432
db_user = odoo
db_password = <secret>
db_name = production
dbfilter = ^production$
proxy_mode = True
All nodes must run the same Odoo version, the same addons at the same commit, and the same Python dependencies. A node with a different module version will happily write data the others cannot read.
The filestore and sessions
Odoo stores attachments under data_dir/filestore/<db> and HTTP sessions under data_dir/sessions. If each node has its own data_dir, an attachment uploaded through one node is missing on the others, and a visitor whose next request lands elsewhere is logged out. The options:
- A shared filesystem (NFS, CephFS, a managed file share) mounted at the same
data_diron every node. Simple, but the share itself becomes a dependency to make redundant. - Object storage for attachments through a community module that stores the filestore in S3-compatible storage, plus a shared session store. More moving parts, no shared filesystem.
- Sticky sessions at the load balancer hide the session problem but not the filestore one, and a node failure still logs its users out.
The load balancer
Put a load balancer or a proxy with active health checks in front of the nodes. Odoo 16 and later answer GET /web/health with a small JSON body when the worker is up; older versions can be checked on /web/login. Route /websocket (Odoo 16 and later) or /longpolling (earlier versions) to the gevent port of the same node as the rest of the request. Odoo's bus uses PostgreSQL LISTEN/NOTIFY, so live notifications work across nodes without extra configuration.
DNS round robin with health checks (several A records, removed when a node fails) is a lighter alternative when a managed load balancer is not available, at the cost of slower failover because of caching.
Scheduled jobs
Odoo locks each scheduled action in the database before running it, so crons running on every node do not execute the same job twice. They do multiply database connections and background load. A common choice is to run crons on one node and set max_cron_threads = 0 on the rest.
Upgrades
A module upgrade (-u) must run on one node while no other node holds a registry on the database, or Odoo's lock timeouts abort it halfway. Stop the other nodes, upgrade on one, then start the rest with the new code. Rolling upgrades are not safe for module updates.
Connections
Each node opens its own pool of connections, up to db_maxconn per worker process. Add nodes and the total grows quickly. Raise max_connections on PostgreSQL to match, or put PgBouncer in front of it in transaction mode.
Making the database survive a server loss
Streaming replication in brief
PostgreSQL ships its write-ahead log (WAL) to a standby server that replays it continuously. The standby is a byte-for-byte copy, so it needs the same PostgreSQL major version and architecture as the primary. With hot_standby = on it also answers read-only queries.
On the primary, the defaults since PostgreSQL 10 already allow replication (wal_level = replica, max_wal_senders = 10). Create a dedicated role and allow it in pg_hba.conf from the standby's address only:
CREATE ROLE replicator WITH LOGIN REPLICATION PASSWORD '<secret>';
SELECT pg_create_physical_replication_slot('standby1', true);
hostssl replication replicator 203.0.113.20/32 scram-sha-256
Replication slots and the WAL cap
A replication slot makes the primary keep every WAL segment the standby has not received yet, so a standby that was down for an hour catches up instead of needing a full copy. The danger is the opposite case: a standby that never comes back pins WAL forever and fills the primary's disk. On PostgreSQL 13 and later, cap it:
ALTER SYSTEM SET max_slot_wal_keep_size = '20GB';
SELECT pg_reload_conf();
A standby that falls further behind than the cap loses its slot (wal_status = lost in pg_replication_slots) and has to be copied again, but the primary stays up. Drop slots for standbys you retire.
Taking the copy
On the standby server, with PostgreSQL stopped and an empty data directory:
pg_basebackup -h 203.0.113.10 -U replicator -D /var/lib/postgresql/data \
-X stream -S standby1 -c fast -R
-R writes primary_conninfo and an empty standby.signal file (PostgreSQL 12 and later), which is what makes the server start as a standby. Start it, then check on the primary:
SELECT application_name, state, replay_lag FROM pg_stat_replication;
SELECT slot_name, active, wal_status,
pg_current_wal_lsn() - restart_lsn AS behind_bytes
FROM pg_replication_slots;
A standby refuses to start if max_connections, max_worker_processes, max_wal_senders, max_prepared_transactions or max_locks_per_transaction are lower than on the primary. Keep its settings identical, and restart it after raising any of them on the primary.
Asynchronous or synchronous
Replication is asynchronous by default: a commit returns before the standby has it, and a crash can lose the last few seconds. Setting synchronous_standby_names with synchronous_commit = on closes that gap, but then every commit waits for the standby, and if the only standby goes down, writes stop. For a single standby, asynchronous replication with monitored lag is the usual compromise. Synchronous replication makes sense with two or more standbys and ANY 1 (...).
Failover: promoting the standby
When the primary is lost, the standby becomes the new primary:
SELECT pg_promote();
After that it accepts writes on a new timeline, and the old primary's data has diverged from it. Three things decide whether the failover is clean.
Fence the old primary first
The worst outcome is two writable primaries: the old one comes back after a network blip, some app nodes still write to it, and you now have two databases with different orders and invoices. Before or at promotion, make sure the old primary cannot accept writes: stop its PostgreSQL, cut its network, or have its app nodes repointed and its own Odoo stopped. Never let a demoted primary start PostgreSQL in read-write mode again.
Repoint every app node
Each node's db_host must change to the new primary, or a virtual IP or DNS name used by all nodes must move. Then restart the nodes, because Odoo keeps open connections to the old address. Writes the old primary accepted after the standby's last replayed position are gone. On asynchronous replication that is the price of failover, and it is why promotion should be a decision, not a reflex.
Automatic or manual
Tools such as Patroni, repmgr or pg_auto_failover detect a failed primary and promote a standby automatically. They need a consensus store (etcd, Consul) or a monitor node, at least three voting members to avoid split brain, and careful testing. For a single Odoo database, a manual promotion with a clear runbook and good alerts is often the safer choice: an automatic failover triggered by a transient network problem causes exactly the two-primary situation fencing is meant to prevent.
After the failover
The old primary cannot simply rejoin. Either rebuild it as a new standby with a fresh pg_basebackup, or rewind it with pg_rewind (which needs wal_log_hints = on or data checksums enabled beforehand). Until then the cluster has no standby, so treat rebuilding it as part of the incident, not a later task.
Read replicas for Odoo itself
Odoo 18 can send some read-only requests to a replica through the db_replica_host and db_replica_port options. That spreads read load but is a different goal from failover: the replica still has to be promoted, and the configuration still has to change, when the primary is lost.
Testing it
High availability that has never failed over is a hypothesis. Schedule drills:
- Stop one app node during working hours and confirm users notice nothing beyond a reload.
- Watch replication lag under a heavy job, such as a large import or a month-end report.
- On a copy of the setup, stop the primary, promote the standby, repoint the nodes, and time it. Then rebuild the old primary as a standby.
- Restore last night's backup somewhere. Replication copies mistakes too: a deleted table is deleted on the standby within seconds, so backups stay necessary.
Doing this with CICDoo
CICDoo builds this cluster from projects on your own servers. A primary project runs Odoo and the database; app nodes on other servers run Odoo against it; a replica is an app node that also keeps a streaming standby of the primary's database, through its own replication slot with a WAL cap. The primary's host keeps pg_hba.conf, TLS between servers and a firewall rule limited to the cluster's addresses, and reports each replica as copying, streaming, lagging or broken. Traffic for the site's name is spread over the members that pass health checks, cluster upgrades stop the nodes before -u runs on the primary, and Promote to primary fails the cluster over to a replica: its standby is promoted, every other member is repointed and restarted, and the old primary is fenced out of traffic and comes back as an app node without a database of its own. Promotion is always a human decision. See High availability with app nodes, self-hosted Odoo with CICDoo, or talk to us.