Backups, upgrades and operations

Odoo high availability and database failover

Keep Odoo up when a server dies: stateless app nodes behind a load balancer, a shared filestore, a PostgreSQL streaming replica and a safe failover.

CICDoo Engineering Updated 8 min read

The short answer

Odoo high availability has two halves. The Odoo side is easy to duplicate: run the same code on several servers against one database, share the filestore and sessions, and put a load balancer with health checks in front. The database is the hard half: keep a PostgreSQL streaming replica on another server, through a replication slot with a cap on retained WAL, and promote it when the primary is lost. Promotion is a one-way door, so fence the old primary first and decide whether a human or a tool like Patroni pulls the trigger.

On this page
  1. What "high availability" means for Odoo
  2. Running several Odoo app nodes
  3. Making the database survive a server loss
  4. Failover: promoting the standby
  5. Testing it
  6. Doing this with CICDoo

What "high availability" means for Odoo

An Odoo deployment fails in two different places, and each needs its own answer:

  1. The Odoo process or its server. A crashed worker, a full disk, a kernel panic, a host that a provider reboots. Odoo itself keeps no state in memory that matters between requests, so the answer is more copies of it.
  2. The PostgreSQL database. Every record lives there, and there is exactly one writable copy. The answer is a second, continuously updated copy that can take over.

Most outages people call "Odoo is down" are the first kind. The second kind is rarer and far more expensive, because a lost database server without a standby means restoring from the last backup and losing everything since.

Before building anything, write down two numbers: how long you can be down (recovery time objective) and how much recent data you can afford to lose (recovery point objective). Nightly backups alone give you a recovery point of up to a day. A streaming replica brings it down to seconds.

Running several Odoo app nodes

Odoo app nodes are interchangeable as long as everything they share lives outside them.

One database, one dbfilter

Every node points at the same PostgreSQL server and database, with identical odoo.conf database settings:

db_host = 10.0.0.10
db_port = 5432
db_user = odoo
db_password = <secret>
db_name = production
dbfilter = ^production$
proxy_mode = True

All nodes must run the same Odoo version, the same addons at the same commit, and the same Python dependencies. A node with a different module version will happily write data the others cannot read.

The filestore and sessions

Odoo stores attachments under data_dir/filestore/<db> and HTTP sessions under data_dir/sessions. If each node has its own data_dir, an attachment uploaded through one node is missing on the others, and a visitor whose next request lands elsewhere is logged out. The options:

  • A shared filesystem (NFS, CephFS, a managed file share) mounted at the same data_dir on every node. Simple, but the share itself becomes a dependency to make redundant.
  • Object storage for attachments through a community module that stores the filestore in S3-compatible storage, plus a shared session store. More moving parts, no shared filesystem.
  • Sticky sessions at the load balancer hide the session problem but not the filestore one, and a node failure still logs its users out.

The load balancer

Put a load balancer or a proxy with active health checks in front of the nodes. Odoo 16 and later answer GET /web/health with a small JSON body when the worker is up; older versions can be checked on /web/login. Route /websocket (Odoo 16 and later) or /longpolling (earlier versions) to the gevent port of the same node as the rest of the request. Odoo's bus uses PostgreSQL LISTEN/NOTIFY, so live notifications work across nodes without extra configuration.

DNS round robin with health checks (several A records, removed when a node fails) is a lighter alternative when a managed load balancer is not available, at the cost of slower failover because of caching.

Scheduled jobs

Odoo locks each scheduled action in the database before running it, so crons running on every node do not execute the same job twice. They do multiply database connections and background load. A common choice is to run crons on one node and set max_cron_threads = 0 on the rest.

Upgrades

A module upgrade (-u) must run on one node while no other node holds a registry on the database, or Odoo's lock timeouts abort it halfway. Stop the other nodes, upgrade on one, then start the rest with the new code. Rolling upgrades are not safe for module updates.

Connections

Each node opens its own pool of connections, up to db_maxconn per worker process. Add nodes and the total grows quickly. Raise max_connections on PostgreSQL to match, or put PgBouncer in front of it in transaction mode.

Making the database survive a server loss

Streaming replication in brief

PostgreSQL ships its write-ahead log (WAL) to a standby server that replays it continuously. The standby is a byte-for-byte copy, so it needs the same PostgreSQL major version and architecture as the primary. With hot_standby = on it also answers read-only queries.

On the primary, the defaults since PostgreSQL 10 already allow replication (wal_level = replica, max_wal_senders = 10). Create a dedicated role and allow it in pg_hba.conf from the standby's address only:

CREATE ROLE replicator WITH LOGIN REPLICATION PASSWORD '<secret>';
SELECT pg_create_physical_replication_slot('standby1', true);
hostssl  replication  replicator  203.0.113.20/32  scram-sha-256

Replication slots and the WAL cap

A replication slot makes the primary keep every WAL segment the standby has not received yet, so a standby that was down for an hour catches up instead of needing a full copy. The danger is the opposite case: a standby that never comes back pins WAL forever and fills the primary's disk. On PostgreSQL 13 and later, cap it:

ALTER SYSTEM SET max_slot_wal_keep_size = '20GB';
SELECT pg_reload_conf();

A standby that falls further behind than the cap loses its slot (wal_status = lost in pg_replication_slots) and has to be copied again, but the primary stays up. Drop slots for standbys you retire.

Taking the copy

On the standby server, with PostgreSQL stopped and an empty data directory:

pg_basebackup -h 203.0.113.10 -U replicator -D /var/lib/postgresql/data \
  -X stream -S standby1 -c fast -R

-R writes primary_conninfo and an empty standby.signal file (PostgreSQL 12 and later), which is what makes the server start as a standby. Start it, then check on the primary:

SELECT application_name, state, replay_lag FROM pg_stat_replication;
SELECT slot_name, active, wal_status,
       pg_current_wal_lsn() - restart_lsn AS behind_bytes
  FROM pg_replication_slots;

A standby refuses to start if max_connections, max_worker_processes, max_wal_senders, max_prepared_transactions or max_locks_per_transaction are lower than on the primary. Keep its settings identical, and restart it after raising any of them on the primary.

Asynchronous or synchronous

Replication is asynchronous by default: a commit returns before the standby has it, and a crash can lose the last few seconds. Setting synchronous_standby_names with synchronous_commit = on closes that gap, but then every commit waits for the standby, and if the only standby goes down, writes stop. For a single standby, asynchronous replication with monitored lag is the usual compromise. Synchronous replication makes sense with two or more standbys and ANY 1 (...).

Failover: promoting the standby

When the primary is lost, the standby becomes the new primary:

SELECT pg_promote();

After that it accepts writes on a new timeline, and the old primary's data has diverged from it. Three things decide whether the failover is clean.

Fence the old primary first

The worst outcome is two writable primaries: the old one comes back after a network blip, some app nodes still write to it, and you now have two databases with different orders and invoices. Before or at promotion, make sure the old primary cannot accept writes: stop its PostgreSQL, cut its network, or have its app nodes repointed and its own Odoo stopped. Never let a demoted primary start PostgreSQL in read-write mode again.

Repoint every app node

Each node's db_host must change to the new primary, or a virtual IP or DNS name used by all nodes must move. Then restart the nodes, because Odoo keeps open connections to the old address. Writes the old primary accepted after the standby's last replayed position are gone. On asynchronous replication that is the price of failover, and it is why promotion should be a decision, not a reflex.

Automatic or manual

Tools such as Patroni, repmgr or pg_auto_failover detect a failed primary and promote a standby automatically. They need a consensus store (etcd, Consul) or a monitor node, at least three voting members to avoid split brain, and careful testing. For a single Odoo database, a manual promotion with a clear runbook and good alerts is often the safer choice: an automatic failover triggered by a transient network problem causes exactly the two-primary situation fencing is meant to prevent.

After the failover

The old primary cannot simply rejoin. Either rebuild it as a new standby with a fresh pg_basebackup, or rewind it with pg_rewind (which needs wal_log_hints = on or data checksums enabled beforehand). Until then the cluster has no standby, so treat rebuilding it as part of the incident, not a later task.

Read replicas for Odoo itself

Odoo 18 can send some read-only requests to a replica through the db_replica_host and db_replica_port options. That spreads read load but is a different goal from failover: the replica still has to be promoted, and the configuration still has to change, when the primary is lost.

Testing it

High availability that has never failed over is a hypothesis. Schedule drills:

  • Stop one app node during working hours and confirm users notice nothing beyond a reload.
  • Watch replication lag under a heavy job, such as a large import or a month-end report.
  • On a copy of the setup, stop the primary, promote the standby, repoint the nodes, and time it. Then rebuild the old primary as a standby.
  • Restore last night's backup somewhere. Replication copies mistakes too: a deleted table is deleted on the standby within seconds, so backups stay necessary.

Doing this with CICDoo

CICDoo builds this cluster from projects on your own servers. A primary project runs Odoo and the database; app nodes on other servers run Odoo against it; a replica is an app node that also keeps a streaming standby of the primary's database, through its own replication slot with a WAL cap. The primary's host keeps pg_hba.conf, TLS between servers and a firewall rule limited to the cluster's addresses, and reports each replica as copying, streaming, lagging or broken. Traffic for the site's name is spread over the members that pass health checks, cluster upgrades stop the nodes before -u runs on the primary, and Promote to primary fails the cluster over to a replica: its standby is promoted, every other member is repointed and restarted, and the old primary is fenced out of traffic and comes back as an app node without a database of its own. Promotion is always a human decision. See High availability with app nodes, self-hosted Odoo with CICDoo, or talk to us.

Part of Self-hosted Odoo, done right

FAQ

Questions, answered

Can Odoo run on several servers at once?

Yes. Odoo app servers are stateless between requests, so several can serve one database at the same time behind a load balancer, as long as they run identical code and share the filestore and sessions.

Does Odoo support database failover?

Odoo itself does not manage the database. Failover is done at the PostgreSQL level: keep a streaming replica on another server, promote it with pg_promote() when the primary is lost, and point every Odoo node at the new primary.

Do I need sticky sessions for an Odoo cluster?

Not if the session directory is shared between nodes. Sticky sessions only hide the session problem; attachments still need a shared filestore or object storage.

Should Odoo crons run on every node?

They can, because Odoo locks each scheduled action before running it, but every node adds database connections and background load. Running crons on one node with max_cron_threads = 0 on the others is a common setup.

How much data can I lose in a PostgreSQL failover?

With asynchronous replication, the transactions committed on the primary that the standby had not replayed yet, usually a few seconds. Synchronous replication avoids that loss at the cost of commit latency and of blocking writes when no synchronous standby is reachable.

Should failover be automatic?

Only with a proper consensus setup such as Patroni with etcd and at least three members, plus fencing. For a single Odoo database a manual promotion with monitoring and a rehearsed runbook avoids promoting on a false alarm and ending up with two primaries.

Is a replica a replacement for backups?

No. A replica copies every change within seconds, including mistakes such as deleted records or a broken module upgrade. Keep regular backups with retention alongside it.

Run Odoo like this, without doing it by hand

CICDoo turns every step in this guide into a push to Git, on servers you own. Talk to an engineer about your setup.

No per-seat fees. Your servers stay yours.