OpenLDAP N-way multi-master replication (docs/replication.md) required an operator to hand-set LDAP_SERVER_ID (unique per site) and LDAP_REPLICATION_HOSTS (every OTHER site's LDAP URL, kept in sync by hand across every node) -- real coordination work, and easy to get wrong or let drift as sites are added. Automates the coordination the master is already in a position to do: - SiteSpoke gets ldapServerId, auto-assigned (next free from 2 upward, 1 reserved for the master) at registration and reused across re-registrations -- same pattern as jump-host's mesh index. - ldapHost is derived from each site's already-known HTTP(S) endpoint (same hostname, port 636) rather than a separately-configured field that could drift from it. - New utils/ldap_replication.js (nextFreeLdapServerId, ldapHostFor), shared between the spoke-facing GET /api/site/ldap-peers (Bearer site join key, returns this caller's own ID + every peer) and the master-local GET /directory-admin/ldap-replication-config (computes its own config directly from SiteSpoke, no HTTP round-trip needed). Verified against real running containers (docker-compose.multisite-e2e.yml): after a real join, the master's computed config correctly includes the spoke as a peer with an assigned ID, and the spoke's own fetched config matches that ID and correctly excludes itself from its own peer list. Known limitation, documented in docs/replication.md: the master's own LDAP_REPLICATION_HOSTS only gets recomputed when ITS setup.sh is re-run (or an admin re-applies it directly) -- there's no live push to an already-running master when a new spoke joins. A spoke's own config is re-checked on every setup.sh run, which is the common/recurring event; the master side is a documented manual step for now rather than a live hot-reload (which would need OpenLDAP's dynamic cn=config backend -- a bigger change, deliberately out of scope here to avoid risking a live directory's LDAP replication on undertested config).
4.1 KiB
layout, title
| layout | title |
|---|---|
| default | Geo-Location Scaling (Replication) |
Geo-Location Scaling (Replication)
SSO Manager is built to be a self-contained identity provider, but if you have multiple physical sites, you may want a local copy of the directory at each site to ensure low latency and high availability.
Why and when to use this?
- High Availability (HA): If your primary site goes completely offline, your other sites can still authenticate users locally without depending on a WAN link.
- Low Latency: Applications at a remote site can bind directly to their local LDAP server (
localhostor LAN IP) instead of traversing the internet to query the primary site, making logins blazing fast. - Independent Failure Domains: By replicating only the LDAP directory (the source of truth) and keeping session state (Redis) independent, you prevent complex "split-brain" scenarios in the web UI. A failure at Site A won't bring down Site B.
By default, the sso-manager Docker container runs a single, independent OpenLDAP instance. However, you can enable N-Way Multi-Master Replication via environment variables.
How it works
In an N-Way Multi-Master setup, every site runs a fully active OpenLDAP server (slapd).
- Reads and Writes anywhere: A user can change their password or update their profile at Site A, Site B, or Site C.
- Conflict Resolution: OpenLDAP's
syncreplengine uses Context Sequence Numbers (CSN) to track changes. If Site A goes offline and a user changes their password at Site B, Site A will automatically pull the newest changes the moment it rejoins the cluster. - Independent Redis: Session data, API Tokens, and OAuth Clients are stored in Redis. By design, Redis is NOT replicated in this geographic setup. This ensures that a failure at Site A never causes Site B's Redis to become read-only, which would break the web UI at Site B. OAuth clients must be configured per-site.
Configuration
The container's entrypoint reads two environment variables to configure this
-- LDAP_SERVER_ID (a unique integer for this node) and
LDAP_REPLICATION_HOSTS (a space-separated list of every other node's
LDAP URL) -- and, when both are set, automatically loads the syncprov
module, enables mirrormode, and generates the necessary syncrepl blocks
in /etc/openldap/slapd.conf.
If you're using theta-suite's setup.sh, you don't set these by hand.
The master assigns each spoke a unique LDAP_SERVER_ID at join time (the
same way it assigns a WireGuard mesh index), and LDAP_REPLICATION_HOSTS is
derived automatically from every site's already-known HTTPS endpoint
(ldaps://<same-host>:636) -- see GET /api/site/ldap-peers (spoke) and
GET /api/directory-admin/ldap-replication-config (master), and
theta-suite's bootstrap/site-ldap-register.js, which re-checks on every
setup.sh run since the peer list changes as new spokes join.
Setting the two env vars directly still works (e.g. a non-theta-suite
deployment) -- example using three manually-configured nodes:
Site 1
LDAP_SERVER_ID=1
LDAP_REPLICATION_HOSTS="ldaps://sso.site2.com:636 ldaps://sso.site3.com:636"
Site 2
LDAP_SERVER_ID=2
LDAP_REPLICATION_HOSTS="ldaps://sso.site1.com:636 ldaps://sso.site3.com:636"
Site 3
LDAP_SERVER_ID=3
LDAP_REPLICATION_HOSTS="ldaps://sso.site1.com:636 ldaps://sso.site2.com:636"
A known limitation of the automatic path: the master's own
LDAP_REPLICATION_HOSTS only gets recomputed when its setup.sh is
re-run (or the operator re-applies it directly) -- there's no live push
telling the master's already-running container about a spoke that joined
five minutes ago. A spoke's own config, by contrast, is re-checked and
applied on every setup.sh run there, which is the common/recurring event.
Re-run setup.sh on the master after bringing up a new spoke to pick up the
new peer and restart replication with it.
User Locations
When creating or editing a user, you can specify their Location (Site). This maps directly to the standard LDAP l (localityName) attribute, allowing you to track which physical site a user belongs to natively within the directory.