d486fb946b
OpenLDAP N-way multi-master replication (docs/replication.md) required an operator to hand-set LDAP_SERVER_ID (unique per site) and LDAP_REPLICATION_HOSTS (every OTHER site's LDAP URL, kept in sync by hand across every node) -- real coordination work, and easy to get wrong or let drift as sites are added. Automates the coordination the master is already in a position to do: - SiteSpoke gets ldapServerId, auto-assigned (next free from 2 upward, 1 reserved for the master) at registration and reused across re-registrations -- same pattern as jump-host's mesh index. - ldapHost is derived from each site's already-known HTTP(S) endpoint (same hostname, port 636) rather than a separately-configured field that could drift from it. - New utils/ldap_replication.js (nextFreeLdapServerId, ldapHostFor), shared between the spoke-facing GET /api/site/ldap-peers (Bearer site join key, returns this caller's own ID + every peer) and the master-local GET /directory-admin/ldap-replication-config (computes its own config directly from SiteSpoke, no HTTP round-trip needed). Verified against real running containers (docker-compose.multisite-e2e.yml): after a real join, the master's computed config correctly includes the spoke as a peer with an assigned ID, and the spoke's own fetched config matches that ID and correctly excludes itself from its own peer list. Known limitation, documented in docs/replication.md: the master's own LDAP_REPLICATION_HOSTS only gets recomputed when ITS setup.sh is re-run (or an admin re-applies it directly) -- there's no live push to an already-running master when a new spoke joins. A spoke's own config is re-checked on every setup.sh run, which is the common/recurring event; the master side is a documented manual step for now rather than a live hot-reload (which would need OpenLDAP's dynamic cn=config backend -- a bigger change, deliberately out of scope here to avoid risking a live directory's LDAP replication on undertested config).
75 lines
4.1 KiB
Markdown
75 lines
4.1 KiB
Markdown
---
|
|
layout: default
|
|
title: Geo-Location Scaling (Replication)
|
|
---
|
|
|
|
# Geo-Location Scaling (Replication)
|
|
|
|
SSO Manager is built to be a self-contained identity provider, but if you have multiple physical sites, you may want a local copy of the directory at each site to ensure low latency and high availability.
|
|
|
|
## Why and when to use this?
|
|
- **High Availability (HA)**: If your primary site goes completely offline, your other sites can still authenticate users locally without depending on a WAN link.
|
|
- **Low Latency**: Applications at a remote site can bind directly to their local LDAP server (`localhost` or LAN IP) instead of traversing the internet to query the primary site, making logins blazing fast.
|
|
- **Independent Failure Domains**: By replicating only the LDAP directory (the source of truth) and keeping session state (Redis) independent, you prevent complex "split-brain" scenarios in the web UI. A failure at Site A won't bring down Site B.
|
|
|
|
By default, the `sso-manager` Docker container runs a single, independent OpenLDAP instance. However, you can enable **N-Way Multi-Master Replication** via environment variables.
|
|
|
|
## How it works
|
|
|
|
In an N-Way Multi-Master setup, every site runs a fully active OpenLDAP server (`slapd`).
|
|
- **Reads and Writes anywhere**: A user can change their password or update their profile at Site A, Site B, or Site C.
|
|
- **Conflict Resolution**: OpenLDAP's `syncrepl` engine uses Context Sequence Numbers (CSN) to track changes. If Site A goes offline and a user changes their password at Site B, Site A will automatically pull the newest changes the moment it rejoins the cluster.
|
|
- **Independent Redis**: Session data, API Tokens, and OAuth Clients are stored in Redis. By design, Redis is NOT replicated in this geographic setup. This ensures that a failure at Site A never causes Site B's Redis to become read-only, which would break the web UI at Site B. OAuth clients must be configured per-site.
|
|
|
|
## Configuration
|
|
|
|
The container's entrypoint reads two environment variables to configure this
|
|
-- `LDAP_SERVER_ID` (a unique integer for this node) and
|
|
`LDAP_REPLICATION_HOSTS` (a space-separated list of every **other** node's
|
|
LDAP URL) -- and, when both are set, automatically loads the `syncprov`
|
|
module, enables `mirrormode`, and generates the necessary `syncrepl` blocks
|
|
in `/etc/openldap/slapd.conf`.
|
|
|
|
**If you're using `theta-suite`'s `setup.sh`, you don't set these by hand.**
|
|
The master assigns each spoke a unique `LDAP_SERVER_ID` at join time (the
|
|
same way it assigns a WireGuard mesh index), and `LDAP_REPLICATION_HOSTS` is
|
|
derived automatically from every site's already-known HTTPS endpoint
|
|
(`ldaps://<same-host>:636`) -- see `GET /api/site/ldap-peers` (spoke) and
|
|
`GET /api/directory-admin/ldap-replication-config` (master), and
|
|
`theta-suite`'s `bootstrap/site-ldap-register.js`, which re-checks on every
|
|
`setup.sh` run since the peer list changes as new spokes join.
|
|
|
|
Setting the two env vars directly still works (e.g. a non-`theta-suite`
|
|
deployment) -- example using three manually-configured nodes:
|
|
|
|
**Site 1**
|
|
```env
|
|
LDAP_SERVER_ID=1
|
|
LDAP_REPLICATION_HOSTS="ldaps://sso.site2.com:636 ldaps://sso.site3.com:636"
|
|
```
|
|
|
|
**Site 2**
|
|
```env
|
|
LDAP_SERVER_ID=2
|
|
LDAP_REPLICATION_HOSTS="ldaps://sso.site1.com:636 ldaps://sso.site3.com:636"
|
|
```
|
|
|
|
**Site 3**
|
|
```env
|
|
LDAP_SERVER_ID=3
|
|
LDAP_REPLICATION_HOSTS="ldaps://sso.site1.com:636 ldaps://sso.site2.com:636"
|
|
```
|
|
|
|
**A known limitation of the automatic path**: the *master's* own
|
|
`LDAP_REPLICATION_HOSTS` only gets recomputed when its `setup.sh` is
|
|
re-run (or the operator re-applies it directly) -- there's no live push
|
|
telling the master's already-running container about a spoke that joined
|
|
five minutes ago. A spoke's own config, by contrast, is re-checked and
|
|
applied on every `setup.sh` run there, which is the common/recurring event.
|
|
Re-run `setup.sh` on the master after bringing up a new spoke to pick up the
|
|
new peer and restart replication with it.
|
|
|
|
## User Locations
|
|
|
|
When creating or editing a user, you can specify their **Location (Site)**. This maps directly to the standard LDAP `l` (localityName) attribute, allowing you to track which physical site a user belongs to natively within the directory.
|