Compare commits

..

13 Commits

Author SHA1 Message Date
wmantly 7efd3fd6bd Merge pull request #210 from theta42/release-v2.8.0
release(v2.8.0): promotion/LDAP-orphan fix, site-slug unification, LDAP status UI
2026-08-10 20:59:31 -07:00
wmantly 9f285c960b release(v2.8.0): promotion/LDAP-orphan fix, site-slug unification, LDAP status UI 2026-08-10 23:57:04 -04:00
wmantly d75daf81b7 Merge pull request #209 from theta42/feat-ldap-status-ui
feat(directory): LDAP replication status + per-spoke detail on the Multi-Site modal
2026-08-10 20:56:32 -07:00
wmantly 6861a113d2 fix(directory): unify the Directory's site slug with the multi-site replication identity
Two previously-unrelated "site slug" concepts existed: the Directory
catalog's site Resource (site.slug, what group names and the resource
tree actually use -- e.g. "E2E Site" / site_e2e) vs. site_config.js's
siteSlug (the multi-site replication identity shown on the Multi-Site
modal's "Local Site Slug" row, sourced only from a separately-set
SITE_SLUG env var). Nothing ever kept them in sync -- a real
deployment could show a real site name in the Directory tree and the
literal "site-default" fallback on the Multi-Site modal for the exact
same node, which is exactly what a live demo of the modal surfaced.

Synced at the source: POST /resources now sets site_config's siteSlug
to match, the moment this node's own site Resource is first created
(bootstrap.js's initial call). Only for a still-default master --
never overwrites a real multi-site identity a join/promote has
already established, and never touches a spoke's identity (the
master's to assign via registration, not this node's own resource
creation to decide).

Verified against a real running container via
docker-compose.multisite-e2e.yml's existing site-creation step.
2026-08-10 23:54:05 -04:00
wmantly 4542c055bb feat(directory): LDAP replication status + per-spoke detail on the Multi-Site modal
The modal previously showed zero LDAP replication status -- no
ServerID, no MMR active/inactive indicator, nothing -- and only
aggregate counts, never per-spoke detail (no-inbound flag, relay
note, assigned ldapServerId). An operator had no way to tell whether
replication was actually configured/working without SSHing in.

Added utils/ldap_replication.js's currentSlapdServerId(), which reads
the ACTUAL running ServerID straight from this node's own slapd.conf
-- distinct from what GET /ldap-peers / /ldap-replication-config
currently ADVERTISE for it, which can genuinely disagree right after a
promotion or a new spoke joining (OpenLDAP's static config only
reloads at process start). GET /directory-admin/site-status now
surfaces both plus a `stale` flag, and a full spokes list (not just a
count) with each one's endpoint, LDAP ServerID, and relay path.

directory.ejs renders this as an "LDAP Replication (MMR)" status row
(ServerID, peer count, a "needs setup.sh re-run" warning when stale)
and a Registered Spokes table.

Verified two ways: real running containers via
docker-compose.multisite-e2e.yml (new site-status assertions), and an
actual browser session against the promoted node -- screenshotted the
rendered modal showing the real registered spoke with its assigned
ldapServerId and the correctly-surfaced "not configured (standalone)"
MMR state (this test node never ran site-ldap-register.js against it,
so the mismatch itself is the expected, documented behavior).
2026-08-10 23:48:20 -04:00
wmantly bd2205f1a9 Merge pull request #208 from theta42/docs-site-join-staleness
docs(site-join): fix stale claims about relay/mesh routing/LDAP
2026-08-10 20:31:52 -07:00
wmantly d5e0d61546 docs(site-join): fix stale "not yet built" claims about relay/mesh routing/LDAP
Three items listed as unbuilt had actually shipped: cross-component
routing over the mesh, no-inbound relay automation, and (as of this
session) OpenLDAP MMR auto-config. Replaced with what's actually true
today, plus the two known LDAP-replication caveats around promotion
(covered in the code comments added alongside the fix for the
promotion/demote SiteSpoke-orphaning bug).
2026-08-10 23:27:45 -04:00
wmantly 0dfb69cc9a Merge pull request #207 from theta42/fix-promotion-ldap-orphan
fix(multi-site): promotion no longer orphans the demoted old master's LDAP replication
2026-08-10 20:24:29 -07:00
wmantly dae0361e82 fix(multi-site): promotion no longer orphans the demoted old master's LDAP replication
Found while auditing the new LDAP MMR auto-config for gaps: neither
POST /site-promote nor POST /demote ever touched SiteSpoke. Two real
problems:

1. The demoted old master got a fresh masterJoinKey but was never
   registered as a spoke of the new master -- no SiteSpoke row, no
   ldapServerId, invisible to GET /ldap-peers's peer list. It also
   structurally could not self-heal: POST /join refuses re-join for a
   node that's already a spoke, and separately requires a fresh
   install (siteIsFresh()) -- neither true for a former master with
   real users/agents. Fixed: /demote now registers itself with the new
   master immediately (POST /spokes), the same way a real join does,
   deriving its own endpoint from stack.selfUrl (override) or
   https://stack.ssoHost (the normal case).

2. The promoted node's live OpenLDAP ServerID doesn't change --
   GET /ldap-replication-config starts advertising 1 for it
   immediately (derived purely from cfg.isMaster), but nothing
   restarts slapd with that value (OpenLDAP's static slapd.conf only
   reloads at process start, and this app has no safe way to restart
   its own container). Can't be fixed in-process; surfaced instead --
   /site-promote's response now includes ldapReplicationNote telling
   the operator to re-run setup.sh promptly.

Verified against real running containers (docker-compose.multisite-e2e.yml):
after promotion, the demoted old master correctly appears in the new
master's LDAP peer list with a real assigned ldapServerId.
2026-08-10 23:21:59 -04:00
wmantly eef7852b69 Merge pull request #206 from theta42/release-v2.7.0
release(v2.7.0): auto-assigned LDAP ServerID + replication hosts
2026-08-10 20:07:42 -07:00
wmantly 0ac0c045ec release(v2.7.0): auto-assigned LDAP ServerID + replication hosts 2026-08-10 23:03:33 -04:00
wmantly dfb819a715 Merge pull request #205 from theta42/feat-ldap-mmr-auto-config
feat(multi-site): auto-assign LDAP ServerID + replication hosts at join time
2026-08-10 19:53:42 -07:00
wmantly d486fb946b feat(multi-site): auto-assign LDAP ServerID + replication hosts at join time
OpenLDAP N-way multi-master replication (docs/replication.md) required
an operator to hand-set LDAP_SERVER_ID (unique per site) and
LDAP_REPLICATION_HOSTS (every OTHER site's LDAP URL, kept in sync by
hand across every node) -- real coordination work, and easy to get
wrong or let drift as sites are added.

Automates the coordination the master is already in a position to do:
- SiteSpoke gets ldapServerId, auto-assigned (next free from 2 upward,
  1 reserved for the master) at registration and reused across
  re-registrations -- same pattern as jump-host's mesh index.
- ldapHost is derived from each site's already-known HTTP(S) endpoint
  (same hostname, port 636) rather than a separately-configured field
  that could drift from it.
- New utils/ldap_replication.js (nextFreeLdapServerId, ldapHostFor),
  shared between the spoke-facing GET /api/site/ldap-peers (Bearer
  site join key, returns this caller's own ID + every peer) and the
  master-local GET /directory-admin/ldap-replication-config (computes
  its own config directly from SiteSpoke, no HTTP round-trip needed).

Verified against real running containers (docker-compose.multisite-e2e.yml):
after a real join, the master's computed config correctly includes the
spoke as a peer with an assigned ID, and the spoke's own fetched
config matches that ID and correctly excludes itself from its own
peer list.

Known limitation, documented in docs/replication.md: the master's own
LDAP_REPLICATION_HOSTS only gets recomputed when ITS setup.sh is
re-run (or an admin re-applies it directly) -- there's no live push to
an already-running master when a new spoke joins. A spoke's own config
is re-checked on every setup.sh run, which is the common/recurring
event; the master side is a documented manual step for now rather than
a live hot-reload (which would need OpenLDAP's dynamic cn=config
backend -- a bigger change, deliberately out of scope here to avoid
risking a live directory's LDAP replication on undertested config).
2026-08-10 22:51:04 -04:00
12 changed files with 466 additions and 21 deletions
+14
View File
@@ -1,3 +1,17 @@
# v2.8.0 - 2026-08-11
### Fixed
- **Promotion no longer orphans the demoted old master's LDAP replication.** Neither `/site-promote` nor `/demote` touched `SiteSpoke` -- the demoted old master got a fresh join key but no `SiteSpoke` entry on the new master (no `ldapServerId`, invisible to the peer list), and structurally could never self-heal via `/join` (refuses re-join for a node that's already a spoke). `/demote` now registers itself with the new master immediately, deriving its own endpoint from `stack.selfUrl`/`stack.ssoHost`. `/site-promote`'s response also now surfaces that the promoted node's own OpenLDAP ServerID needs a `setup.sh` re-run to actually apply.
- **The Directory's site slug and the multi-site replication identity are unified.** These were two unrelated values that happened to share a name -- a deployment could show a real site name in the Directory catalog and the literal `site-default` fallback on the Multi-Site modal for the same node. `POST /resources` now syncs `site_config`'s `siteSlug` to match the moment this node's own site Resource is first created, for a still-default master only.
### Added
- **LDAP replication status + per-spoke detail on the Multi-Site modal.** New `utils/ldap_replication.js`'s `currentSlapdServerId()` reads the actual running ServerID from this node's own `slapd.conf` -- distinct from what the API currently advertises, which can genuinely disagree right after a promotion or a new spoke joining. `GET /directory-admin/site-status` now surfaces both plus a `stale` flag and a full spokes list (endpoint, assigned `ldapServerId`, relay path), not just an aggregate count.
# v2.7.0 - 2026-08-11
### Added
- **OpenLDAP N-way multi-master replication auto-config.** `SiteSpoke.ldapServerId` is now auto-assigned at registration (next free from 2 upward, 1 reserved for the master -- same pattern as jump-host's mesh index), and each site's LDAP URL is derived from its already-known HTTP(S) endpoint rather than a separately-configured field. New `utils/ldap_replication.js`, `GET /api/site/ldap-peers` (spoke-facing, Bearer site join key) and `GET /directory-admin/ldap-replication-config` (master-local). Operators no longer hand-maintain `LDAP_SERVER_ID`/`LDAP_REPLICATION_HOSTS` for a `theta-suite`-joined cluster (see `theta-suite`'s `bootstrap/site-ldap-register.js`). Verified against real running containers (`docker-compose.multisite-e2e.yml`).
# v2.6.0 - 2026-08-11
### Fixed
+6
View File
@@ -25,6 +25,11 @@ services:
- LDAP_ADMIN_PASS=secret
- ORG_NAME=E2E Master
- app_oauth__jwtSecret=e2e-multisite-master-jwt-secret
# POST /demote's self-registration (routes/api_site.js) needs a real
# reachable endpoint for this container; stack.selfUrl overrides the
# normal https://<stack.ssoHost> derivation, which isn't reachable
# here (plain HTTP, no TLS/proxy in front, non-443 port).
- app_stack__selfUrl=http://master:3001
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://localhost:3001/health >/dev/null 2>&1"]
interval: 2s
@@ -42,6 +47,7 @@ services:
- LDAP_ADMIN_PASS=secret
- ORG_NAME=E2E Spoke
- app_oauth__jwtSecret=e2e-multisite-spoke-jwt-secret
- app_stack__selfUrl=http://spoke:3001
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://localhost:3001/health >/dev/null 2>&1"]
interval: 2s
+27 -8
View File
@@ -23,32 +23,51 @@ In an N-Way Multi-Master setup, every site runs a fully active OpenLDAP server (
## Configuration
To enable replication, you must pass two environment variables to the `sso-manager` container:
The container's entrypoint reads two environment variables to configure this
-- `LDAP_SERVER_ID` (a unique integer for this node) and
`LDAP_REPLICATION_HOSTS` (a space-separated list of every **other** node's
LDAP URL) -- and, when both are set, automatically loads the `syncprov`
module, enables `mirrormode`, and generates the necessary `syncrepl` blocks
in `/etc/openldap/slapd.conf`.
1. `LDAP_SERVER_ID`: A unique integer for this node (e.g., `1`, `2`, `3`). This MUST be unique across the cluster.
2. `LDAP_REPLICATION_HOSTS`: A space-separated list of the LDAP URLs of all **other** nodes in the cluster.
**If you're using `theta-suite`'s `setup.sh`, you don't set these by hand.**
The master assigns each spoke a unique `LDAP_SERVER_ID` at join time (the
same way it assigns a WireGuard mesh index), and `LDAP_REPLICATION_HOSTS` is
derived automatically from every site's already-known HTTPS endpoint
(`ldaps://<same-host>:636`) -- see `GET /api/site/ldap-peers` (spoke) and
`GET /api/directory-admin/ldap-replication-config` (master), and
`theta-suite`'s `bootstrap/site-ldap-register.js`, which re-checks on every
`setup.sh` run since the peer list changes as new spokes join.
### Example using `theta-env` / Docker Compose
Setting the two env vars directly still works (e.g. a non-`theta-suite`
deployment) -- example using three manually-configured nodes:
**Site 1 (`setup.env` or `docker-compose.yml`)**
**Site 1**
```env
LDAP_SERVER_ID=1
LDAP_REPLICATION_HOSTS="ldaps://sso.site2.com:636 ldaps://sso.site3.com:636"
```
**Site 2 (`setup.env` or `docker-compose.yml`)**
**Site 2**
```env
LDAP_SERVER_ID=2
LDAP_REPLICATION_HOSTS="ldaps://sso.site1.com:636 ldaps://sso.site3.com:636"
```
**Site 3 (`setup.env` or `docker-compose.yml`)**
**Site 3**
```env
LDAP_SERVER_ID=3
LDAP_REPLICATION_HOSTS="ldaps://sso.site1.com:636 ldaps://sso.site2.com:636"
```
Once configured, the container's entrypoint will automatically load the `syncprov` module, enable `mirrormode`, and generate the necessary `syncrepl` blocks in `/etc/openldap/slapd.conf`.
**A known limitation of the automatic path**: the *master's* own
`LDAP_REPLICATION_HOSTS` only gets recomputed when its `setup.sh` is
re-run (or the operator re-applies it directly) -- there's no live push
telling the master's already-running container about a spoke that joined
five minutes ago. A spoke's own config, by contrast, is re-checked and
applied on every `setup.sh` run there, which is the common/recurring event.
Re-run `setup.sh` on the master after bringing up a new spoke to pick up the
new peer and restart replication with it.
## User Locations
+28 -6
View File
@@ -135,11 +135,33 @@ already joined reports "already a spoke" and setup continues (idempotent).
## Not yet built
- Traffic between sites (join/export/resync) still goes over the open
network path that already reaches the target — it does not route over the
WireGuard mesh `theta-gateway` can now establish (see `MULTI_SITE_SPEC.md`).
- A no-inbound spoke (no public IP at all) still can't join — the mechanism
for a master to relay through the mesh to such a spoke is verified as
working, but nothing automates creating that route yet.
- OpenBao secret replication covers only the agent-signing key; LDAP admin
creds, JWT secret, and other per-deployment secrets aren't synced.
- A promoted spoke's own OpenLDAP `ServerID` doesn't apply live -- `POST
/site-promote` starts advertising `1` for it immediately
(`GET /directory-admin/ldap-replication-config`), but nothing restarts
`slapd` with that value automatically (its static `slapd.conf` is only
read at process start). Re-run `setup.sh` on the newly-promoted node
promptly after promotion to actually apply it.
- The master's own `LDAP_REPLICATION_HOSTS` peer list only recomputes on
its next `setup.sh` run, not live the instant a new spoke joins -- same
re-run-`setup.sh` caveat as above, just triggered by a join instead of a
promotion.
## Shipped since the above was last stale
- Traffic between sites (`utils/site_replicate.js`'s resync push) prefers a
registered spoke's WireGuard mesh IP over the open internet when one's on
file, falling back to the public endpoint on failure.
- A no-inbound spoke (no public IP at all) CAN join: `noInbound`/`meshIp`/
`publicHost` on `POST /api/site/join` drive `utils/proxy_client.js`, which
auto-creates/updates the relay route on the master's own `theta-proxy`.
Mesh peering between the two jump-hosts is still a manual, one-time step
(see `theta-suite`'s `spoke.env.example` for the operator-facing side).
A spoke with zero inbound *and* zero outbound path still can't join --
the join itself needs to reach the master's API directly.
- OpenLDAP N-way multi-master replication now auto-configures on join --
the master assigns each spoke a unique `LDAP_SERVER_ID` and derives every
site's `ldaps://` URL automatically (`GET /api/site/ldap-peers`,
`GET /directory-admin/ldap-replication-config`). See
`docs/replication.md`.
+9 -1
View File
@@ -39,7 +39,15 @@ class SiteSpoke extends Model {
noInbound: { type: 'boolean', default: false },
meshIp: { type: 'string' },
publicHost: { type: 'string' },
relayNote: { type: 'string' }
relayNote: { type: 'string' },
// OpenLDAP multi-master replication (docs/replication.md): a unique
// small integer this spoke's slapd.conf ServerID must use. Assigned
// once at registration (see api_site.js's nextFreeLdapServerId),
// reused on re-registration -- a spoke that re-registers after a
// restart must not get bumped to a new ID, same reasoning as
// jump-host's meshIndex. The master reserves 1 for itself, never
// assigned here.
ldapServerId: { type: 'integer' }
};
toPublic() {
+2 -2
View File
@@ -1,6 +1,6 @@
{
"name": "t42-theta-directory",
"version": "2.6.0",
"version": "2.8.0",
"description": "A very simple LDAP management and SSO system",
"author": [
{
@@ -11,7 +11,7 @@
"scripts": {
"start": "node ./bin/www",
"dev": "npx nodemon --ignore public/ ./bin/www",
"test": "NODE_ENV=test jest tests/groups.test.js tests/subtypes.test.js tests/site_join.test.js tests/site_config.test.js tests/site_replicate.test.js tests/proxy_client.test.js tests/reconciler.test.js tests/nmap_plugin.test.js tests/jump_client.test.js --forceExit"
"test": "NODE_ENV=test jest tests/groups.test.js tests/subtypes.test.js tests/site_join.test.js tests/site_config.test.js tests/site_replicate.test.js tests/proxy_client.test.js tests/reconciler.test.js tests/nmap_plugin.test.js tests/jump_client.test.js tests/ldap_replication.test.js --forceExit"
},
"jest": {
"testEnvironment": "node",
+81 -2
View File
@@ -365,6 +365,25 @@ router.post('/resources', async (req, res, next) => {
const ancestorSite = await Resource.findAncestorSiteSlug(r.id);
if (r.kind === 'site') {
await ensureSiteGroups(r.slug, req.user.dn, r.name, r.id);
// Two previously-unrelated "site slug" concepts: this Resource's own
// slug (the Directory catalog's site container -- what every group
// name and the resource tree actually use) vs. site_config.js's
// siteSlug (the multi-site replication identity shown on the
// Multi-Site modal, sourced only from a separately-set SITE_SLUG env
// var). They coincidentally share the name "site slug" but nothing
// ever kept them in sync -- a real deployment could show "E2E Site"
// in the Directory tree and "site-default" on the Multi-Site modal
// for the exact same node. Sync them here, the moment this node's own
// site Resource is created (bootstrap.js's first call), so there's
// one real identity instead of two that can drift apart. Only for a
// still-default master: never overwrite a real multi-site identity a
// join/promote has already established, and a spoke's replication
// identity is the master's to assign, not this node's own resource
// creation to decide.
const cfg = siteConfig.get();
if (cfg.isMaster && cfg.siteSlug === 'site-default') {
siteConfig.save({ siteSlug: r.slug });
}
} else if (gKind && ancestorSite) {
await ensureSiteGroups(ancestorSite, req.user.dn, r.name); // backfill site tier if missing
await provisionResourceGroups(r, gKind, ancestorSite, req.user.dn);
@@ -962,6 +981,7 @@ const siteConfig = require('../utils/site_config');
const { siteIsFresh } = require('../utils/site_join');
const { Agent } = require('../models/agent');
const { SiteSpoke } = require('../models/site_spoke');
const { ldapHostFor, currentSlapdServerId } = require('../utils/ldap_replication');
// probeMasterHealth checks whether this (spoke) node can reach its master over
// the site join key. The master's /api/site/ping is deliberately lightweight.
@@ -1001,13 +1021,29 @@ router.get('/site-status', async (req, res, next) => {
// via an older bootstrap, or the UI form before it grew the field) is
// fully joined but silently stuck on the one-time snapshot, which was
// otherwise invisible anywhere in the UI.
const registeredSpokesCount = cfg.isMaster ? await SiteSpoke.list().then(l => l.length).catch(() => 0) : 0;
const allSpokes = cfg.isMaster ? await SiteSpoke.list().catch(() => []) : [];
const registeredSpokesCount = allSpokes.length;
// Real gateway-to-gateway mesh peer count from jump-host's own registry
// (utils/jump_client.js), not this app's unrelated WireGuard
// roaming-client Resources. count is null (not 0) when the query
// couldn't run at all -- the UI distinguishes "0 gateways" from "can't
// tell" instead of showing a misleading zero.
const gateways = await jumpClient.getGatewayCount();
// LDAP MMR status (docs/replication.md): configuredServerId is read
// straight from THIS node's own live slapd.conf; advertisedServerId is
// what GET /ldap-peers / /ldap-replication-config currently hand out for
// it. These can genuinely disagree -- a promotion or a newly-joined
// spoke changes the advertised value immediately, but OpenLDAP's static
// config only reloads at process start, so a mismatch means "re-run
// setup.sh here" rather than "something's broken". Only computed for the
// master (a spoke's advertised ID lives on the master, not locally, and
// querying it here would mean another WAN round-trip on every page load).
const configuredServerId = currentSlapdServerId();
const ldap = cfg.isMaster
? { configuredServerId, advertisedServerId: 1, stale: configuredServerId !== null && configuredServerId !== 1, peersCount: allSpokes.filter(s => s.ldapServerId).length }
: { configuredServerId, advertisedServerId: null, stale: null, peersCount: null };
res.json({
status: 'ok',
config: {
@@ -1023,11 +1059,44 @@ router.get('/site-status', async (req, res, next) => {
sitesCount: sites.length,
sites: sites.map(s => ({ id: s.id, name: s.name, slug: s.slug })),
gatewaysCount: gateways.count,
gatewaysNote: gateways.note
gatewaysNote: gateways.note,
ldap,
// Per-spoke detail (master only) -- endpoint/siteSlug/noInbound/
// relayNote/ldapServerId, not just an aggregate count, so an operator
// can actually see what's registered instead of only "N spokes".
spokes: allSpokes.map(s => ({
siteSlug: s.siteSlug, endpoint: s.endpoint, noInbound: !!s.noInbound,
relayNote: s.relayNote || null, ldapServerId: s.ldapServerId || null,
lastSeenOn: s.last_seen_on || null
}))
});
} catch (err) { next(err); }
});
// OpenLDAP multi-master replication config for THIS node (docs/replication.md).
// Master-only: the master already has every registered spoke's info locally
// (SiteSpoke), so it can compute its own ServerID (always 1) + full peer
// list without an HTTP round-trip. A spoke gets its config from the master
// directly instead (GET /api/site/ldap-peers -- see bootstrap/
// site-ldap-register.js in theta-suite, which calls whichever of the two
// applies to this node's role).
router.get('/ldap-replication-config', async (req, res, next) => {
try {
const cfg = siteConfig.get();
if (!cfg.isMaster) {
return res.status(400).json({ status: 'error', message: 'this node is a spoke -- fetch replication config from the master via GET /api/site/ldap-peers instead' });
}
const spokes = await SiteSpoke.list();
const peers = [];
for (const s of spokes) {
if (!s.ldapServerId) continue;
const host = ldapHostFor(s.endpoint);
if (host) peers.push({ ldapServerId: s.ldapServerId, ldapHost: host });
}
res.json({ status: 'ok', ldapServerId: 1, peers });
} catch (err) { next(err); }
});
router.post('/site-promote', async (req, res, next) => {
try {
// god_admin privilege check. This used to read req.user.groups, which
@@ -1091,6 +1160,16 @@ router.post('/site-promote', async (req, res, next) => {
status: 'ok',
message: 'Node successfully promoted to Master Site',
handoff: handoffNote,
// This node's own OpenLDAP ServerID stays whatever it was as a spoke
// (e.g. 2) until `setup.sh` is re-run here -- GET
// /ldap-replication-config will immediately start advertising 1 for
// this node (the master's reserved ID) since that's derived purely
// from cfg.isMaster, but nothing restarts slapd with the new value
// automatically (OpenLDAP's static slapd.conf is only read at process
// start, and this app has no safe way to restart its own container).
// Surfaced here + on the Multi-Site modal so an operator promoting a
// site knows to re-run setup.sh promptly, not just assume it's done.
ldapReplicationNote: 'Re-run setup.sh on this node to apply its new LDAP ServerID (1) and pick up the current spoke peer list -- OpenLDAP config only reloads at process start.',
config: {
isMaster: true,
masterUrl: '',
+82 -1
View File
@@ -42,6 +42,8 @@ function logAudit(action, details) {
console.log(JSON.stringify({ timestamp: new Date().toISOString(), component: 'site', action, ...details }));
}
const { nextFreeLdapServerId, ldapHostFor } = require('../utils/ldap_replication');
// slurpLdif dumps the local LDAP tree with slapcat (the sso-manager container
// carries an OpenLDAP build with slapcat on PATH).
async function slurpLdif() {
@@ -136,6 +138,7 @@ router.post('/spokes', async (req, res, next) => {
endpoint,
pushToken: SiteSpoke.generatePushToken(),
created_on: now,
ldapServerId: await nextFreeLdapServerId(),
...patch
});
}
@@ -160,6 +163,47 @@ router.post('/spokes', async (req, res, next) => {
} catch (e) { next(e); }
});
// ── LDAP replication peer list (SPOKE-callable, Bearer site join key) ───────
// OpenLDAP multi-master replication (docs/replication.md) needs each site to
// know its own ServerID plus every OTHER site's LDAPS URL. The master
// coordinates ID assignment (nextFreeLdapServerId, above); this is how a
// spoke asks "what's my ID, and who are my peers" -- called by
// theta-suite's bootstrap/site-ldap-register.js on every setup.sh run, not
// just once at join time, since the peer list changes as other spokes join.
// Same join-key auth as /spokes (a spoke already has this stored from its
// own join). `endpoint` identifies the CALLER so it can be excluded from its
// own peer list -- same identity SiteSpoke.list() keys registration on.
router.get('/ldap-peers', async (req, res, next) => {
try {
const auth = req.headers.authorization || '';
const rawKey = auth.startsWith('Bearer ') ? auth.slice(7).trim() : '';
const key = await SiteJoinKey.authenticate(rawKey);
if (!key) return res.status(401).json({ status: 'error', message: 'invalid or revoked site join key' });
const callerEndpoint = req.query.endpoint;
if (!callerEndpoint) {
return res.status(400).json({ status: 'error', message: 'endpoint query param is required' });
}
const cfg = siteConfig.get();
const masterHost = ldapHostFor(cfg.masterUrl || req.protocol + '://' + req.get('host'));
const spokes = await SiteSpoke.list();
const caller = spokes.find((s) => s.endpoint === callerEndpoint);
if (!caller || !caller.ldapServerId) {
return res.status(404).json({ status: 'error', message: 'this endpoint is not a registered spoke -- register via POST /api/site/spokes first' });
}
const peers = [{ ldapServerId: 1, ldapHost: masterHost }];
for (const s of spokes) {
if (s.endpoint === callerEndpoint || !s.ldapServerId) continue;
const host = ldapHostFor(s.endpoint);
if (host) peers.push({ ldapServerId: s.ldapServerId, ldapHost: host });
}
res.json({ status: 'ok', ldapServerId: caller.ldapServerId, peers });
} catch (e) { next(e); }
});
// ── Resync (SPOKE side, Bearer pushToken; no admin session) ─────────────────
// The receiving end of utils/site_replicate.js's fire-and-forget push: the
// master pings this when its catalog changes. Deliberately just
@@ -216,7 +260,44 @@ router.post('/demote', async (req, res, next) => {
const base = String(newMasterUrl).replace(/\/+$/, '');
siteConfig.save({ isMaster: false, masterUrl: base, masterJoinKey: newJoinKey });
logAudit('demoted', { demotedBy: key.keyPrefix, newMasterUrl: base });
res.json({ status: 'ok', message: 'Demoted to spoke of ' + base });
// Register with the new master immediately, the same way a real /join
// does (POST /spokes) -- without this, a demoted former master was
// orphaned: it had a masterJoinKey but no SiteSpoke entry on the new
// master (so no ldapServerId, no live replication push target), and
// structurally could never self-heal via /join (which refuses re-join
// for a node that's already a spoke, and requires a fresh install --
// neither true for a former master with real users/agents). Best-effort:
// failing to register here must not fail the demotion itself, same
// reasoning as a normal join's optional live-replication registration.
let registrationNote = 'not attempted (no stack.ssoHost/stack.selfUrl configured to register with)';
// stack.selfUrl is a full-URL override (scheme + port) for environments
// where "https://<ssoHost>" isn't the real reachable address -- the
// multisite e2e test harness (plain HTTP, docker-network hostnames,
// no TLS/proxy in front) is exactly that case; every real deployment
// just relies on the ssoHost derivation.
const selfUrl = (conf.stack && conf.stack.selfUrl) || (conf.stack && conf.stack.ssoHost && `https://${conf.stack.ssoHost}`);
if (selfUrl) {
try {
const regResp = await fetch(base + '/api/site/spokes', {
method: 'POST',
headers: { Authorization: 'Bearer ' + newJoinKey, 'Content-Type': 'application/json' },
body: JSON.stringify({ endpoint: selfUrl, siteSlug: cfg.siteSlug })
});
if (regResp.ok) {
const regBody = await regResp.json();
if (regBody.pushToken) siteConfig.save({ replicationPushToken: regBody.pushToken });
registrationNote = 'registered as a spoke of the new master';
} else {
registrationNote = 'registration failed: HTTP ' + regResp.status;
}
} catch (e) {
registrationNote = 'registration failed: ' + e.message;
}
}
logAudit('demoted_self_registered', { newMasterUrl: base, registrationNote });
res.json({ status: 'ok', message: 'Demoted to spoke of ' + base, registration: { note: registrationNote } });
} catch (e) { next(e); }
});
+44
View File
@@ -0,0 +1,44 @@
require('./setup');
const { SiteSpoke } = require('../models/site_spoke');
const { nextFreeLdapServerId, ldapHostFor } = require('../utils/ldap_replication');
describe('ldap_replication', () => {
beforeEach(async () => {
const all = await SiteSpoke.list();
for (const s of all) await s.delete();
});
describe('ldapHostFor', () => {
test('derives ldaps://<host>:636 from an http(s) endpoint, ignoring its own port', () => {
expect(ldapHostFor('https://sso.site2.example.com')).toBe('ldaps://sso.site2.example.com:636');
expect(ldapHostFor('https://sso.site2.example.com:8443')).toBe('ldaps://sso.site2.example.com:636');
expect(ldapHostFor('http://sso.site3.example.com')).toBe('ldaps://sso.site3.example.com:636');
});
test('returns null for an unparseable endpoint', () => {
expect(ldapHostFor('not-a-url')).toBeNull();
expect(ldapHostFor('')).toBeNull();
});
});
describe('nextFreeLdapServerId', () => {
test('starts at 2 (1 is reserved for the master) when no spokes are registered', async () => {
await expect(nextFreeLdapServerId()).resolves.toBe(2);
});
test('picks the lowest free id, not just the next highest', async () => {
const now = Math.floor(Date.now() / 1000);
await SiteSpoke.create({ id: 'a', endpoint: 'https://a.example.com', pushToken: 'tok-a', created_on: now, ldapServerId: 2 });
await SiteSpoke.create({ id: 'b', endpoint: 'https://b.example.com', pushToken: 'tok-b', created_on: now, ldapServerId: 4 });
await expect(nextFreeLdapServerId()).resolves.toBe(3);
});
test('ignores spokes with no ldapServerId assigned yet', async () => {
const now = Math.floor(Date.now() / 1000);
await SiteSpoke.create({ id: 'c', endpoint: 'https://c.example.com', pushToken: 'tok-c', created_on: now });
await expect(nextFreeLdapServerId()).resolves.toBe(2);
});
});
});
+68
View File
@@ -0,0 +1,68 @@
'use strict';
const fs = require('fs');
// OpenLDAP multi-master replication (docs/replication.md) config derivation,
// shared between routes/api_site.js (the spoke-facing side: assigns a
// ServerID at registration, serves GET /api/site/ldap-peers) and
// routes/api_directory_admin.js (the master-local side: GET
// /directory-admin/ldap-replication-config computes the master's own
// replication config from the same SiteSpoke registry, no HTTP round-trip
// needed since it already has the data).
const { SiteSpoke } = require('../models/site_spoke');
// The master reserves ServerID 1 for itself; every spoke gets the lowest
// free ID from 2 upward, assigned once at registration and reused across
// re-registrations (SiteSpoke.ldapServerId is only ever set on first
// create). Small max, matching mesh_gateway.js's mesh index -- nothing in
// the OpenLDAP protocol requires a small ServerID, but this deployment's
// docs/examples always have.
const MAX_LDAP_SERVER_ID = 4094;
async function nextFreeLdapServerId() {
const spokes = await SiteSpoke.list();
const used = new Set(spokes.map((s) => s.ldapServerId).filter(Boolean));
for (let i = 2; i <= MAX_LDAP_SERVER_ID; i++) {
if (!used.has(i)) return i;
}
throw new Error(`LDAP server ID space exhausted (max ${MAX_LDAP_SERVER_ID} spokes)`);
}
// A site's LDAP replication URL, derived from its already-known HTTP(S)
// endpoint rather than requiring a separately-configured field: same
// hostname, LDAPS port 636 -- exactly the convention docs/replication.md's
// own worked examples already use (ldaps://sso.site2.com:636 alongside
// https://sso.site2.com). No new config an operator has to keep in sync.
function ldapHostFor(endpoint) {
try {
const host = new URL(endpoint).hostname;
return `ldaps://${host}:636`;
} catch (e) {
return null;
}
}
const SLAPD_CONF_PATH = process.env.SLAPD_CONF_PATH || '/etc/openldap/slapd.conf';
// The ServerID this node's OpenLDAP is ACTUALLY running with right now, read
// straight from slapd.conf (the same file docker-entrypoint.sh writes
// `ServerID <n>` into). This can genuinely differ from what
// GET /ldap-peers / /ldap-replication-config currently ADVERTISE for this
// node -- OpenLDAP's static slapd.conf is only read at process start, so a
// promotion or a new spoke joining doesn't retroactively change what's
// already running until `setup.sh` restarts the container. Surfaced on the
// Multi-Site modal so an operator can see "configured X, but slapd is still
// running Y" instead of assuming replication is live because the API says so.
function currentSlapdServerId() {
let contents;
try {
contents = fs.readFileSync(SLAPD_CONF_PATH, 'utf8');
} catch (e) {
return null;
}
const m = contents.match(/^ServerID\s+(\d+)/m);
return m ? Number(m[1]) : null;
}
module.exports = { MAX_LDAP_SERVER_ID, nextFreeLdapServerId, ldapHostFor, currentSlapdServerId };
+52 -1
View File
@@ -3319,6 +3319,55 @@
} catch (e) { console.error('Failed to fetch site status:', e); }
}
// ldap: { configuredServerId, advertisedServerId, stale, peersCount } from
// GET /directory-admin/site-status (routes/api_directory_admin.js).
// configuredServerId is read from THIS node's live slapd.conf;
// advertisedServerId (master only) is what the API currently hands spokes.
// They can genuinely disagree right after a promotion or a new spoke
// joining -- OpenLDAP's static config only reloads at process start.
function renderLdapStatus(ldap) {
if (!ldap) return '<span class="text-muted">unknown</span>';
if (ldap.configuredServerId == null) {
return '<span class="badge bg-secondary"><i class="fa-solid fa-circle-minus me-1"></i> Not configured (standalone)</span>';
}
let html = '<span class="badge bg-dark">ServerID ' + esc(ldap.configuredServerId) + '</span>';
if (ldap.peersCount != null) {
html += ' <span class="badge bg-primary">' + ldap.peersCount + ' peer' + (ldap.peersCount === 1 ? '' : 's') + '</span>';
}
if (ldap.stale) {
html += ' <span class="badge bg-warning text-dark" title="This node now advertises ServerID ' + esc(ldap.advertisedServerId) +
' (e.g. after a promotion), but slapd is still running with ' + esc(ldap.configuredServerId) +
' -- re-run setup.sh here to apply it."><i class="fa-solid fa-triangle-exclamation me-1"></i> Needs setup.sh re-run</span>';
}
return html;
}
// spokes: [{siteSlug, endpoint, noInbound, relayNote, ldapServerId, lastSeenOn}]
// Per-spoke detail (master only) so an operator can see what's actually
// registered instead of only an aggregate count.
function renderSpokesTable(spokes) {
if (!spokes || !spokes.length) return '';
const rows = spokes.map(function(s) {
return '<tr>' +
'<td><code>' + esc(s.siteSlug || '?') + '</code></td>' +
'<td class="small">' + esc(s.endpoint) + '</td>' +
'<td>' + (s.ldapServerId != null ? '<span class="badge bg-dark">' + esc(s.ldapServerId) + '</span>' : '<span class="text-muted small">unassigned</span>') + '</td>' +
'<td>' + (s.noInbound
? '<span class="badge bg-info text-dark" title="' + esc(s.relayNote || '') + '"><i class="fa-solid fa-diagram-project me-1"></i> Relayed</span>'
: '<span class="text-muted small">direct</span>') + '</td>' +
'</tr>';
}).join('');
return '<div class="card mt-3">' +
'<div class="card-header py-2 fw-bold small"><i class="fa-solid fa-diagram-project me-1"></i> Registered Spokes</div>' +
'<div class="card-body p-0">' +
'<table class="table table-sm table-hover mb-0">' +
'<thead><tr><th>Site</th><th>Endpoint</th><th>LDAP ServerID</th><th>Path</th></tr></thead>' +
'<tbody>' + rows + '</tbody>' +
'</table>' +
'</div>' +
'</div>';
}
async function openSiteStatusModal() {
try {
const res = await app.api.get('directory-admin/site-status');
@@ -3351,12 +3400,14 @@
'<tr><th>Theta Gateways:</th><td>' + (res.gatewaysCount == null
? '<span class="badge bg-secondary" title="' + esc(res.gatewaysNote || 'not configured') + '"><i class="fa-solid fa-question me-1"></i> Unknown (jump-host integration not configured)</span>'
: '<span class="badge bg-dark">' + res.gatewaysCount + ' active gateway' + (res.gatewaysCount === 1 ? '' : 's') + '</span>') + '</td></tr>' +
'<tr><th>LDAP Replication (MMR):</th><td>' + renderLdapStatus(res.ldap) + '</td></tr>' +
'</table>' +
'</div>' +
'</div>' +
'<div class="alert alert-secondary small mb-3">' +
'<i class="fa-solid fa-network-wired me-1"></i> <strong>WireGuard Gateway Mesh & NETMAP</strong>: Inter-site routing operates via <code>theta-gateway</code> subnets (<code>10.x.0.0/16</code>) with default NETMAP shadow translations (<code>10.x.168.0/24 &rarr; 192.168.1.0/24</code>).' +
'</div>';
'</div>' +
(isMaster ? renderSpokesTable(res.spokes) : '');
// Fresh install (no users/resources yet): offer to JOIN an existing
// master site instead of seeding a new directory.
+53
View File
@@ -172,6 +172,12 @@ async function main() {
});
if (siteRes.status !== 200) fail(`seeding pre-join site on master failed: ${siteRes.status} ${JSON.stringify(siteRes.body)}`);
step('Verifying the Directory site Resource\'s slug synced into the multi-site replication identity');
const { body: masterCfgAfterSite } = await api(MASTER_URL, '/api/site/config', { token: masterToken });
if (masterCfgAfterSite.config.siteSlug !== 'site_e2e') {
fail(`expected site_config's siteSlug to sync to the new site Resource's slug (site_e2e), got ${JSON.stringify(masterCfgAfterSite.config.siteSlug)}`);
}
const seedRes = await api(MASTER_URL, '/api/directory-admin/resources', {
method: 'POST',
token: masterToken,
@@ -205,6 +211,43 @@ async function main() {
if (spokeCfg.config.isMaster !== false) fail(`spoke should be isMaster:false after join, got ${JSON.stringify(spokeCfg.config)}`);
if (!spokeCfg.config.masterUrl) fail('spoke should have masterUrl set after join');
step('Verifying the master computed its own LDAP replication config (ServerID 1 + the new spoke as a peer)');
const { body: masterLdapCfg } = await api(MASTER_URL, '/api/directory-admin/ldap-replication-config', { token: masterToken });
if (masterLdapCfg.ldapServerId !== 1) fail(`master's own ldapServerId should be 1, got ${JSON.stringify(masterLdapCfg)}`);
const spokePeer = (masterLdapCfg.peers || []).find(p => p.ldapHost === 'ldaps://spoke:636');
if (!spokePeer || typeof spokePeer.ldapServerId !== 'number') {
fail(`master's peer list should include the spoke at ldaps://spoke:636 with an assigned ldapServerId, got ${JSON.stringify(masterLdapCfg.peers)}`);
}
step('Verifying the spoke can fetch its own assigned LDAP ServerID + peer list from the master');
const spokeLdapPeersResp = await fetch(`${MASTER_URL}/api/site/ldap-peers?endpoint=${encodeURIComponent('http://spoke:3001')}`, {
headers: { Authorization: 'Bearer ' + joinKey }
});
const spokeLdapCfg = await spokeLdapPeersResp.json();
if (spokeLdapPeersResp.status !== 200) fail(`GET /api/site/ldap-peers failed: ${spokeLdapPeersResp.status} ${JSON.stringify(spokeLdapCfg)}`);
if (spokeLdapCfg.ldapServerId !== spokePeer.ldapServerId) {
fail(`spoke's own reported ldapServerId (${spokeLdapCfg.ldapServerId}) should match what the master's peer list assigned it (${spokePeer.ldapServerId})`);
}
const masterAsPeer = (spokeLdapCfg.peers || []).find(p => p.ldapServerId === 1);
if (!masterAsPeer || masterAsPeer.ldapHost !== 'ldaps://master:636') {
fail(`spoke's peer list should include the master (ServerID 1, ldaps://master:636), got ${JSON.stringify(spokeLdapCfg.peers)}`);
}
const selfInOwnPeerList = (spokeLdapCfg.peers || []).some(p => p.ldapServerId === spokeLdapCfg.ldapServerId);
if (selfInOwnPeerList) fail(`spoke's own peer list should not include itself, got ${JSON.stringify(spokeLdapCfg.peers)}`);
step('Verifying the master\'s own site-status surfaces LDAP status + per-spoke detail (Multi-Site modal data)');
const { body: masterStatus } = await api(MASTER_URL, '/api/directory-admin/site-status', { token: masterToken });
if (!masterStatus.ldap || masterStatus.ldap.advertisedServerId !== 1) {
fail(`master's site-status should report ldap.advertisedServerId 1, got ${JSON.stringify(masterStatus.ldap)}`);
}
if (masterStatus.ldap.peersCount !== 1) {
fail(`master's site-status should report exactly 1 LDAP peer (the spoke), got ${JSON.stringify(masterStatus.ldap)}`);
}
const statusSpokeEntry = (masterStatus.spokes || []).find(s => s.endpoint === 'http://spoke:3001');
if (!statusSpokeEntry || typeof statusSpokeEntry.ldapServerId !== 'number') {
fail(`master's site-status spokes list should include the spoke with an ldapServerId, got ${JSON.stringify(masterStatus.spokes)}`);
}
step('Verifying the spoke adopted the master\'s pre-join catalog');
const spokeResources = await api(SPOKE_URL, '/api/directory-admin/resources', { token: spokeToken });
const adopted = (spokeResources.body.results || spokeResources.body.resources || spokeResources.body || []);
@@ -257,6 +300,16 @@ async function main() {
if (promoteRes.body.handoff !== 'previous master demoted') {
fail(`expected the old master to be demoted as part of promotion, got handoff=${JSON.stringify(promoteRes.body.handoff)}`);
}
if (!promoteRes.body.ldapReplicationNote) {
fail('expected /site-promote to surface a note that this node\'s LDAP ServerID needs a setup.sh re-run to apply');
}
step('Verifying the demoted old master auto-registered itself as a real spoke of the new master (not orphaned)');
const { body: newMasterLdapCfg } = await api(SPOKE_URL, '/api/directory-admin/ldap-replication-config', { token: spokeToken });
const oldMasterAsPeer = (newMasterLdapCfg.peers || []).find(p => p.ldapHost === 'ldaps://master:636');
if (!oldMasterAsPeer || typeof oldMasterAsPeer.ldapServerId !== 'number') {
fail(`the demoted old master should appear as a registered peer with an assigned ldapServerId, got ${JSON.stringify(newMasterLdapCfg.peers)}`);
}
step('Verifying the newly-promoted node is master');
const { body: newMasterCfg } = await api(SPOKE_URL, '/api/site/config', { token: spokeToken });