Closes the gap where all of this session's new server-side capability (live replication, coordinated promotion, identical signing keys) had no UI at all -- an operator using the Master Site modal had no way to know any of it existed or was working. - Master Site modal: new "Live Replication" row (spoke) shows whether this join actually registered for live updates or is stuck on a one-time snapshot; new "Registered Spokes" row (master) shows how many spokes are receiving live pushes. - Join form: new "this site's own reachable URL" field, prefilled from window.location.origin, wired to the selfUrl the join API already supported but the UI never sent -- a UI-driven join previously NEVER registered for live replication, only the setup.sh bootstrap path did. The success toast now reports whether live replication actually activated, not just "joined". - Promote button: success toast now surfaces the handoff result (old master demoted / unreachable / no previous master), so the operator sees immediately whether the coordinated demotion actually happened. - GET /api/site/config no longer returns masterJoinKey or replicationPushToken in the response -- found while wiring this up: live credentials were being sent straight to the browser for every admin session. Replaced with boolean derivatives (hasMasterJoinKey, liveReplication). - GET /api/directory-admin/site-status gained liveReplication (spoke) and registeredSpokesCount (master) so the modal has something to render. Verified by actually driving it in a real browser against a live container (not just code review): logged in, opened the modal, saw the new rows, minted a real join key end-to-end, no console errors. docs/site-join.md rewritten to cover live replication, signing-key sync, coordinated promotion/demote, and the new endpoints -- it previously only described the v2.2.0-v2.3.0 one-time-snapshot behavior.
8.3 KiB
Multi-Site: Joining a Spoke to the Master Directory
The Directory can be deployed across multiple sites. The master site holds
single write authority for the shared catalog; spoke sites run a read-only
copy for local latency and autonomy (see the root MULTI_SITE_SPEC.md for the
full architecture). This page covers the server endpoints that make a spoke
"join" an existing master.
Status: server endpoints + UI + setup.sh wiring, live replication, coordinated promotion. A fresh bring-up can adopt a master directory via the Directory UI or via
setup.env; a joined spoke is read-only with live WAN health, stays in sync after joining (not just a one-time snapshot), and can be promoted to master with the old master demoted as part of the same action.
The flow
- On the master, an admin mints a site join key (
stj_…, shown once, stored hashed, revocable) — Directory → the Master Site modal → Site Join Keys. - On the spoke (a fresh install), either:
- UI: Directory → the Master Site modal → Join an Existing Site, or
- setup.sh: set
CFG_MASTER_DIRECTORY_URL+CFG_MASTER_DIRECTORY_JOIN_KEYinsetup.envbefore the first run.
- The spoke pulls the master's directory export (LDAP tree + resource
catalog + agent-signing key), imports it, and persists its own spoke role
(
isMaster: false,masterUrl,siteSlug) in/config/site.json. - If the spoke also knows its own reachable URL (
selfUrl—setup.shpasseshttps://$CFG_SSO_HOSTautomatically), it registers itself with the master (POST /api/site/spokes) so the master can push live updates back to it afterward — see Live replication below. WithoutselfUrlthe join still succeeds; the spoke just stays a one-time snapshot.
Joining is allowed only on a fresh install (no users beyond the bootstrap admin, no enrolled agents) — the join endpoint enforces this, so a populated directory can never be merged into a master's.
Live replication (not a one-time snapshot)
A registered spoke stays in sync: every successful catalog write on the
master fires a fire-and-forget push (utils/site_replicate.js) at every
registered spoke, concurrently — one unreachable spoke never blocks or delays
delivery to another. The spoke's POST /api/site/resync handler (called by
that push) re-runs the same export-pull-and-import logic used at join time,
so there's exactly one tested code path for "make my catalog match the
master's," not a separate diff-application mechanism.
The agent-signing key travels the same path: POST /api/site/export
best-effort includes it, and the spoke adopts it via agent_keys.adopt() on
both join and every resync. Every site holding the same signing key means any
site's sso-manager-node can validly sign a command for any agent enrolled
at any other site — a deliberate tradeoff (see MULTI_SITE_SPEC.md §2)
accepted for this deployment's small, trusted scale. Don't extend this
pattern to a larger/adversarial-tenant deployment without revisiting it.
Coordinated master promotion
POST /api/directory-admin/site-promote (god_admin only) promotes this
node to master as one coordinated action, not a manual two-step
demote-then-promote:
- If this node currently has a master on file, it mints a fresh join key and
calls that master's
POST /api/site/demote(authenticated with the join key this node already holds), handing over the new key so the demoted node can keep talking to the new master afterward. - This step is best-effort — an unreachable old master (the WAN-outage
scenario this whole control exists for) never blocks the local promotion.
The response's
handofffield reports what happened ("previous master demoted", an HTTP failure, or "unreachable, promoted locally anyway") so the operator can reconcile it manually if needed. - Every known spoke gets a fire-and-forget
master-promotedresync ping so they pick up the new master on their next sync.
The Master Site modal's Promote to Master button surfaces the handoff
result in a toast so the operator sees immediately whether the old master was
actually reached.
Endpoints
| Method | Path | Purpose |
|---|---|---|
GET |
/api/site/join-keys |
List keys (prefix + usage only; never the key) |
POST |
/api/site/join-keys |
Mint one — returned once |
POST |
/api/site/join-keys/:id/revoke |
Stop it accepting new joins |
DELETE |
/api/site/join-keys/:id |
Remove it |
GET |
/api/site/config |
Current role (isMaster, masterUrl, siteSlug) |
POST |
/api/site/export |
Master directory export incl. agent-signing key (Bearer stj_ key) |
POST |
/api/site/ping |
Lightweight master reachability probe (Bearer stj_ key) |
POST |
/api/site/join |
Adopt a master directory + register for live replication (admin session) |
POST |
/api/site/spokes |
Register a spoke's endpoint for live replication (Bearer stj_ key, called by the spoke right after join) |
POST |
/api/site/resync |
Re-pull the master's export (Bearer the spoke's own pushToken, called by the master's fire-and-forget push) |
POST |
/api/site/demote |
Step down to spoke of a new master (Bearer stj_ key, called by the newly-promoted node) |
POST |
/api/directory-admin/site-promote |
Promote this node to master, coordinating demotion of the old one (god_admin session) |
Behavior after joining (spoke)
- Read-only: directory-write requests (resources, edges, groups, secrets,
grants, driver actions, discovery merges) are rejected with
403pointing at the master. Writes must go to the master. - WAN health:
site-statuspings the master over the stored site join key and reportswanConnected; the Master Site modal shows live Online/Offline. - Role persists:
isMaster/masterUrl/siteSluglive in/config/site.json(the env varsIS_MASTER/MASTER_URL/SITE_SLUGonly seed the defaults), so a restart never silently reverts a spoke to master.
Deployment (setup.sh)
setup.env carries the intent so the join runs only on a fresh bring-up:
# Honored ONLY on first run; re-runs ignore it once ./config/ exists.
CFG_MASTER_DIRECTORY_URL=https://sso.master.example.com
CFG_MASTER_DIRECTORY_JOIN_KEY=stj_9f2e...
setup.sh runs bootstrap/site-join.js inside the sso-manager container after
the bootstrap; it logs in as the admin and calls /api/site/join. A node that
already joined reports "already a spoke" and setup continues (idempotent).
Security
- Join keys are single-use-intent credentials: shown once, stored as a SHA-256 hash, revocable/expirable — the same model as agent join keys.
- The export/ping endpoints return only the directory tree/catalog (no admin secrets) and require a valid join key.
- Join is admin-gated on the spoke, key-gated on the master, and fresh-install gated on both sides.
- The join key is stored on the spoke only so it can reach the master for WAN health (and, in a later layer, write-proxy).
pushToken(the credential a spoke stores so it can recognize a legitimate resync push from its master) is minted fresh per spoke registration and, by design, kept in retrievable form on the master — unlike a join key, it's a credential the master must keep presenting, not just verifying, so it can't be one-way hashed. Comparemodels/site_spoke.js's doc comment for why that's the correct tradeoff, not an oversight.- Every site sharing one agent-signing key (see Live replication above)
means a compromised spoke — including the smallest, least-secured one — has
the same agent-command authority as the master. Accepted for this
deployment's scale; see
MULTI_SITE_SPEC.md§2 before reusing this pattern somewhere that assumption doesn't hold.
Not yet built
- Traffic between sites (join/export/resync) still goes over the open
network path that already reaches the target — it does not route over the
WireGuard mesh
theta-gatewaycan now establish (seeMULTI_SITE_SPEC.md). - A no-inbound spoke (no public IP at all) still can't join — the mechanism for a master to relay through the mesh to such a spoke is verified as working, but nothing automates creating that route yet.
- OpenBao secret replication covers only the agent-signing key; LDAP admin creds, JWT secret, and other per-deployment secrets aren't synced.