feat(multi-site): UI for live replication, promotion handoff, signing key

Closes the gap where all of this session's new server-side capability
(live replication, coordinated promotion, identical signing keys) had
no UI at all -- an operator using the Master Site modal had no way to
know any of it existed or was working.

- Master Site modal: new "Live Replication" row (spoke) shows whether
  this join actually registered for live updates or is stuck on a
  one-time snapshot; new "Registered Spokes" row (master) shows how
  many spokes are receiving live pushes.
- Join form: new "this site's own reachable URL" field, prefilled from
  window.location.origin, wired to the selfUrl the join API already
  supported but the UI never sent -- a UI-driven join previously NEVER
  registered for live replication, only the setup.sh bootstrap path did.
  The success toast now reports whether live replication actually
  activated, not just "joined".
- Promote button: success toast now surfaces the handoff result (old
  master demoted / unreachable / no previous master), so the operator
  sees immediately whether the coordinated demotion actually happened.
- GET /api/site/config no longer returns masterJoinKey or
  replicationPushToken in the response -- found while wiring this up:
  live credentials were being sent straight to the browser for every
  admin session. Replaced with boolean derivatives
  (hasMasterJoinKey, liveReplication).
- GET /api/directory-admin/site-status gained liveReplication (spoke)
  and registeredSpokesCount (master) so the modal has something to render.

Verified by actually driving it in a real browser against a live
container (not just code review): logged in, opened the modal, saw the
new rows, minted a real join key end-to-end, no console errors.

docs/site-join.md rewritten to cover live replication, signing-key sync,
coordinated promotion/demote, and the new endpoints -- it previously only
described the v2.2.0-v2.3.0 one-time-snapshot behavior.
This commit is contained in:
2026-08-10 18:30:48 -04:00
parent 9c604f0258
commit dc3d760d2b
4 changed files with 119 additions and 13 deletions
+75 -6
View File
@@ -6,9 +6,7 @@ copy for local latency and autonomy (see the root `MULTI_SITE_SPEC.md` for the
full architecture). This page covers the server endpoints that make a spoke
"join" an existing master.
> Status: **server endpoints + UI + setup.sh wiring.** A fresh bring-up can
> adopt a master directory via the Directory UI or via `setup.env`, and a
> joined spoke is read-only with live WAN health.
> Status: **server endpoints + UI + setup.sh wiring, live replication, coordinated promotion.** A fresh bring-up can adopt a master directory via the Directory UI or via `setup.env`; a joined spoke is read-only with live WAN health, stays in sync after joining (not just a one-time snapshot), and can be promoted to master with the old master demoted as part of the same action.
## The flow
@@ -19,13 +17,58 @@ full architecture). This page covers the server endpoints that make a spoke
- **setup.sh**: set `CFG_MASTER_DIRECTORY_URL` + `CFG_MASTER_DIRECTORY_JOIN_KEY`
in `setup.env` before the first run.
3. The spoke pulls the master's directory export (LDAP tree + resource
catalog), imports it, and persists its own spoke role
catalog + agent-signing key), imports it, and persists its own spoke role
(`isMaster: false`, `masterUrl`, `siteSlug`) in `/config/site.json`.
4. If the spoke also knows its own reachable URL (`selfUrl``setup.sh` passes
`https://$CFG_SSO_HOST` automatically), it registers itself with the master
(`POST /api/site/spokes`) so the master can push live updates back to it
afterward — see **Live replication** below. Without `selfUrl` the join still
succeeds; the spoke just stays a one-time snapshot.
Joining is allowed only on a **fresh install** (no users beyond the bootstrap
admin, no enrolled agents) — the join endpoint enforces this, so a populated
directory can never be merged into a master's.
## Live replication (not a one-time snapshot)
A registered spoke stays in sync: every successful catalog write on the
master fires a fire-and-forget push (`utils/site_replicate.js`) at every
registered spoke, concurrently — one unreachable spoke never blocks or delays
delivery to another. The spoke's `POST /api/site/resync` handler (called by
that push) re-runs the same export-pull-and-import logic used at join time,
so there's exactly one tested code path for "make my catalog match the
master's," not a separate diff-application mechanism.
The agent-signing key travels the same path: `POST /api/site/export`
best-effort includes it, and the spoke adopts it via `agent_keys.adopt()` on
both join and every resync. Every site holding the same signing key means any
site's `sso-manager-node` can validly sign a command for any agent enrolled
at any other site — a deliberate tradeoff (see `MULTI_SITE_SPEC.md` §2)
accepted for this deployment's small, trusted scale. Don't extend this
pattern to a larger/adversarial-tenant deployment without revisiting it.
## Coordinated master promotion
`POST /api/directory-admin/site-promote` (`god_admin` only) promotes this
node to master as **one coordinated action**, not a manual two-step
demote-then-promote:
1. If this node currently has a master on file, it mints a fresh join key and
calls that master's `POST /api/site/demote` (authenticated with the join
key this node already holds), handing over the new key so the demoted node
can keep talking to the new master afterward.
2. This step is **best-effort** — an unreachable old master (the WAN-outage
scenario this whole control exists for) never blocks the local promotion.
The response's `handoff` field reports what happened
(`"previous master demoted"`, an HTTP failure, or "unreachable, promoted
locally anyway") so the operator can reconcile it manually if needed.
3. Every known spoke gets a fire-and-forget `master-promoted` resync ping so
they pick up the new master on their next sync.
The Master Site modal's **Promote to Master** button surfaces the `handoff`
result in a toast so the operator sees immediately whether the old master was
actually reached.
## Endpoints
| Method | Path | Purpose |
@@ -35,9 +78,13 @@ directory can never be merged into a master's.
| `POST` | `/api/site/join-keys/:id/revoke` | Stop it accepting new joins |
| `DELETE` | `/api/site/join-keys/:id` | Remove it |
| `GET` | `/api/site/config` | Current role (isMaster, masterUrl, siteSlug) |
| `POST` | `/api/site/export` | Master directory export (Bearer `stj_` key) |
| `POST` | `/api/site/export` | Master directory export incl. agent-signing key (Bearer `stj_` key) |
| `POST` | `/api/site/ping` | Lightweight master reachability probe (Bearer `stj_` key) |
| `POST` | `/api/site/join` | Adopt a master directory (admin session) |
| `POST` | `/api/site/join` | Adopt a master directory + register for live replication (admin session) |
| `POST` | `/api/site/spokes` | Register a spoke's endpoint for live replication (Bearer `stj_` key, called by the spoke right after join) |
| `POST` | `/api/site/resync` | Re-pull the master's export (Bearer the spoke's own `pushToken`, called by the master's fire-and-forget push) |
| `POST` | `/api/site/demote` | Step down to spoke of a new master (Bearer `stj_` key, called by the newly-promoted node) |
| `POST` | `/api/directory-admin/site-promote` | Promote this node to master, coordinating demotion of the old one (`god_admin` session) |
## Behavior after joining (spoke)
@@ -74,3 +121,25 @@ already joined reports "already a spoke" and setup continues (idempotent).
gated on both sides.
- The join key is stored on the spoke only so it can reach the master for WAN
health (and, in a later layer, write-proxy).
- `pushToken` (the credential a spoke stores so it can recognize a legitimate
resync push from its master) is minted fresh per spoke registration and, by
design, kept in retrievable form on the master — unlike a join key, it's a
credential the master must keep *presenting*, not just verifying, so it
can't be one-way hashed. Compare `models/site_spoke.js`'s doc comment for
why that's the correct tradeoff, not an oversight.
- Every site sharing one agent-signing key (see **Live replication** above)
means a compromised spoke — including the smallest, least-secured one — has
the same agent-command authority as the master. Accepted for this
deployment's scale; see `MULTI_SITE_SPEC.md` §2 before reusing this pattern
somewhere that assumption doesn't hold.
## Not yet built
- Traffic between sites (join/export/resync) still goes over the open
network path that already reaches the target — it does not route over the
WireGuard mesh `theta-gateway` can now establish (see `MULTI_SITE_SPEC.md`).
- A no-inbound spoke (no public IP at all) still can't join — the mechanism
for a master to relay through the mesh to such a spoke is verified as
working, but nothing automates creating that route yet.
- OpenBao secret replication covers only the agent-signing key; LDAP admin
creds, JWT secret, and other per-deployment secrets aren't synced.