feat(multi-site): UI for live replication, promotion handoff, signing key
Closes the gap where all of this session's new server-side capability (live replication, coordinated promotion, identical signing keys) had no UI at all -- an operator using the Master Site modal had no way to know any of it existed or was working. - Master Site modal: new "Live Replication" row (spoke) shows whether this join actually registered for live updates or is stuck on a one-time snapshot; new "Registered Spokes" row (master) shows how many spokes are receiving live pushes. - Join form: new "this site's own reachable URL" field, prefilled from window.location.origin, wired to the selfUrl the join API already supported but the UI never sent -- a UI-driven join previously NEVER registered for live replication, only the setup.sh bootstrap path did. The success toast now reports whether live replication actually activated, not just "joined". - Promote button: success toast now surfaces the handoff result (old master demoted / unreachable / no previous master), so the operator sees immediately whether the coordinated demotion actually happened. - GET /api/site/config no longer returns masterJoinKey or replicationPushToken in the response -- found while wiring this up: live credentials were being sent straight to the browser for every admin session. Replaced with boolean derivatives (hasMasterJoinKey, liveReplication). - GET /api/directory-admin/site-status gained liveReplication (spoke) and registeredSpokesCount (master) so the modal has something to render. Verified by actually driving it in a real browser against a live container (not just code review): logged in, opened the modal, saw the new rows, minted a real join key end-to-end, no console errors. docs/site-join.md rewritten to cover live replication, signing-key sync, coordinated promotion/demote, and the new endpoints -- it previously only described the v2.2.0-v2.3.0 one-time-snapshot behavior.
This commit is contained in:
+75
-6
@@ -6,9 +6,7 @@ copy for local latency and autonomy (see the root `MULTI_SITE_SPEC.md` for the
|
||||
full architecture). This page covers the server endpoints that make a spoke
|
||||
"join" an existing master.
|
||||
|
||||
> Status: **server endpoints + UI + setup.sh wiring.** A fresh bring-up can
|
||||
> adopt a master directory via the Directory UI or via `setup.env`, and a
|
||||
> joined spoke is read-only with live WAN health.
|
||||
> Status: **server endpoints + UI + setup.sh wiring, live replication, coordinated promotion.** A fresh bring-up can adopt a master directory via the Directory UI or via `setup.env`; a joined spoke is read-only with live WAN health, stays in sync after joining (not just a one-time snapshot), and can be promoted to master with the old master demoted as part of the same action.
|
||||
|
||||
## The flow
|
||||
|
||||
@@ -19,13 +17,58 @@ full architecture). This page covers the server endpoints that make a spoke
|
||||
- **setup.sh**: set `CFG_MASTER_DIRECTORY_URL` + `CFG_MASTER_DIRECTORY_JOIN_KEY`
|
||||
in `setup.env` before the first run.
|
||||
3. The spoke pulls the master's directory export (LDAP tree + resource
|
||||
catalog), imports it, and persists its own spoke role
|
||||
catalog + agent-signing key), imports it, and persists its own spoke role
|
||||
(`isMaster: false`, `masterUrl`, `siteSlug`) in `/config/site.json`.
|
||||
4. If the spoke also knows its own reachable URL (`selfUrl` — `setup.sh` passes
|
||||
`https://$CFG_SSO_HOST` automatically), it registers itself with the master
|
||||
(`POST /api/site/spokes`) so the master can push live updates back to it
|
||||
afterward — see **Live replication** below. Without `selfUrl` the join still
|
||||
succeeds; the spoke just stays a one-time snapshot.
|
||||
|
||||
Joining is allowed only on a **fresh install** (no users beyond the bootstrap
|
||||
admin, no enrolled agents) — the join endpoint enforces this, so a populated
|
||||
directory can never be merged into a master's.
|
||||
|
||||
## Live replication (not a one-time snapshot)
|
||||
|
||||
A registered spoke stays in sync: every successful catalog write on the
|
||||
master fires a fire-and-forget push (`utils/site_replicate.js`) at every
|
||||
registered spoke, concurrently — one unreachable spoke never blocks or delays
|
||||
delivery to another. The spoke's `POST /api/site/resync` handler (called by
|
||||
that push) re-runs the same export-pull-and-import logic used at join time,
|
||||
so there's exactly one tested code path for "make my catalog match the
|
||||
master's," not a separate diff-application mechanism.
|
||||
|
||||
The agent-signing key travels the same path: `POST /api/site/export`
|
||||
best-effort includes it, and the spoke adopts it via `agent_keys.adopt()` on
|
||||
both join and every resync. Every site holding the same signing key means any
|
||||
site's `sso-manager-node` can validly sign a command for any agent enrolled
|
||||
at any other site — a deliberate tradeoff (see `MULTI_SITE_SPEC.md` §2)
|
||||
accepted for this deployment's small, trusted scale. Don't extend this
|
||||
pattern to a larger/adversarial-tenant deployment without revisiting it.
|
||||
|
||||
## Coordinated master promotion
|
||||
|
||||
`POST /api/directory-admin/site-promote` (`god_admin` only) promotes this
|
||||
node to master as **one coordinated action**, not a manual two-step
|
||||
demote-then-promote:
|
||||
|
||||
1. If this node currently has a master on file, it mints a fresh join key and
|
||||
calls that master's `POST /api/site/demote` (authenticated with the join
|
||||
key this node already holds), handing over the new key so the demoted node
|
||||
can keep talking to the new master afterward.
|
||||
2. This step is **best-effort** — an unreachable old master (the WAN-outage
|
||||
scenario this whole control exists for) never blocks the local promotion.
|
||||
The response's `handoff` field reports what happened
|
||||
(`"previous master demoted"`, an HTTP failure, or "unreachable, promoted
|
||||
locally anyway") so the operator can reconcile it manually if needed.
|
||||
3. Every known spoke gets a fire-and-forget `master-promoted` resync ping so
|
||||
they pick up the new master on their next sync.
|
||||
|
||||
The Master Site modal's **Promote to Master** button surfaces the `handoff`
|
||||
result in a toast so the operator sees immediately whether the old master was
|
||||
actually reached.
|
||||
|
||||
## Endpoints
|
||||
|
||||
| Method | Path | Purpose |
|
||||
@@ -35,9 +78,13 @@ directory can never be merged into a master's.
|
||||
| `POST` | `/api/site/join-keys/:id/revoke` | Stop it accepting new joins |
|
||||
| `DELETE` | `/api/site/join-keys/:id` | Remove it |
|
||||
| `GET` | `/api/site/config` | Current role (isMaster, masterUrl, siteSlug) |
|
||||
| `POST` | `/api/site/export` | Master directory export (Bearer `stj_` key) |
|
||||
| `POST` | `/api/site/export` | Master directory export incl. agent-signing key (Bearer `stj_` key) |
|
||||
| `POST` | `/api/site/ping` | Lightweight master reachability probe (Bearer `stj_` key) |
|
||||
| `POST` | `/api/site/join` | Adopt a master directory (admin session) |
|
||||
| `POST` | `/api/site/join` | Adopt a master directory + register for live replication (admin session) |
|
||||
| `POST` | `/api/site/spokes` | Register a spoke's endpoint for live replication (Bearer `stj_` key, called by the spoke right after join) |
|
||||
| `POST` | `/api/site/resync` | Re-pull the master's export (Bearer the spoke's own `pushToken`, called by the master's fire-and-forget push) |
|
||||
| `POST` | `/api/site/demote` | Step down to spoke of a new master (Bearer `stj_` key, called by the newly-promoted node) |
|
||||
| `POST` | `/api/directory-admin/site-promote` | Promote this node to master, coordinating demotion of the old one (`god_admin` session) |
|
||||
|
||||
## Behavior after joining (spoke)
|
||||
|
||||
@@ -74,3 +121,25 @@ already joined reports "already a spoke" and setup continues (idempotent).
|
||||
gated on both sides.
|
||||
- The join key is stored on the spoke only so it can reach the master for WAN
|
||||
health (and, in a later layer, write-proxy).
|
||||
- `pushToken` (the credential a spoke stores so it can recognize a legitimate
|
||||
resync push from its master) is minted fresh per spoke registration and, by
|
||||
design, kept in retrievable form on the master — unlike a join key, it's a
|
||||
credential the master must keep *presenting*, not just verifying, so it
|
||||
can't be one-way hashed. Compare `models/site_spoke.js`'s doc comment for
|
||||
why that's the correct tradeoff, not an oversight.
|
||||
- Every site sharing one agent-signing key (see **Live replication** above)
|
||||
means a compromised spoke — including the smallest, least-secured one — has
|
||||
the same agent-command authority as the master. Accepted for this
|
||||
deployment's scale; see `MULTI_SITE_SPEC.md` §2 before reusing this pattern
|
||||
somewhere that assumption doesn't hold.
|
||||
|
||||
## Not yet built
|
||||
|
||||
- Traffic between sites (join/export/resync) still goes over the open
|
||||
network path that already reaches the target — it does not route over the
|
||||
WireGuard mesh `theta-gateway` can now establish (see `MULTI_SITE_SPEC.md`).
|
||||
- A no-inbound spoke (no public IP at all) still can't join — the mechanism
|
||||
for a master to relay through the mesh to such a spoke is verified as
|
||||
working, but nothing automates creating that route yet.
|
||||
- OpenBao secret replication covers only the agent-signing key; LDAP admin
|
||||
creds, JWT secret, and other per-deployment secrets aren't synced.
|
||||
|
||||
Reference in New Issue
Block a user