Files
sso-manager-node/docs/site-join.md
T
wmantly dc3d760d2b feat(multi-site): UI for live replication, promotion handoff, signing key
Closes the gap where all of this session's new server-side capability
(live replication, coordinated promotion, identical signing keys) had
no UI at all -- an operator using the Master Site modal had no way to
know any of it existed or was working.

- Master Site modal: new "Live Replication" row (spoke) shows whether
  this join actually registered for live updates or is stuck on a
  one-time snapshot; new "Registered Spokes" row (master) shows how
  many spokes are receiving live pushes.
- Join form: new "this site's own reachable URL" field, prefilled from
  window.location.origin, wired to the selfUrl the join API already
  supported but the UI never sent -- a UI-driven join previously NEVER
  registered for live replication, only the setup.sh bootstrap path did.
  The success toast now reports whether live replication actually
  activated, not just "joined".
- Promote button: success toast now surfaces the handoff result (old
  master demoted / unreachable / no previous master), so the operator
  sees immediately whether the coordinated demotion actually happened.
- GET /api/site/config no longer returns masterJoinKey or
  replicationPushToken in the response -- found while wiring this up:
  live credentials were being sent straight to the browser for every
  admin session. Replaced with boolean derivatives
  (hasMasterJoinKey, liveReplication).
- GET /api/directory-admin/site-status gained liveReplication (spoke)
  and registeredSpokesCount (master) so the modal has something to render.

Verified by actually driving it in a real browser against a live
container (not just code review): logged in, opened the modal, saw the
new rows, minted a real join key end-to-end, no console errors.

docs/site-join.md rewritten to cover live replication, signing-key sync,
coordinated promotion/demote, and the new endpoints -- it previously only
described the v2.2.0-v2.3.0 one-time-snapshot behavior.
2026-08-10 18:30:48 -04:00

146 lines
8.3 KiB
Markdown

# Multi-Site: Joining a Spoke to the Master Directory
The Directory can be deployed across multiple sites. The **master** site holds
single write authority for the shared catalog; **spoke** sites run a read-only
copy for local latency and autonomy (see the root `MULTI_SITE_SPEC.md` for the
full architecture). This page covers the server endpoints that make a spoke
"join" an existing master.
> Status: **server endpoints + UI + setup.sh wiring, live replication, coordinated promotion.** A fresh bring-up can adopt a master directory via the Directory UI or via `setup.env`; a joined spoke is read-only with live WAN health, stays in sync after joining (not just a one-time snapshot), and can be promoted to master with the old master demoted as part of the same action.
## The flow
1. On the **master**, an admin mints a **site join key** (`stj_…`, shown once,
stored hashed, revocable) — Directory → the Master Site modal → **Site Join Keys**.
2. On the **spoke** (a fresh install), either:
- **UI**: Directory → the Master Site modal → **Join an Existing Site**, or
- **setup.sh**: set `CFG_MASTER_DIRECTORY_URL` + `CFG_MASTER_DIRECTORY_JOIN_KEY`
in `setup.env` before the first run.
3. The spoke pulls the master's directory export (LDAP tree + resource
catalog + agent-signing key), imports it, and persists its own spoke role
(`isMaster: false`, `masterUrl`, `siteSlug`) in `/config/site.json`.
4. If the spoke also knows its own reachable URL (`selfUrl``setup.sh` passes
`https://$CFG_SSO_HOST` automatically), it registers itself with the master
(`POST /api/site/spokes`) so the master can push live updates back to it
afterward — see **Live replication** below. Without `selfUrl` the join still
succeeds; the spoke just stays a one-time snapshot.
Joining is allowed only on a **fresh install** (no users beyond the bootstrap
admin, no enrolled agents) — the join endpoint enforces this, so a populated
directory can never be merged into a master's.
## Live replication (not a one-time snapshot)
A registered spoke stays in sync: every successful catalog write on the
master fires a fire-and-forget push (`utils/site_replicate.js`) at every
registered spoke, concurrently — one unreachable spoke never blocks or delays
delivery to another. The spoke's `POST /api/site/resync` handler (called by
that push) re-runs the same export-pull-and-import logic used at join time,
so there's exactly one tested code path for "make my catalog match the
master's," not a separate diff-application mechanism.
The agent-signing key travels the same path: `POST /api/site/export`
best-effort includes it, and the spoke adopts it via `agent_keys.adopt()` on
both join and every resync. Every site holding the same signing key means any
site's `sso-manager-node` can validly sign a command for any agent enrolled
at any other site — a deliberate tradeoff (see `MULTI_SITE_SPEC.md` §2)
accepted for this deployment's small, trusted scale. Don't extend this
pattern to a larger/adversarial-tenant deployment without revisiting it.
## Coordinated master promotion
`POST /api/directory-admin/site-promote` (`god_admin` only) promotes this
node to master as **one coordinated action**, not a manual two-step
demote-then-promote:
1. If this node currently has a master on file, it mints a fresh join key and
calls that master's `POST /api/site/demote` (authenticated with the join
key this node already holds), handing over the new key so the demoted node
can keep talking to the new master afterward.
2. This step is **best-effort** — an unreachable old master (the WAN-outage
scenario this whole control exists for) never blocks the local promotion.
The response's `handoff` field reports what happened
(`"previous master demoted"`, an HTTP failure, or "unreachable, promoted
locally anyway") so the operator can reconcile it manually if needed.
3. Every known spoke gets a fire-and-forget `master-promoted` resync ping so
they pick up the new master on their next sync.
The Master Site modal's **Promote to Master** button surfaces the `handoff`
result in a toast so the operator sees immediately whether the old master was
actually reached.
## Endpoints
| Method | Path | Purpose |
| :--- | :--- | :--- |
| `GET` | `/api/site/join-keys` | List keys (prefix + usage only; never the key) |
| `POST` | `/api/site/join-keys` | Mint one — returned **once** |
| `POST` | `/api/site/join-keys/:id/revoke` | Stop it accepting new joins |
| `DELETE` | `/api/site/join-keys/:id` | Remove it |
| `GET` | `/api/site/config` | Current role (isMaster, masterUrl, siteSlug) |
| `POST` | `/api/site/export` | Master directory export incl. agent-signing key (Bearer `stj_` key) |
| `POST` | `/api/site/ping` | Lightweight master reachability probe (Bearer `stj_` key) |
| `POST` | `/api/site/join` | Adopt a master directory + register for live replication (admin session) |
| `POST` | `/api/site/spokes` | Register a spoke's endpoint for live replication (Bearer `stj_` key, called by the spoke right after join) |
| `POST` | `/api/site/resync` | Re-pull the master's export (Bearer the spoke's own `pushToken`, called by the master's fire-and-forget push) |
| `POST` | `/api/site/demote` | Step down to spoke of a new master (Bearer `stj_` key, called by the newly-promoted node) |
| `POST` | `/api/directory-admin/site-promote` | Promote this node to master, coordinating demotion of the old one (`god_admin` session) |
## Behavior after joining (spoke)
- **Read-only**: directory-write requests (resources, edges, groups, secrets,
grants, driver actions, discovery merges) are rejected with `403` pointing at
the master. Writes must go to the master.
- **WAN health**: `site-status` pings the master over the stored site join key
and reports `wanConnected`; the Master Site modal shows live Online/Offline.
- **Role persists**: `isMaster`/`masterUrl`/`siteSlug` live in `/config/site.json`
(the env vars `IS_MASTER`/`MASTER_URL`/`SITE_SLUG` only seed the defaults), so
a restart never silently reverts a spoke to master.
## Deployment (setup.sh)
`setup.env` carries the intent so the join runs only on a **fresh** bring-up:
```
# Honored ONLY on first run; re-runs ignore it once ./config/ exists.
CFG_MASTER_DIRECTORY_URL=https://sso.master.example.com
CFG_MASTER_DIRECTORY_JOIN_KEY=stj_9f2e...
```
`setup.sh` runs `bootstrap/site-join.js` inside the sso-manager container after
the bootstrap; it logs in as the admin and calls `/api/site/join`. A node that
already joined reports "already a spoke" and setup continues (idempotent).
## Security
- Join keys are single-use-intent credentials: shown once, stored as a SHA-256
hash, revocable/expirable — the same model as agent join keys.
- The export/ping endpoints return only the directory tree/catalog (no admin
secrets) and require a valid join key.
- Join is admin-gated on the spoke, key-gated on the master, and fresh-install
gated on both sides.
- The join key is stored on the spoke only so it can reach the master for WAN
health (and, in a later layer, write-proxy).
- `pushToken` (the credential a spoke stores so it can recognize a legitimate
resync push from its master) is minted fresh per spoke registration and, by
design, kept in retrievable form on the master — unlike a join key, it's a
credential the master must keep *presenting*, not just verifying, so it
can't be one-way hashed. Compare `models/site_spoke.js`'s doc comment for
why that's the correct tradeoff, not an oversight.
- Every site sharing one agent-signing key (see **Live replication** above)
means a compromised spoke — including the smallest, least-secured one — has
the same agent-command authority as the master. Accepted for this
deployment's scale; see `MULTI_SITE_SPEC.md` §2 before reusing this pattern
somewhere that assumption doesn't hold.
## Not yet built
- Traffic between sites (join/export/resync) still goes over the open
network path that already reaches the target — it does not route over the
WireGuard mesh `theta-gateway` can now establish (see `MULTI_SITE_SPEC.md`).
- A no-inbound spoke (no public IP at all) still can't join — the mechanism
for a master to relay through the mesh to such a spoke is verified as
working, but nothing automates creating that route yet.
- OpenBao secret replication covers only the agent-signing key; LDAP admin
creds, JWT secret, and other per-deployment secrets aren't synced.