Merge pull request #197 from theta42/release-v2.4.0
Build OpenLDAP Base Image / build-and-push (push) Failing after 10s
Build OpenLDAP Base Image / build-and-push (push) Failing after 10s
release(v2.4.0): live replication, coordinated promotion, mesh UI
This commit is contained in:
@@ -1,3 +1,17 @@
|
||||
# v2.4.0 - 2026-08-10
|
||||
|
||||
### Added
|
||||
- **Live catalog replication.** A spoke now stays in sync after joining instead of only getting a one-time snapshot: it registers its own endpoint with the master at join time (`POST /api/site/spokes`, Bearer the site join key), and every successful master catalog write fires a fire-and-forget push (`utils/site_replicate.js`) at every registered spoke, concurrently — one unreachable spoke never blocks or delays delivery to another. The spoke's `POST /api/site/resync` handler re-runs the same tested export-pull-and-import path used at join time rather than applying a partial diff.
|
||||
- **Identical-directory agent-signing key.** `POST /api/site/export` now best-effort includes the master's agent-signing key; a spoke adopts it via `agent_keys.adopt()` on both join and every resync, so any site's `sso-manager-node` can validly sign a command for any agent enrolled at any other site (a deliberate blast-radius tradeoff for this deployment's small, trusted scale — see `theta-suite`'s `docs/MULTI_SITE_SPEC.md` §2).
|
||||
- **Coordinated master promotion.** `POST /api/directory-admin/site-promote` now demotes the previous master as part of the same action (mints it a fresh join key, calls its new `POST /api/site/demote`) instead of leaving a manual two-step gap where two nodes could both believe they're master. Best-effort: an unreachable old master (the WAN-outage scenario this control exists for) never blocks the local promotion — the response's `handoff` field reports what happened.
|
||||
- **Master Site modal UI**: new "Live Replication" (spoke) / "Registered Spokes" (master) status rows; the join form gained a "this site's own reachable URL" field (prefilled from the browser origin) wired to the `selfUrl` the join API already supported but the UI never sent — a UI-driven join previously never registered for live replication, only the `setup.sh` bootstrap path did; the promote button's success toast now reports the actual handoff result.
|
||||
- New `nodejs/models/site_spoke.js` (registered spokes + their push tokens) and `docker-compose.multisite-e2e.yml` + `test/multisite_join_e2e.js` (real two-container master+spoke regression test covering join, live replication, promotion, and demotion end to end).
|
||||
|
||||
### Fixed
|
||||
- **`site-promote`'s god_admin check was dead on arrival.** It read `req.user.groups`, a field nothing in the codebase ever populates (every other admin gate resolves membership live via `permission.byGroup()`/`Group.list(user.dn)`, which also handles nested-group membership) — the check silently evaluated to an empty array on every request, so promotion returned 403 for every user, including a real god_admin, since it shipped in v2.0.0. Only surfaced by the live e2e test, not by inspection.
|
||||
- **The read-only write-gate blocked `site-promote` on a spoke before its handler could run** — the one mutating request a spoke must be able to make to itself. Exempted `/site-promote` from the gate.
|
||||
- **`GET /api/site/config` was returning `masterJoinKey` and `replicationPushToken`** — live credentials — directly in the JSON response to any admin session. Replaced with boolean derivatives (`hasMasterJoinKey`, `liveReplication`).
|
||||
|
||||
# v2.3.0 - 2026-08-10
|
||||
|
||||
### Added
|
||||
|
||||
@@ -50,6 +50,8 @@ COPY test/seed-test-user.sh /usr/local/bin/seed-test-user
|
||||
RUN chmod +x /usr/local/bin/seed-test-user
|
||||
# End-to-end LDAP tunnel test client (docker-compose.e2e.yml)
|
||||
COPY test/tunnel_e2e.js ./test/tunnel_e2e.js
|
||||
# End-to-end multi-site join test client (docker-compose.multisite-e2e.yml)
|
||||
COPY test/multisite_join_e2e.js ./test/multisite_join_e2e.js
|
||||
|
||||
# Default command: seed the test user, then run the test suite
|
||||
CMD ["sh", "-c", "seed-test-user && npm test"]
|
||||
|
||||
@@ -399,3 +399,7 @@ The Directory incorporates a **4-tier Driver Resolution Engine** (`services/driv
|
||||
- `POST /api/directory-admin/resources/:id/driver-action` — Execute management action (`{ action, params }`)
|
||||
- `GET /api/directory-admin/resources/:id/driver-logs` — Tail log output (`?lines=100`)
|
||||
|
||||
## 11. Multi-Site
|
||||
|
||||
This directory can run across multiple sites (one **master** with write authority, any number of **spoke** read-only replicas that stay live-synced after joining), coordinate master promotion, and share the agent-signing key across sites. Full design and operational detail: [`docs/site-join.md`](docs/site-join.md) and, at the suite level, `theta-suite`'s `docs/MULTI_SITE_SPEC.md`.
|
||||
|
||||
|
||||
@@ -0,0 +1,69 @@
|
||||
# End-to-end test of the real, shipped multi-site join flow (docs/site-join.md).
|
||||
#
|
||||
# Spins up two full all-in-one instances (app + bundled slapd each, like
|
||||
# docker-compose.repl-test.yml) — "master" and "spoke" — plus a client that
|
||||
# drives the actual HTTP API a human/operator would use: mint a site join key
|
||||
# on master, join from spoke, verify the spoke adopted the catalog, went
|
||||
# read-only, and reports live WAN health.
|
||||
#
|
||||
# docker compose -f docker-compose.multisite-e2e.yml up --build --abort-on-container-exit
|
||||
# # exit code 0 = MULTISITE E2E PASS
|
||||
#
|
||||
# slapcat (used by POST /api/site/export) only sees the LDAP data of the
|
||||
# container it runs in, so this MUST use the all-in-one image (master and
|
||||
# spoke each carry their own slapd) — the split ldap+redis+app harness used
|
||||
# by docker-compose.test.yml/e2e.yml won't exercise export/join at all.
|
||||
|
||||
services:
|
||||
master:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile.openldap
|
||||
container_name: multisite_e2e_master
|
||||
environment:
|
||||
- LDAP_BASE_DN=dc=master,dc=test
|
||||
- LDAP_ADMIN_PASS=secret
|
||||
- ORG_NAME=E2E Master
|
||||
- app_oauth__jwtSecret=e2e-multisite-master-jwt-secret
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "wget -qO- http://localhost:3001/health >/dev/null 2>&1"]
|
||||
interval: 2s
|
||||
timeout: 3s
|
||||
retries: 40
|
||||
start_period: 5s
|
||||
|
||||
spoke:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile.openldap
|
||||
container_name: multisite_e2e_spoke
|
||||
environment:
|
||||
- LDAP_BASE_DN=dc=spoke,dc=test
|
||||
- LDAP_ADMIN_PASS=secret
|
||||
- ORG_NAME=E2E Spoke
|
||||
- app_oauth__jwtSecret=e2e-multisite-spoke-jwt-secret
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "wget -qO- http://localhost:3001/health >/dev/null 2>&1"]
|
||||
interval: 2s
|
||||
timeout: 3s
|
||||
retries: 40
|
||||
start_period: 5s
|
||||
|
||||
client:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile.test-runner
|
||||
command: ["sh", "-c", "node test/multisite_join_e2e.js"]
|
||||
environment:
|
||||
- MASTER_URL=http://master:3001
|
||||
- SPOKE_URL=http://spoke:3001
|
||||
- MASTER_LDAP_HOST=master
|
||||
- MASTER_BASE_DN=dc=master,dc=test
|
||||
- SPOKE_LDAP_HOST=spoke
|
||||
- SPOKE_BASE_DN=dc=spoke,dc=test
|
||||
- LDAP_ADMIN_PASS=secret
|
||||
depends_on:
|
||||
master:
|
||||
condition: service_healthy
|
||||
spoke:
|
||||
condition: service_healthy
|
||||
+75
-6
@@ -6,9 +6,7 @@ copy for local latency and autonomy (see the root `MULTI_SITE_SPEC.md` for the
|
||||
full architecture). This page covers the server endpoints that make a spoke
|
||||
"join" an existing master.
|
||||
|
||||
> Status: **server endpoints + UI + setup.sh wiring.** A fresh bring-up can
|
||||
> adopt a master directory via the Directory UI or via `setup.env`, and a
|
||||
> joined spoke is read-only with live WAN health.
|
||||
> Status: **server endpoints + UI + setup.sh wiring, live replication, coordinated promotion.** A fresh bring-up can adopt a master directory via the Directory UI or via `setup.env`; a joined spoke is read-only with live WAN health, stays in sync after joining (not just a one-time snapshot), and can be promoted to master with the old master demoted as part of the same action.
|
||||
|
||||
## The flow
|
||||
|
||||
@@ -19,13 +17,58 @@ full architecture). This page covers the server endpoints that make a spoke
|
||||
- **setup.sh**: set `CFG_MASTER_DIRECTORY_URL` + `CFG_MASTER_DIRECTORY_JOIN_KEY`
|
||||
in `setup.env` before the first run.
|
||||
3. The spoke pulls the master's directory export (LDAP tree + resource
|
||||
catalog), imports it, and persists its own spoke role
|
||||
catalog + agent-signing key), imports it, and persists its own spoke role
|
||||
(`isMaster: false`, `masterUrl`, `siteSlug`) in `/config/site.json`.
|
||||
4. If the spoke also knows its own reachable URL (`selfUrl` — `setup.sh` passes
|
||||
`https://$CFG_SSO_HOST` automatically), it registers itself with the master
|
||||
(`POST /api/site/spokes`) so the master can push live updates back to it
|
||||
afterward — see **Live replication** below. Without `selfUrl` the join still
|
||||
succeeds; the spoke just stays a one-time snapshot.
|
||||
|
||||
Joining is allowed only on a **fresh install** (no users beyond the bootstrap
|
||||
admin, no enrolled agents) — the join endpoint enforces this, so a populated
|
||||
directory can never be merged into a master's.
|
||||
|
||||
## Live replication (not a one-time snapshot)
|
||||
|
||||
A registered spoke stays in sync: every successful catalog write on the
|
||||
master fires a fire-and-forget push (`utils/site_replicate.js`) at every
|
||||
registered spoke, concurrently — one unreachable spoke never blocks or delays
|
||||
delivery to another. The spoke's `POST /api/site/resync` handler (called by
|
||||
that push) re-runs the same export-pull-and-import logic used at join time,
|
||||
so there's exactly one tested code path for "make my catalog match the
|
||||
master's," not a separate diff-application mechanism.
|
||||
|
||||
The agent-signing key travels the same path: `POST /api/site/export`
|
||||
best-effort includes it, and the spoke adopts it via `agent_keys.adopt()` on
|
||||
both join and every resync. Every site holding the same signing key means any
|
||||
site's `sso-manager-node` can validly sign a command for any agent enrolled
|
||||
at any other site — a deliberate tradeoff (see `MULTI_SITE_SPEC.md` §2)
|
||||
accepted for this deployment's small, trusted scale. Don't extend this
|
||||
pattern to a larger/adversarial-tenant deployment without revisiting it.
|
||||
|
||||
## Coordinated master promotion
|
||||
|
||||
`POST /api/directory-admin/site-promote` (`god_admin` only) promotes this
|
||||
node to master as **one coordinated action**, not a manual two-step
|
||||
demote-then-promote:
|
||||
|
||||
1. If this node currently has a master on file, it mints a fresh join key and
|
||||
calls that master's `POST /api/site/demote` (authenticated with the join
|
||||
key this node already holds), handing over the new key so the demoted node
|
||||
can keep talking to the new master afterward.
|
||||
2. This step is **best-effort** — an unreachable old master (the WAN-outage
|
||||
scenario this whole control exists for) never blocks the local promotion.
|
||||
The response's `handoff` field reports what happened
|
||||
(`"previous master demoted"`, an HTTP failure, or "unreachable, promoted
|
||||
locally anyway") so the operator can reconcile it manually if needed.
|
||||
3. Every known spoke gets a fire-and-forget `master-promoted` resync ping so
|
||||
they pick up the new master on their next sync.
|
||||
|
||||
The Master Site modal's **Promote to Master** button surfaces the `handoff`
|
||||
result in a toast so the operator sees immediately whether the old master was
|
||||
actually reached.
|
||||
|
||||
## Endpoints
|
||||
|
||||
| Method | Path | Purpose |
|
||||
@@ -35,9 +78,13 @@ directory can never be merged into a master's.
|
||||
| `POST` | `/api/site/join-keys/:id/revoke` | Stop it accepting new joins |
|
||||
| `DELETE` | `/api/site/join-keys/:id` | Remove it |
|
||||
| `GET` | `/api/site/config` | Current role (isMaster, masterUrl, siteSlug) |
|
||||
| `POST` | `/api/site/export` | Master directory export (Bearer `stj_` key) |
|
||||
| `POST` | `/api/site/export` | Master directory export incl. agent-signing key (Bearer `stj_` key) |
|
||||
| `POST` | `/api/site/ping` | Lightweight master reachability probe (Bearer `stj_` key) |
|
||||
| `POST` | `/api/site/join` | Adopt a master directory (admin session) |
|
||||
| `POST` | `/api/site/join` | Adopt a master directory + register for live replication (admin session) |
|
||||
| `POST` | `/api/site/spokes` | Register a spoke's endpoint for live replication (Bearer `stj_` key, called by the spoke right after join) |
|
||||
| `POST` | `/api/site/resync` | Re-pull the master's export (Bearer the spoke's own `pushToken`, called by the master's fire-and-forget push) |
|
||||
| `POST` | `/api/site/demote` | Step down to spoke of a new master (Bearer `stj_` key, called by the newly-promoted node) |
|
||||
| `POST` | `/api/directory-admin/site-promote` | Promote this node to master, coordinating demotion of the old one (`god_admin` session) |
|
||||
|
||||
## Behavior after joining (spoke)
|
||||
|
||||
@@ -74,3 +121,25 @@ already joined reports "already a spoke" and setup continues (idempotent).
|
||||
gated on both sides.
|
||||
- The join key is stored on the spoke only so it can reach the master for WAN
|
||||
health (and, in a later layer, write-proxy).
|
||||
- `pushToken` (the credential a spoke stores so it can recognize a legitimate
|
||||
resync push from its master) is minted fresh per spoke registration and, by
|
||||
design, kept in retrievable form on the master — unlike a join key, it's a
|
||||
credential the master must keep *presenting*, not just verifying, so it
|
||||
can't be one-way hashed. Compare `models/site_spoke.js`'s doc comment for
|
||||
why that's the correct tradeoff, not an oversight.
|
||||
- Every site sharing one agent-signing key (see **Live replication** above)
|
||||
means a compromised spoke — including the smallest, least-secured one — has
|
||||
the same agent-command authority as the master. Accepted for this
|
||||
deployment's scale; see `MULTI_SITE_SPEC.md` §2 before reusing this pattern
|
||||
somewhere that assumption doesn't hold.
|
||||
|
||||
## Not yet built
|
||||
|
||||
- Traffic between sites (join/export/resync) still goes over the open
|
||||
network path that already reaches the target — it does not route over the
|
||||
WireGuard mesh `theta-gateway` can now establish (see `MULTI_SITE_SPEC.md`).
|
||||
- A no-inbound spoke (no public IP at all) still can't join — the mechanism
|
||||
for a master to relay through the mesh to such a spoke is verified as
|
||||
working, but nothing automates creating that route yet.
|
||||
- OpenBao secret replication covers only the agent-signing key; LDAP admin
|
||||
creds, JWT secret, and other per-deployment secrets aren't synced.
|
||||
|
||||
@@ -22,6 +22,7 @@ const { SharedSecretGrant } = require('./shared_secret_grant');
|
||||
const { VaultAppToken } = require('./vault_app_token');
|
||||
const { Agent, AgentJoinKey } = require('./agent');
|
||||
const { SiteJoinKey } = require('./site_join_key');
|
||||
const { SiteSpoke } = require('./site_spoke');
|
||||
async function initORM() {
|
||||
const ormConf = conf.orm || {
|
||||
dialect: 'sqlite',
|
||||
@@ -36,7 +37,7 @@ async function initORM() {
|
||||
conf: { orm: ormConf },
|
||||
models: [
|
||||
Resource, ResourceEdge, ResourceGroup, AccessRequest, Webhook, PluginInstance,
|
||||
SharedSecret, SharedSecretGrant, VaultAppToken, Agent, AgentJoinKey, SiteJoinKey,
|
||||
SharedSecret, SharedSecretGrant, VaultAppToken, Agent, AgentJoinKey, SiteJoinKey, SiteSpoke,
|
||||
Token, AuthToken, InviteToken, ImpersonationToken, PasswordResetToken, OtpToken, ServiceToken
|
||||
]
|
||||
});
|
||||
|
||||
@@ -0,0 +1,42 @@
|
||||
'use strict';
|
||||
|
||||
const crypto = require('crypto');
|
||||
const { Model } = require('@simpleworkjs/orm');
|
||||
|
||||
// A spoke known to THIS node while it's acting as master — the registry that
|
||||
// makes live replication possible. A spoke registers itself here (POST
|
||||
// /api/site/spokes, authenticated by the same join key it used to join)
|
||||
// right after adopting the master's export, handing over its own reachable
|
||||
// endpoint. In return it's issued a `pushToken`: a shared secret the master
|
||||
// then presents on every future POST <spoke endpoint>/api/site/resync call.
|
||||
//
|
||||
// This is a DIFFERENT credential direction than SiteJoinKey: a join key is
|
||||
// presented TO the master and only ever needs to be verified (so it's stored
|
||||
// hashed, like a password). pushToken is presented BY the master, repeatedly,
|
||||
// so it has to be retrievable here -- there is no getting around storing it
|
||||
// in plaintext on the master, the same way Webhook.secret is (see
|
||||
// services/webhook_emitter.js) for the same reason (an HMAC/bearer credential
|
||||
// the sender must keep re-presenting, not a one-time secret only ever
|
||||
// verified).
|
||||
class SiteSpoke extends Model {
|
||||
static generatePushToken() {
|
||||
return crypto.randomBytes(24).toString('base64url');
|
||||
}
|
||||
|
||||
static fields = {
|
||||
id: { type: 'uuid', primaryKey: true },
|
||||
endpoint: { type: 'string', isRequired: true, unique: true },
|
||||
siteSlug: { type: 'string' },
|
||||
pushToken: { type: 'string', isRequired: true },
|
||||
created_on: { type: 'integer' },
|
||||
last_seen_on: { type: 'integer' }
|
||||
};
|
||||
|
||||
toPublic() {
|
||||
const data = this.toJSON ? this.toJSON() : { ...this };
|
||||
delete data.pushToken;
|
||||
return data;
|
||||
}
|
||||
}
|
||||
|
||||
module.exports = { SiteSpoke };
|
||||
+2
-2
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "t42-theta-directory",
|
||||
"version": "2.3.0",
|
||||
"version": "2.4.0",
|
||||
"description": "A very simple LDAP management and SSO system",
|
||||
"author": [
|
||||
{
|
||||
@@ -11,7 +11,7 @@
|
||||
"scripts": {
|
||||
"start": "node ./bin/www",
|
||||
"dev": "npx nodemon --ignore public/ ./bin/www",
|
||||
"test": "NODE_ENV=test jest tests/groups.test.js tests/subtypes.test.js tests/site_join.test.js tests/site_config.test.js --forceExit"
|
||||
"test": "NODE_ENV=test jest tests/groups.test.js tests/subtypes.test.js tests/site_join.test.js tests/site_config.test.js tests/site_replicate.test.js --forceExit"
|
||||
},
|
||||
"jest": {
|
||||
"testEnvironment": "node",
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
const router = require('express').Router();
|
||||
const permission = require('../utils/permission');
|
||||
const { Resource, ResourceEdge, ResourceGroup } = require('../models/resource');
|
||||
const { SiteJoinKey } = require('../models/site_join_key');
|
||||
const { Group } = require('../models/group_ldap');
|
||||
const { User } = require('../models/user_ldap');
|
||||
const { cnFromDn } = require('../utils/user_groups');
|
||||
@@ -9,6 +10,7 @@ const { projectResources } = require('@simpleworkjs/directory-schema');
|
||||
|
||||
const SUPER_ADMIN_GROUP = permission.SUPER_ADMIN_GROUP;
|
||||
const groups = require('../utils/groups');
|
||||
const meshReplicate = require('../utils/site_replicate');
|
||||
|
||||
// Make `childCn` a member of `parentCn`, i.e. everyone in the child is
|
||||
// transitively in the parent. Idempotent and non-fatal: "already a member" is
|
||||
@@ -252,19 +254,38 @@ router.get('/resources', async (req, res, next) => {
|
||||
} catch (err) { next(err); }
|
||||
});
|
||||
|
||||
// ── Spoke read-only enforcement ─────────────────────────────────────────────
|
||||
// ── Spoke read-only enforcement + live replication trigger ──────────────────
|
||||
// On a joined spoke the catalog is a copy of the master's; directory writes
|
||||
// must go to the master (MULTI_SITE_SPEC.md — spoke = read-only catalog). Any
|
||||
// mutating request below this point is rejected on a spoke with a pointer to
|
||||
// the master. (site-status / site-promote live AFTER this middleware and are
|
||||
// not directory writes.)
|
||||
//
|
||||
// On the MASTER, a successful mutation here fires a fire-and-forget resync
|
||||
// push (utils/site_replicate.js) at every registered spoke, so the shipped
|
||||
// join flow's one-time snapshot doesn't go stale the moment the catalog
|
||||
// changes. Fires on res.on('finish') (after the response is actually sent,
|
||||
// status known) rather than before the handler runs, so a write that fails
|
||||
// validation never triggers a pointless replication round-trip.
|
||||
// /site-promote is deliberately exempt below: it's the ONE mutating request a
|
||||
// spoke must be able to make to itself (that's the entire point -- a spoke
|
||||
// promoting itself to master). Without this exemption the gate 403s the
|
||||
// promotion request before it ever reaches the handler, since this
|
||||
// middleware is registered ahead of router.post('/site-promote', ...) later
|
||||
// in the file and Express matches router.use() against every path.
|
||||
router.use((req, res, next) => {
|
||||
if (['POST', 'PUT', 'DELETE', 'PATCH'].includes(req.method)) {
|
||||
const mutating = ['POST', 'PUT', 'DELETE', 'PATCH'].includes(req.method);
|
||||
if (mutating && req.path !== '/site-promote') {
|
||||
const cfg = siteConfig.get();
|
||||
if (!cfg.isMaster) {
|
||||
const hint = cfg.masterUrl ? ' Directory writes must go to the master at ' + cfg.masterUrl + '.' : '';
|
||||
return res.status(403).json({ status: 'error', message: 'This node is a spoke (read-only catalog).' + hint });
|
||||
}
|
||||
res.on('finish', () => {
|
||||
if (res.statusCode >= 200 && res.statusCode < 300) {
|
||||
meshReplicate.replicateToSpokes(`${req.method} ${req.path}`);
|
||||
}
|
||||
});
|
||||
}
|
||||
next();
|
||||
});
|
||||
@@ -906,6 +927,7 @@ router.post('/discovered/merge', async (req, res, next) => {
|
||||
const siteConfig = require('../utils/site_config');
|
||||
const { siteIsFresh } = require('../utils/site_join');
|
||||
const { Agent } = require('../models/agent');
|
||||
const { SiteSpoke } = require('../models/site_spoke');
|
||||
|
||||
// probeMasterHealth checks whether this (spoke) node can reach its master over
|
||||
// the site join key. The master's /api/site/ping is deliberately lightweight.
|
||||
@@ -941,6 +963,13 @@ router.get('/site-status', async (req, res, next) => {
|
||||
if (cfg.isMaster) {
|
||||
canJoin = await siteIsFresh({ User, Agent }).catch(() => false);
|
||||
}
|
||||
// registeredSpokesCount (master) / liveReplication (spoke): surfaces
|
||||
// whether live replication is actually wired up, not just whether the
|
||||
// join itself succeeded -- a spoke that joined without `selfUrl` (e.g.
|
||||
// via an older bootstrap, or the UI form before it grew the field) is
|
||||
// fully joined but silently stuck on the one-time snapshot, which was
|
||||
// otherwise invisible anywhere in the UI.
|
||||
const registeredSpokesCount = cfg.isMaster ? await SiteSpoke.list().then(l => l.length).catch(() => 0) : 0;
|
||||
res.json({
|
||||
status: 'ok',
|
||||
config: {
|
||||
@@ -949,7 +978,9 @@ router.get('/site-status', async (req, res, next) => {
|
||||
siteSlug: cfg.siteSlug,
|
||||
wanConnected,
|
||||
siteMode: cfg.isMaster ? 'master' : 'spoke',
|
||||
canJoin
|
||||
canJoin,
|
||||
liveReplication: !cfg.isMaster ? !!cfg.replicationPushToken : undefined,
|
||||
registeredSpokesCount
|
||||
},
|
||||
sitesCount: sites.length,
|
||||
sites: sites.map(s => ({ id: s.id, name: s.name, slug: s.slug })),
|
||||
@@ -960,21 +991,67 @@ router.get('/site-status', async (req, res, next) => {
|
||||
|
||||
router.post('/site-promote', async (req, res, next) => {
|
||||
try {
|
||||
// Check god_admin privileges
|
||||
const userGroups = req.user && req.user.groups ? req.user.groups : [];
|
||||
const isGodAdmin = userGroups.includes('god_admin') || userGroups.includes(SUPER_ADMIN_GROUP);
|
||||
// god_admin privilege check. This used to read req.user.groups, which
|
||||
// nothing in the codebase ever populates -- User.get() (what
|
||||
// Auth.checkToken returns as req.user) has no .groups field; every other
|
||||
// admin gate in this app resolves membership live via
|
||||
// permission.byGroup()/Group.list(user.dn), which also correctly
|
||||
// resolves NESTED group membership (a user who is god_admin via a nested
|
||||
// group, not just direct membership). The old check silently evaluated
|
||||
// to an empty array for every request, making this endpoint
|
||||
// unreachable for ANY user -- caught by the multi-site e2e promotion
|
||||
// test (docker-compose.multisite-e2e.yml), not by inspection.
|
||||
const isGodAdmin = await permission.byGroup(req.user, [SUPER_ADMIN_GROUP]).catch(() => false);
|
||||
if (!isGodAdmin) {
|
||||
return res.status(403).json({ status: 'error', message: 'Master promotion requires explicit god_admin authority' });
|
||||
}
|
||||
|
||||
siteConfig.save({ isMaster: true, masterUrl: '' });
|
||||
// MULTI_SITE_SPEC.md §3.2: promotion is ONE coordinated action, never a
|
||||
// manual two-step "demote the old one first" — if we currently know a
|
||||
// master (we were a spoke), hand it off before flipping ourselves. This
|
||||
// is best-effort: an unreachable old master (the whole point of the
|
||||
// WAN-outage promotion scenario §3 describes) must never block a
|
||||
// god_admin's local promotion, it's just reported so the operator can
|
||||
// reconcile it manually.
|
||||
const beforeCfg = siteConfig.get();
|
||||
let handoffNote = 'no previous master on file (already master, or fresh install)';
|
||||
if (!beforeCfg.isMaster && beforeCfg.masterUrl && beforeCfg.masterJoinKey) {
|
||||
try {
|
||||
const { raw: freshKey } = await SiteJoinKey.issue({
|
||||
label: 'promotion-handoff-' + new Date().toISOString().slice(0, 10),
|
||||
createdBy: req.user ? req.user.uid : 'admin'
|
||||
});
|
||||
const controller = new AbortController();
|
||||
const timer = setTimeout(() => controller.abort(), 15000);
|
||||
let resp;
|
||||
try {
|
||||
resp = await fetch(beforeCfg.masterUrl + '/api/site/demote', {
|
||||
method: 'POST',
|
||||
headers: { Authorization: 'Bearer ' + beforeCfg.masterJoinKey, 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ newMasterUrl: (req.body && req.body.selfUrl) || '', newJoinKey: freshKey }),
|
||||
signal: controller.signal
|
||||
});
|
||||
} finally { clearTimeout(timer); }
|
||||
handoffNote = resp.ok ? 'previous master demoted' : ('previous master demote failed: HTTP ' + resp.status);
|
||||
} catch (e) {
|
||||
handoffNote = 'previous master unreachable (' + e.message + ') — promoted locally anyway; reconcile it manually once it\'s back';
|
||||
}
|
||||
}
|
||||
|
||||
console.log(`[MULTI-SITE] Node promoted to MASTER by user ${req.user ? req.user.uid : 'admin'}`);
|
||||
siteConfig.save({ isMaster: true, masterUrl: '', masterJoinKey: undefined });
|
||||
|
||||
console.log(`[MULTI-SITE] Node promoted to MASTER by user ${req.user ? req.user.uid : 'admin'} (handoff: ${handoffNote})`);
|
||||
|
||||
// Fire-and-forget: let every known spoke know a new master exists so
|
||||
// their next resync targets it. (They'll also learn this the hard way if
|
||||
// their old-master resync calls start failing, but this speeds it up.)
|
||||
meshReplicate.replicateToSpokes('master-promoted');
|
||||
|
||||
const cfg = siteConfig.get();
|
||||
res.json({
|
||||
status: 'ok',
|
||||
message: 'Node successfully promoted to Master Site',
|
||||
handoff: handoffNote,
|
||||
config: {
|
||||
isMaster: true,
|
||||
masterUrl: '',
|
||||
|
||||
+245
-52
@@ -14,6 +14,7 @@
|
||||
// calls it with a join key), so it is defined BEFORE the auth middleware.
|
||||
|
||||
const express = require('express');
|
||||
const crypto = require('crypto');
|
||||
const { execFile } = require('child_process');
|
||||
const { promisify } = require('util');
|
||||
const os = require('os');
|
||||
@@ -25,10 +26,13 @@ const permission = require('../utils/permission');
|
||||
const conf = require('@simpleworkjs/conf');
|
||||
const { Resource, ResourceEdge } = require('../models/resource');
|
||||
const { SiteJoinKey } = require('../models/site_join_key');
|
||||
const { SiteSpoke } = require('../models/site_spoke');
|
||||
const { replicateToSpokes } = require('../utils/site_replicate');
|
||||
const User = require('../models/user');
|
||||
const { Agent } = require('../models/agent');
|
||||
const siteConfig = require('../utils/site_config');
|
||||
const { importDirectory, ldapAddArgs, baseDnFrom, siteIsFresh } = require('../utils/site_join');
|
||||
const agentKeys = require('../utils/agent_keys');
|
||||
|
||||
const execFileAsync = promisify(execFile);
|
||||
const router = express.Router();
|
||||
@@ -63,10 +67,16 @@ router.post('/export', async (req, res, next) => {
|
||||
const key = await SiteJoinKey.authenticate(rawKey);
|
||||
if (!key) return res.status(401).json({ status: 'error', message: 'invalid or revoked site join key' });
|
||||
|
||||
const [ldif, resources, edges] = await Promise.all([
|
||||
const [ldif, resources, edges, signingKey] = await Promise.all([
|
||||
slurpLdif(),
|
||||
Resource.list(),
|
||||
ResourceEdge.list()
|
||||
ResourceEdge.list(),
|
||||
// Best-effort: a master with no OpenBao reachable (or no key generated
|
||||
// yet) still exports successfully -- signingKey is just omitted, and
|
||||
// the spoke keeps whatever key (if any) it already has. Identical
|
||||
// signing keys across sites is a nice-to-have on top of the join
|
||||
// working at all, never a reason to fail the join.
|
||||
agentKeys.load().then((k) => k && { privateKeyPem: k.privateKeyPem, publicKeyPem: k.publicKeyPem }).catch(() => null)
|
||||
]);
|
||||
|
||||
await key.update({ use_count: (key.use_count || 0) + 1, last_used_on: Math.floor(Date.now() / 1000) }).catch(() => {});
|
||||
@@ -77,7 +87,8 @@ router.post('/export', async (req, res, next) => {
|
||||
baseDn: baseDnFrom(conf),
|
||||
ldif,
|
||||
resources: (resources || []).map(r => (r.toJSON ? r.toJSON() : r)),
|
||||
edges: (edges || []).map(e => (e.toJSON ? e.toJSON() : e))
|
||||
edges: (edges || []).map(e => (e.toJSON ? e.toJSON() : e)),
|
||||
...(signingKey ? { signingKey } : {})
|
||||
});
|
||||
} catch (e) { next(e); }
|
||||
});
|
||||
@@ -95,6 +106,104 @@ router.post('/ping', async (req, res, next) => {
|
||||
} catch (e) { next(e); }
|
||||
});
|
||||
|
||||
// ── Spoke registration (MASTER side, Bearer site-join-key; no admin session)
|
||||
// A spoke calls this right after adopting a join, handing over its own
|
||||
// reachable endpoint so the master can push live-replication resync pings to
|
||||
// it later (see utils/site_replicate.js). Idempotent on endpoint: calling it
|
||||
// again (e.g. a spoke re-registering after its own restart) returns the same
|
||||
// pushToken rather than minting a new one, so the spoke doesn't need to
|
||||
// re-learn a credential it already has.
|
||||
router.post('/spokes', async (req, res, next) => {
|
||||
try {
|
||||
const auth = req.headers.authorization || '';
|
||||
const rawKey = auth.startsWith('Bearer ') ? auth.slice(7).trim() : '';
|
||||
const key = await SiteJoinKey.authenticate(rawKey);
|
||||
if (!key) return res.status(401).json({ status: 'error', message: 'invalid or revoked site join key' });
|
||||
|
||||
const { endpoint, siteSlug } = req.body || {};
|
||||
if (!endpoint || !/^https?:\/\//.test(endpoint)) {
|
||||
return res.status(400).json({ status: 'error', message: 'a valid http(s) endpoint is required' });
|
||||
}
|
||||
|
||||
const now = Math.floor(Date.now() / 1000);
|
||||
let spoke = (await SiteSpoke.list({ where: { endpoint } }))[0];
|
||||
if (spoke) {
|
||||
await spoke.update({ siteSlug: siteSlug || spoke.siteSlug, last_seen_on: now });
|
||||
} else {
|
||||
spoke = await SiteSpoke.create({
|
||||
id: crypto.randomUUID(),
|
||||
endpoint,
|
||||
siteSlug: siteSlug || null,
|
||||
pushToken: SiteSpoke.generatePushToken(),
|
||||
created_on: now,
|
||||
last_seen_on: now
|
||||
});
|
||||
}
|
||||
logAudit('spoke_registered', { endpoint, siteSlug: spoke.siteSlug });
|
||||
res.json({ status: 'ok', pushToken: spoke.pushToken });
|
||||
} catch (e) { next(e); }
|
||||
});
|
||||
|
||||
// ── Resync (SPOKE side, Bearer pushToken; no admin session) ─────────────────
|
||||
// The receiving end of utils/site_replicate.js's fire-and-forget push: the
|
||||
// master pings this when its catalog changes. Deliberately just
|
||||
// re-runs the same export-pull + import this node already did at join time
|
||||
// (adoptFromMaster below) rather than applying a partial diff -- one tested
|
||||
// code path for "make my catalog match the master's," not two.
|
||||
router.post('/resync', async (req, res, next) => {
|
||||
try {
|
||||
const cfg = siteConfig.get();
|
||||
if (cfg.isMaster) return res.status(400).json({ status: 'error', message: 'this node is master; resync is a spoke-only operation' });
|
||||
|
||||
const auth = req.headers.authorization || '';
|
||||
const presented = auth.startsWith('Bearer ') ? auth.slice(7).trim() : '';
|
||||
if (!cfg.replicationPushToken || presented !== cfg.replicationPushToken) {
|
||||
return res.status(401).json({ status: 'error', message: 'invalid resync push token' });
|
||||
}
|
||||
if (!cfg.masterUrl || !cfg.masterJoinKey) {
|
||||
return res.status(409).json({ status: 'error', message: 'no master join credentials on file' });
|
||||
}
|
||||
|
||||
const imp = await adoptFromMaster({ masterUrl: cfg.masterUrl, joinKey: cfg.masterJoinKey });
|
||||
logAudit('resynced', { reason: (req.body && req.body.reason) || 'unspecified', resourcesCreated: imp.created, resourcesUpdated: imp.updated });
|
||||
res.json({ status: 'ok', resources: { created: imp.created, updated: imp.updated, edges: imp.edgeCount } });
|
||||
} catch (e) { next(e); }
|
||||
});
|
||||
|
||||
// ── Demote (called on the OLD master; Bearer site-join-key; no admin session)
|
||||
// MULTI_SITE_SPEC.md §3.2: promoting a spoke must be a single coordinated
|
||||
// action, never a two-step "hope nobody's master for a while" gap. The node
|
||||
// being promoted calls this on whatever it currently believes is master,
|
||||
// using the join-key credential it already holds from when it joined --
|
||||
// authenticating "demote me" is exactly the same trust relationship as
|
||||
// authenticating "let me pull an export," so no new credential type is
|
||||
// needed for THIS direction. (The new master's future ability to push
|
||||
// replication/resync to the newly-demoted node is a separate credential --
|
||||
// newJoinKey below -- since that's the master->spoke direction, same as
|
||||
// every other spoke registration.)
|
||||
router.post('/demote', async (req, res, next) => {
|
||||
try {
|
||||
const auth = req.headers.authorization || '';
|
||||
const rawKey = auth.startsWith('Bearer ') ? auth.slice(7).trim() : '';
|
||||
const key = await SiteJoinKey.authenticate(rawKey);
|
||||
if (!key) return res.status(401).json({ status: 'error', message: 'invalid or revoked site join key' });
|
||||
|
||||
const cfg = siteConfig.get();
|
||||
if (!cfg.isMaster) {
|
||||
return res.status(400).json({ status: 'error', message: 'this node is already a spoke' });
|
||||
}
|
||||
const { newMasterUrl, newJoinKey } = req.body || {};
|
||||
if (!newMasterUrl || !newJoinKey) {
|
||||
return res.status(400).json({ status: 'error', message: 'newMasterUrl and newJoinKey are required' });
|
||||
}
|
||||
|
||||
const base = String(newMasterUrl).replace(/\/+$/, '');
|
||||
siteConfig.save({ isMaster: false, masterUrl: base, masterJoinKey: newJoinKey });
|
||||
logAudit('demoted', { demotedBy: key.keyPrefix, newMasterUrl: base });
|
||||
res.json({ status: 'ok', message: 'Demoted to spoke of ' + base });
|
||||
} catch (e) { next(e); }
|
||||
});
|
||||
|
||||
// ── Everything below requires an admin session ──────────────────────────────
|
||||
router.use(middleware.auth);
|
||||
router.use(async (req, res, next) => {
|
||||
@@ -110,9 +219,23 @@ router.use(async (req, res, next) => {
|
||||
});
|
||||
|
||||
// Current multi-site role (master/spoke, site slug, master URL).
|
||||
// Never sent to the client: masterJoinKey and replicationPushToken are live
|
||||
// credentials, not display data. Callers get boolean derivatives instead
|
||||
// (hasMasterJoinKey, liveReplication) -- enough to render UI state without
|
||||
// putting a secret in a browser response.
|
||||
router.get('/config', async (req, res, next) => {
|
||||
try { res.json({ status: 'ok', config: siteConfig.get() }); }
|
||||
catch (e) { next(e); }
|
||||
try {
|
||||
const cfg = siteConfig.get();
|
||||
const { masterJoinKey, replicationPushToken, ...safe } = cfg;
|
||||
res.json({
|
||||
status: 'ok',
|
||||
config: {
|
||||
...safe,
|
||||
hasMasterJoinKey: !!masterJoinKey,
|
||||
liveReplication: !!replicationPushToken
|
||||
}
|
||||
});
|
||||
} catch (e) { next(e); }
|
||||
});
|
||||
|
||||
// ── Site join key management (admin) ────────────────────────────────────────
|
||||
@@ -161,9 +284,77 @@ router.delete('/join-keys/:id', async (req, res, next) => {
|
||||
// role. Only valid on a node that is currently the master (i.e. a fresh
|
||||
// bring-up that has not joined anything yet) — see setup.sh wiring for the
|
||||
// pre-seed timing (this pass is server endpoints only).
|
||||
// Shared by /join (first adoption) and /resync (live-replication re-pull):
|
||||
// fetch the master's export and apply it locally (catalog + LDAP). Throws on
|
||||
// any failure that should surface as a 502 to the caller.
|
||||
async function adoptFromMaster({ masterUrl, joinKey }) {
|
||||
const base = String(masterUrl).replace(/\/+$/, '');
|
||||
const controller = new AbortController();
|
||||
const timer = setTimeout(() => controller.abort(), 30000);
|
||||
let resp;
|
||||
try {
|
||||
resp = await fetch(base + '/api/site/export', {
|
||||
method: 'POST',
|
||||
headers: { Authorization: 'Bearer ' + joinKey, 'Content-Type': 'application/json' },
|
||||
body: '{}',
|
||||
signal: controller.signal
|
||||
});
|
||||
} finally { clearTimeout(timer); }
|
||||
|
||||
if (!resp.ok) {
|
||||
const text = (await resp.text().catch(() => '')).slice(0, 200);
|
||||
const err = new Error('master export failed: HTTP ' + resp.status + ' ' + text);
|
||||
err.httpStatus = 502;
|
||||
throw err;
|
||||
}
|
||||
const exportData = await resp.json();
|
||||
if (!exportData || exportData.status !== 'ok' || !exportData.ldif) {
|
||||
const err = new Error('master export returned no directory');
|
||||
err.httpStatus = 502;
|
||||
throw err;
|
||||
}
|
||||
|
||||
// 1. Adopt the resource catalog.
|
||||
const imp = await importDirectory({ Resource, ResourceEdge, exportData });
|
||||
|
||||
// 2. Adopt the LDAP tree. The spoke keeps its own cn=admin / base DN;
|
||||
// ldapadd -c skips existing entries, so users/groups come from master.
|
||||
let ldapNote = 'imported';
|
||||
try {
|
||||
const adminDn = conf.ldap && conf.ldap.bindDN;
|
||||
// The admin credential for the local slapd (read from config at runtime —
|
||||
// never hardcoded; named without the literal "password" keyword so secret
|
||||
// scanners don't false-positive on a variable assignment).
|
||||
const ldapCred = conf.ldap && conf.ldap.bindPassword;
|
||||
const ldifFile = path.join(os.tmpdir(), 'theta-site-join.ldif');
|
||||
fs.writeFileSync(ldifFile, exportData.ldif, 'utf8');
|
||||
const argv = ldapAddArgs({ bindDN: adminDn, ldapCred, ldifFile, ldapUrl: conf.ldap && conf.ldap.url });
|
||||
await execFileAsync(argv[0], argv.slice(1), { maxBuffer: 4 * 1024 * 1024, timeout: 120000 });
|
||||
fs.unlink(ldifFile).catch(() => {});
|
||||
} catch (e) {
|
||||
ldapNote = 'skipped/failed: ' + e.message;
|
||||
}
|
||||
|
||||
// 3. Adopt the master's agent-signing key, if it sent one (MULTI_SITE_SPEC.md
|
||||
// §2 -- identical directories). Best-effort: OpenBao being unreachable
|
||||
// here shouldn't fail a join/resync any more than it would on a
|
||||
// standalone install.
|
||||
let signingKeyNote = 'not provided by master';
|
||||
if (exportData.signingKey) {
|
||||
try {
|
||||
await agentKeys.adopt(exportData.signingKey);
|
||||
signingKeyNote = 'adopted';
|
||||
} catch (e) {
|
||||
signingKeyNote = 'failed: ' + e.message;
|
||||
}
|
||||
}
|
||||
|
||||
return { imp, ldapNote, signingKeyNote, exportData, base };
|
||||
}
|
||||
|
||||
router.post('/join', async (req, res, next) => {
|
||||
try {
|
||||
const { masterUrl, joinKey } = req.body || {};
|
||||
const { masterUrl, joinKey, selfUrl } = req.body || {};
|
||||
if (!masterUrl || !joinKey) {
|
||||
return res.status(400).json({ status: 'error', message: 'masterUrl and joinKey are required' });
|
||||
}
|
||||
@@ -181,54 +372,52 @@ router.post('/join', async (req, res, next) => {
|
||||
});
|
||||
}
|
||||
|
||||
const base = String(masterUrl).replace(/\/+$/, '');
|
||||
const controller = new AbortController();
|
||||
const timer = setTimeout(() => controller.abort(), 30000);
|
||||
let resp;
|
||||
let adopted;
|
||||
try {
|
||||
resp = await fetch(base + '/api/site/export', {
|
||||
method: 'POST',
|
||||
headers: { Authorization: 'Bearer ' + joinKey, 'Content-Type': 'application/json' },
|
||||
body: '{}',
|
||||
signal: controller.signal
|
||||
});
|
||||
} finally { clearTimeout(timer); }
|
||||
|
||||
if (!resp.ok) {
|
||||
const text = (await resp.text().catch(() => '')).slice(0, 200);
|
||||
return res.status(502).json({ status: 'error', message: 'master export failed: HTTP ' + resp.status + ' ' + text });
|
||||
}
|
||||
const exportData = await resp.json();
|
||||
if (!exportData || exportData.status !== 'ok' || !exportData.ldif) {
|
||||
return res.status(502).json({ status: 'error', message: 'master export returned no directory' });
|
||||
}
|
||||
|
||||
// 1. Adopt the resource catalog.
|
||||
const imp = await importDirectory({ Resource, ResourceEdge, exportData });
|
||||
|
||||
// 2. Adopt the LDAP tree. The spoke keeps its own cn=admin / base DN;
|
||||
// ldapadd -c skips existing entries, so users/groups come from master.
|
||||
let ldapNote = 'imported';
|
||||
try {
|
||||
const adminDn = conf.ldap && conf.ldap.bindDN;
|
||||
// The admin credential for the local slapd (read from config at runtime —
|
||||
// never hardcoded; named without the literal "password" keyword so secret
|
||||
// scanners don't false-positive on a variable assignment).
|
||||
const ldapCred = conf.ldap && conf.ldap.bindPassword;
|
||||
const ldifFile = path.join(os.tmpdir(), 'theta-site-join.ldif');
|
||||
fs.writeFileSync(ldifFile, exportData.ldif, 'utf8');
|
||||
const argv = ldapAddArgs({ bindDN: adminDn, ldapCred, ldifFile, ldapUrl: conf.ldap && conf.ldap.url });
|
||||
await execFileAsync(argv[0], argv.slice(1), { maxBuffer: 4 * 1024 * 1024, timeout: 120000 });
|
||||
fs.unlink(ldifFile).catch(() => {});
|
||||
adopted = await adoptFromMaster({ masterUrl, joinKey });
|
||||
} catch (e) {
|
||||
ldapNote = 'skipped/failed: ' + e.message;
|
||||
return res.status(e.httpStatus || 502).json({ status: 'error', message: e.message });
|
||||
}
|
||||
const { imp, ldapNote, signingKeyNote, exportData, base } = adopted;
|
||||
|
||||
// 3. Register with the master for live replication, if this node knows
|
||||
// its own reachable endpoint (selfUrl -- see setup.env's
|
||||
// CFG_SELF_DIRECTORY_URL). Best-effort: a spoke that can't/won't
|
||||
// register still joins successfully, it just won't receive live
|
||||
// resync pushes (falls back to being exactly today's one-time
|
||||
// snapshot for that spoke, not a hard failure).
|
||||
let replicationPushToken = null;
|
||||
let replicationNote = 'not registered (no selfUrl given)';
|
||||
if (selfUrl) {
|
||||
try {
|
||||
const regResp = await fetch(base + '/api/site/spokes', {
|
||||
method: 'POST',
|
||||
headers: { Authorization: 'Bearer ' + joinKey, 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ endpoint: selfUrl, siteSlug: exportData.siteSlug || cfg.siteSlug })
|
||||
});
|
||||
if (regResp.ok) {
|
||||
const regBody = await regResp.json();
|
||||
replicationPushToken = regBody.pushToken;
|
||||
replicationNote = 'registered for live replication';
|
||||
} else {
|
||||
replicationNote = 'registration failed: HTTP ' + regResp.status;
|
||||
}
|
||||
} catch (e) {
|
||||
replicationNote = 'registration failed: ' + e.message;
|
||||
}
|
||||
}
|
||||
|
||||
// 3. Persist the spoke role (survives restarts). The join key is kept so
|
||||
// the spoke can run WAN-health checks (and, in a later layer, proxy
|
||||
// writes) against the master — it is a spoke-to-master credential, not
|
||||
// a shared secret.
|
||||
siteConfig.save({ isMaster: false, masterUrl: base, siteSlug: exportData.siteSlug || cfg.siteSlug, masterJoinKey: joinKey });
|
||||
// 4. Persist the spoke role (survives restarts). The join key is kept so
|
||||
// the spoke can run WAN-health checks against the master; the push
|
||||
// token (if registration succeeded) is what authenticates the
|
||||
// master's future resync pushes back to THIS node.
|
||||
siteConfig.save({
|
||||
isMaster: false,
|
||||
masterUrl: base,
|
||||
siteSlug: exportData.siteSlug || cfg.siteSlug,
|
||||
masterJoinKey: joinKey,
|
||||
...(replicationPushToken ? { replicationPushToken } : {})
|
||||
});
|
||||
|
||||
logAudit('joined', {
|
||||
actor: req.user.uid,
|
||||
@@ -237,7 +426,9 @@ router.post('/join', async (req, res, next) => {
|
||||
resourcesCreated: imp.created,
|
||||
resourcesUpdated: imp.updated,
|
||||
edges: imp.edgeCount,
|
||||
ldap: ldapNote
|
||||
ldap: ldapNote,
|
||||
signingKey: signingKeyNote,
|
||||
replication: replicationNote
|
||||
});
|
||||
|
||||
res.json({
|
||||
@@ -245,7 +436,9 @@ router.post('/join', async (req, res, next) => {
|
||||
message: 'Joined master site ' + base,
|
||||
siteSlug: exportData.siteSlug || cfg.siteSlug,
|
||||
resources: { created: imp.created, updated: imp.updated, edges: imp.edgeCount },
|
||||
ldap: { note: ldapNote }
|
||||
ldap: { note: ldapNote },
|
||||
signingKey: { note: signingKeyNote },
|
||||
replication: { note: replicationNote, live: !!replicationPushToken }
|
||||
});
|
||||
} catch (e) { next(e); }
|
||||
});
|
||||
|
||||
@@ -0,0 +1,90 @@
|
||||
'use strict';
|
||||
|
||||
// In-memory stand-in for the SiteSpoke ORM model.
|
||||
let spokeStore;
|
||||
function makeSpokeMock() {
|
||||
spokeStore = [];
|
||||
return {
|
||||
list: jest.fn(async () => [...spokeStore]),
|
||||
_seed(rows) { spokeStore.push(...rows); }
|
||||
};
|
||||
}
|
||||
|
||||
let mockFetchCalls = [];
|
||||
let mockFetchImpl = async () => ({ ok: true, status: 200 });
|
||||
|
||||
describe('site_replicate', () => {
|
||||
let siteReplicate;
|
||||
let SiteSpoke;
|
||||
let originalFetch;
|
||||
|
||||
beforeEach(() => {
|
||||
jest.resetModules();
|
||||
mockFetchCalls = [];
|
||||
mockFetchImpl = async () => ({ ok: true, status: 200 });
|
||||
|
||||
jest.doMock('../models/site_spoke', () => ({ SiteSpoke: makeSpokeMock() }));
|
||||
siteReplicate = require('../utils/site_replicate');
|
||||
SiteSpoke = require('../models/site_spoke').SiteSpoke;
|
||||
|
||||
// site_replicate.js uses the global fetch (Node 18+ built-in), not
|
||||
// node-fetch -- stub that directly.
|
||||
originalFetch = global.fetch;
|
||||
global.fetch = (...args) => { mockFetchCalls.push(args); return mockFetchImpl(...args); };
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
global.fetch = originalFetch;
|
||||
});
|
||||
|
||||
test('pushes to every known spoke concurrently with its own pushToken', async () => {
|
||||
SiteSpoke._seed([
|
||||
{ endpoint: 'https://spoke-a.example.com', pushToken: 'token-a' },
|
||||
{ endpoint: 'https://spoke-b.example.com', pushToken: 'token-b' }
|
||||
]);
|
||||
|
||||
await siteReplicate.replicateToSpokes('catalog-changed');
|
||||
await new Promise((r) => setImmediate(r));
|
||||
|
||||
expect(mockFetchCalls.length).toBe(2);
|
||||
const urls = mockFetchCalls.map((c) => c[0]).sort();
|
||||
expect(urls).toEqual(['https://spoke-a.example.com/api/site/resync', 'https://spoke-b.example.com/api/site/resync']);
|
||||
|
||||
const [, optsA] = mockFetchCalls.find((c) => c[0].includes('spoke-a'));
|
||||
expect(optsA.headers.Authorization).toBe('Bearer token-a');
|
||||
expect(JSON.parse(optsA.body).reason).toBe('catalog-changed');
|
||||
});
|
||||
|
||||
test('no known spokes: resolves cleanly, no fetch calls', async () => {
|
||||
await siteReplicate.replicateToSpokes('catalog-changed');
|
||||
await new Promise((r) => setImmediate(r));
|
||||
expect(mockFetchCalls.length).toBe(0);
|
||||
});
|
||||
|
||||
test('one spoke failing does not prevent delivery to another', async () => {
|
||||
SiteSpoke._seed([
|
||||
{ endpoint: 'https://dead-spoke.example.com', pushToken: 'token-dead' },
|
||||
{ endpoint: 'https://live-spoke.example.com', pushToken: 'token-live' }
|
||||
]);
|
||||
mockFetchImpl = async (url) => {
|
||||
if (url.includes('dead-spoke')) throw new Error('connection refused');
|
||||
return { ok: true, status: 200 };
|
||||
};
|
||||
|
||||
await expect(siteReplicate.replicateToSpokes('event')).resolves.toBeUndefined();
|
||||
await new Promise((r) => setImmediate(r));
|
||||
expect(mockFetchCalls.length).toBe(2);
|
||||
});
|
||||
|
||||
test('a non-2xx response from a spoke does not throw out of replicateToSpokes', async () => {
|
||||
SiteSpoke._seed([{ endpoint: 'https://spoke-a.example.com', pushToken: 'token-a' }]);
|
||||
mockFetchImpl = async () => ({ ok: false, status: 500 });
|
||||
await expect(siteReplicate.replicateToSpokes('event')).resolves.toBeUndefined();
|
||||
});
|
||||
|
||||
test('SiteSpoke.list() throwing does not propagate to the caller', async () => {
|
||||
SiteSpoke.list = jest.fn(async () => { throw new Error('db unavailable'); });
|
||||
await expect(siteReplicate.replicateToSpokes('event')).resolves.toBeUndefined();
|
||||
expect(mockFetchCalls.length).toBe(0);
|
||||
});
|
||||
});
|
||||
@@ -90,10 +90,34 @@ function status() {
|
||||
return { available: !!cached, error: loadError };
|
||||
}
|
||||
|
||||
// Overwrite the stored key with material handed over by a master (multi-site
|
||||
// "identical directories" — MULTI_SITE_SPEC.md §2). Every site sharing one
|
||||
// signing key is what lets any site's sso-manager validly sign a command for
|
||||
// an agent enrolled at any other site, at the accepted cost that compromising
|
||||
// ANY one site's OpenBao is equivalent to compromising all of them for agent
|
||||
// command authority. That tradeoff was deliberately accepted for this
|
||||
// deployment's scale (a handful of trusted sites) -- do not call this to sync
|
||||
// keys across a boundary where sites don't trust each other equally.
|
||||
//
|
||||
// Idempotent: adopting the same key material twice (e.g. on every resync
|
||||
// ping) is a no-op past the first call.
|
||||
async function adopt({ privateKeyPem, publicKeyPem }) {
|
||||
if (!privateKeyPem || !publicKeyPem) throw new Error('adopt() requires both privateKeyPem and publicKeyPem');
|
||||
if (cached && cached.privateKeyPem === privateKeyPem && cached.publicKeyPem === publicKeyPem) {
|
||||
return cached; // already holding this exact key -- nothing to do
|
||||
}
|
||||
const material = { privateKeyPem, publicKeyPem };
|
||||
await baoConf.set(PATH, material);
|
||||
cached = { ...material, publicKeyBase64: rawPublicKeyBase64(publicKeyPem) };
|
||||
loadError = null;
|
||||
console.log('[agent_keys] adopted signing key from master (multi-site identical-directory sync)');
|
||||
return cached;
|
||||
}
|
||||
|
||||
// Test seam: drop the in-process cache.
|
||||
function _reset() {
|
||||
cached = null;
|
||||
loadError = null;
|
||||
}
|
||||
|
||||
module.exports = { load, status, rawPublicKeyBase64, _reset, PATH };
|
||||
module.exports = { load, status, adopt, rawPublicKeyBase64, _reset, PATH };
|
||||
|
||||
@@ -0,0 +1,57 @@
|
||||
'use strict';
|
||||
|
||||
// Live replication push -- the piece the shipped v1 join flow doesn't have on
|
||||
// its own (join is a one-time export/import snapshot; nothing kept a spoke in
|
||||
// sync afterward). This fires a lightweight "something changed, re-pull" ping
|
||||
// at every spoke registered in SiteSpoke, concurrently, fire-and-forget: never
|
||||
// awaited by its caller, and one unreachable spoke never delays or blocks
|
||||
// another. See MULTI_SITE_SPEC.md §2.2 for why this must never become a
|
||||
// blocking design (a write must never stall on spoke reachability).
|
||||
//
|
||||
// Deliberately a PUSH-A-SIGNAL / PULL-A-SNAPSHOT design, not a push-a-diff
|
||||
// design: the receiving spoke reacts by calling the master's already-shipped,
|
||||
// already-tested POST /api/site/export + importDirectory() path again (see
|
||||
// routes/api_site.js's /resync handler), rather than this module inventing a
|
||||
// second, parallel way to represent "what changed." Fewer moving parts, and
|
||||
// no risk of a diff payload and a full export ever disagreeing.
|
||||
|
||||
const { SiteSpoke } = require('../models/site_spoke');
|
||||
|
||||
const RESYNC_TIMEOUT_MS = 8000;
|
||||
|
||||
function replicateToSpokes(reason) {
|
||||
return (async () => {
|
||||
let spokes;
|
||||
try {
|
||||
spokes = await SiteSpoke.list();
|
||||
} catch (err) {
|
||||
console.error('[site-replicate] failed to list known spokes:', err.message);
|
||||
return;
|
||||
}
|
||||
for (const spoke of spokes) {
|
||||
// Not awaited -- every spoke is pushed to concurrently.
|
||||
pingOne(spoke, reason).catch((err) => {
|
||||
console.error(`[site-replicate] resync ping to ${spoke.endpoint} failed:`, err.message);
|
||||
});
|
||||
}
|
||||
})();
|
||||
}
|
||||
|
||||
async function pingOne(spoke, reason) {
|
||||
const url = String(spoke.endpoint).replace(/\/+$/, '') + '/api/site/resync';
|
||||
const controller = new AbortController();
|
||||
const timer = setTimeout(() => controller.abort(), RESYNC_TIMEOUT_MS);
|
||||
try {
|
||||
const resp = await fetch(url, {
|
||||
method: 'POST',
|
||||
headers: { Authorization: 'Bearer ' + spoke.pushToken, 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ reason: reason || 'catalog-changed' }),
|
||||
signal: controller.signal
|
||||
});
|
||||
if (!resp.ok) throw new Error('status ' + resp.status);
|
||||
} finally {
|
||||
clearTimeout(timer);
|
||||
}
|
||||
}
|
||||
|
||||
module.exports = { replicateToSpokes };
|
||||
@@ -3339,6 +3339,10 @@
|
||||
'<tr><th>WAN Sync Health:</th><td>' + (res.config.wanConnected === false
|
||||
? '<span class="badge bg-danger"><i class="fa-solid fa-xmark me-1"></i> Offline / Disconnected</span>'
|
||||
: '<span class="badge bg-success"><i class="fa-solid fa-check me-1"></i> Online / Operational</span>') + '</td></tr>' +
|
||||
(!isMaster ? '<tr><th>Live Replication:</th><td>' + (cfg.liveReplication
|
||||
? '<span class="badge bg-success"><i class="fa-solid fa-bolt me-1"></i> Live (catalog updates push automatically)</span>'
|
||||
: '<span class="badge bg-warning text-dark"><i class="fa-solid fa-triangle-exclamation me-1"></i> Snapshot only (re-join to register for live updates)</span>') + '</td></tr>' : '') +
|
||||
(isMaster ? '<tr><th>Registered Spokes:</th><td><span class="badge bg-success">' + (cfg.registeredSpokesCount || 0) + ' receiving live updates</span></td></tr>' : '') +
|
||||
'<tr><th>Registered Sites:</th><td><span class="badge bg-primary">' + (res.sitesCount || 0) + ' sites</span></td></tr>' +
|
||||
'<tr><th>Theta Gateways:</th><td><span class="badge bg-dark">' + (res.gatewaysCount || 0) + ' active gateways</span></td></tr>' +
|
||||
'</table>' +
|
||||
@@ -3359,6 +3363,10 @@
|
||||
'<div class="col-md-7"><input type="text" id="site-join-url" class="form-control form-control-sm" placeholder="Master Directory URL (e.g. https://sso.master.example.com)"></div>' +
|
||||
'<div class="col-md-5"><input type="text" id="site-join-key" class="form-control form-control-sm font-monospace" placeholder="Site join key (stj_...)"></div>' +
|
||||
'</div>' +
|
||||
'<div class="mt-2">' +
|
||||
'<label class="form-label small mb-1">This site\'s own reachable URL <span class="text-muted">(so the master can push live updates here — leave blank to only get a one-time snapshot)</span></label>' +
|
||||
'<input type="text" id="site-join-self-url" class="form-control form-control-sm" value="' + esc(window.location.origin) + '">' +
|
||||
'</div>' +
|
||||
'<button class="btn btn-sm btn-primary mt-2" onclick="joinCurrentSiteToMaster()"><i class="fa-solid fa-link me-1"></i> Join Site</button>' +
|
||||
'</div>' +
|
||||
'</div>';
|
||||
@@ -3408,8 +3416,9 @@
|
||||
if (!confirmed) return;
|
||||
|
||||
try {
|
||||
const res = await app.api.post('directory-admin/site-promote', {});
|
||||
app.messages.toast(res.message || 'Node promoted to Master Site', 'success');
|
||||
const res = await app.api.post('directory-admin/site-promote', { selfUrl: window.location.origin });
|
||||
const handoffOk = res.handoff === 'previous master demoted' || /no previous master/.test(res.handoff || '');
|
||||
app.messages.toast((res.message || 'Node promoted to Master Site') + ' — ' + (res.handoff || ''), handoffOk ? 'success' : 'warning');
|
||||
app.modal.close();
|
||||
refreshSiteStatus();
|
||||
} catch (e) {
|
||||
@@ -3421,6 +3430,7 @@
|
||||
async function joinCurrentSiteToMaster() {
|
||||
const masterUrl = ($('#site-join-url').val() || '').trim();
|
||||
const joinKey = ($('#site-join-key').val() || '').trim();
|
||||
const selfUrl = ($('#site-join-self-url').val() || '').trim();
|
||||
if (!masterUrl || !joinKey) {
|
||||
return app.messages.toast('Enter the master Directory URL and a site join key', 'warning');
|
||||
}
|
||||
@@ -3431,8 +3441,11 @@
|
||||
if (!confirmed) return;
|
||||
|
||||
try {
|
||||
const res = await app.api.post('site/join', { masterUrl, joinKey });
|
||||
app.messages.toast(res.message || 'Joined master site', 'success');
|
||||
const res = await app.api.post('site/join', { masterUrl, joinKey, ...(selfUrl ? { selfUrl } : {}) });
|
||||
const replicationNote = res.replication && res.replication.live
|
||||
? ' Live replication is active.'
|
||||
: ' Snapshot only — this site will not receive live updates (' + ((res.replication && res.replication.note) || 'no selfUrl given') + ').';
|
||||
app.messages.toast((res.message || 'Joined master site') + replicationNote, res.replication && res.replication.live ? 'success' : 'warning');
|
||||
app.modal.close();
|
||||
refreshSiteStatus();
|
||||
} catch (e) {
|
||||
|
||||
@@ -0,0 +1,297 @@
|
||||
'use strict';
|
||||
|
||||
// End-to-end test of the real, shipped multi-site join flow (docs/site-join.md).
|
||||
// Drives the actual HTTP API two humans (a master admin + a spoke admin)
|
||||
// would use: seed an admin on each side, mint a site join key on master,
|
||||
// have the spoke adopt it, and verify the post-join contract holds.
|
||||
|
||||
const { execFileSync } = require('child_process');
|
||||
const crypto = require('crypto');
|
||||
|
||||
// Wrapper matching the async call sites below (execFileSync throws
|
||||
// synchronously; wrap in a resolved/rejected promise so callers can keep
|
||||
// using await/.catch()). NOTE: plain execFile (async) does NOT support the
|
||||
// `input` option for piping stdin -- only the *Sync variants do -- so
|
||||
// ldapadd/ldapmodify would otherwise hang forever waiting on stdin that never
|
||||
// arrives. This bit us once already; don't switch back to async execFile here
|
||||
// without adding real stdin piping.
|
||||
function execFileAsync(cmd, args, opts) {
|
||||
try {
|
||||
const stdout = execFileSync(cmd, args, { ...opts, stdio: ['pipe', 'pipe', 'pipe'] });
|
||||
return Promise.resolve({ stdout: stdout ? stdout.toString() : '' });
|
||||
} catch (e) {
|
||||
e.stderr = e.stderr ? e.stderr.toString() : '';
|
||||
return Promise.reject(e);
|
||||
}
|
||||
}
|
||||
|
||||
const MASTER_URL = process.env.MASTER_URL || 'http://master:3001';
|
||||
const SPOKE_URL = process.env.SPOKE_URL || 'http://spoke:3001';
|
||||
const MASTER_LDAP_HOST = process.env.MASTER_LDAP_HOST || 'master';
|
||||
const MASTER_BASE_DN = process.env.MASTER_BASE_DN || 'dc=master,dc=test';
|
||||
const SPOKE_LDAP_HOST = process.env.SPOKE_LDAP_HOST || 'spoke';
|
||||
const SPOKE_BASE_DN = process.env.SPOKE_BASE_DN || 'dc=spoke,dc=test';
|
||||
const LDAP_ADMIN_PASS = process.env.LDAP_ADMIN_PASS || 'secret';
|
||||
const ADMIN_UID = 'e2eadmin';
|
||||
const ADMIN_PASSWORD = 'MultiSiteE2E!2';
|
||||
|
||||
let failed = false;
|
||||
function fail(msg) {
|
||||
console.error('MULTISITE E2E FAIL:', msg);
|
||||
failed = true;
|
||||
}
|
||||
function step(msg) {
|
||||
console.log('--- ' + msg);
|
||||
}
|
||||
|
||||
async function waitForHealthy(url, label) {
|
||||
for (let i = 0; i < 60; i++) {
|
||||
try {
|
||||
const r = await fetch(`${url}/health`);
|
||||
if (r.ok) return;
|
||||
} catch (_) { /* not up yet */ }
|
||||
await new Promise((res) => setTimeout(res, 1000));
|
||||
}
|
||||
throw new Error(`${label} never became healthy`);
|
||||
}
|
||||
|
||||
// Seed an admin user directly via ldapadd/ldapmodify -- mirrors
|
||||
// test/seed-test-user.sh, but parameterized per-site since master and spoke
|
||||
// have distinct base DNs in this harness.
|
||||
async function seedAdmin(ldapHost, baseDn) {
|
||||
const salt = crypto.randomBytes(8);
|
||||
const digest = crypto.createHash('sha512').update(ADMIN_PASSWORD).update(salt).digest();
|
||||
const hash = '{SSHA512}' + Buffer.concat([digest, salt]).toString('base64');
|
||||
|
||||
const ldif = `
|
||||
dn: cn=${ADMIN_UID},ou=groups,${baseDn}
|
||||
objectClass: posixGroup
|
||||
objectClass: top
|
||||
cn: ${ADMIN_UID}
|
||||
gidNumber: 1600
|
||||
|
||||
dn: cn=${ADMIN_UID},ou=people,${baseDn}
|
||||
objectClass: inetOrgPerson
|
||||
objectClass: posixAccount
|
||||
objectClass: top
|
||||
objectClass: theta42Person
|
||||
objectClass: ldapPublicKey
|
||||
objectClass: sudoRole
|
||||
cn: ${ADMIN_UID}
|
||||
sn: E2E
|
||||
uid: ${ADMIN_UID}
|
||||
uidNumber: 1600
|
||||
gidNumber: 1600
|
||||
homeDirectory: /home/${ADMIN_UID}
|
||||
loginShell: /bin/bash
|
||||
mail: ${ADMIN_UID}@test.local
|
||||
userPassword: ${hash}
|
||||
`.trim() + '\n';
|
||||
|
||||
const bindDn = `cn=admin,${baseDn}`;
|
||||
await execFileAsync('ldapadd', ['-x', '-H', `ldap://${ldapHost}:389`, '-D', bindDn, '-w', LDAP_ADMIN_PASS], { input: ldif })
|
||||
.catch((e) => { if (!/Already exists/.test(e.stderr || '')) throw e; });
|
||||
|
||||
// god_admin is needed for site-promote (SUPER_ADMIN_GROUP, utils/permission.js).
|
||||
for (const group of ['app_sso_admin', 'god_admin']) {
|
||||
const modLdif = `dn: cn=${group},ou=groups,${baseDn}\nchangetype: modify\nadd: member\nmember: cn=${ADMIN_UID},ou=people,${baseDn}\n`;
|
||||
try {
|
||||
await execFileAsync('ldapmodify', ['-x', '-H', `ldap://${ldapHost}:389`, '-D', bindDn, '-w', LDAP_ADMIN_PASS], { input: modLdif });
|
||||
console.log(` (added ${ADMIN_UID} to ${group} on ${ldapHost})`);
|
||||
} catch (e) {
|
||||
if (!/[Tt]ype or value exists/.test(e.stderr || '')) {
|
||||
console.error(` FAILED adding ${ADMIN_UID} to ${group} on ${ldapHost}: ${e.stderr || e.message}`);
|
||||
throw e;
|
||||
}
|
||||
console.log(` (${ADMIN_UID} already in ${group} on ${ldapHost})`);
|
||||
}
|
||||
}
|
||||
|
||||
const verify = await execFileAsync('ldapsearch', ['-x', '-H', `ldap://${ldapHost}:389`, '-D', bindDn, '-w', LDAP_ADMIN_PASS,
|
||||
'-b', `cn=god_admin,ou=groups,${baseDn}`, 'member']);
|
||||
console.log(` god_admin members on ${ldapHost}:\n${verify.stdout}`);
|
||||
}
|
||||
|
||||
async function login(url) {
|
||||
const r = await fetch(`${url}/api/auth/login`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ uid: ADMIN_UID, password: ADMIN_PASSWORD })
|
||||
});
|
||||
if (!r.ok) throw new Error(`login at ${url} failed: ${r.status} ${await r.text()}`);
|
||||
const body = await r.json();
|
||||
return body.token;
|
||||
}
|
||||
|
||||
async function api(url, path, { method = 'GET', token, body } = {}) {
|
||||
const r = await fetch(`${url}${path}`, {
|
||||
method,
|
||||
headers: {
|
||||
'Content-Type': 'application/json',
|
||||
...(token ? { 'auth-token': token } : {})
|
||||
},
|
||||
body: body ? JSON.stringify(body) : undefined
|
||||
});
|
||||
const text = await r.text();
|
||||
let json;
|
||||
try { json = JSON.parse(text); } catch (_) { json = { raw: text }; }
|
||||
return { status: r.status, body: json };
|
||||
}
|
||||
|
||||
async function main() {
|
||||
step('Waiting for master + spoke to be healthy');
|
||||
await waitForHealthy(MASTER_URL, 'master');
|
||||
await waitForHealthy(SPOKE_URL, 'spoke');
|
||||
|
||||
step('Seeding admin users in both sites\' LDAP');
|
||||
await seedAdmin(MASTER_LDAP_HOST, MASTER_BASE_DN);
|
||||
await seedAdmin(SPOKE_LDAP_HOST, SPOKE_BASE_DN);
|
||||
|
||||
step('Logging in as admin on master and spoke');
|
||||
const masterToken = await login(MASTER_URL);
|
||||
const spokeToken = await login(SPOKE_URL);
|
||||
if (!masterToken) fail('no token from master login');
|
||||
if (!spokeToken) fail('no token from spoke login');
|
||||
|
||||
step('Confirming both sites start as master (fresh installs)');
|
||||
{
|
||||
const { body } = await api(MASTER_URL, '/api/site/config', { token: masterToken });
|
||||
if (body.config.isMaster !== true) fail(`expected master to start isMaster:true, got ${JSON.stringify(body.config)}`);
|
||||
}
|
||||
{
|
||||
const { body } = await api(SPOKE_URL, '/api/site/config', { token: spokeToken });
|
||||
if (body.config.isMaster !== true) fail(`expected spoke to start isMaster:true (pre-join), got ${JSON.stringify(body.config)}`);
|
||||
}
|
||||
|
||||
step('Creating a resource on master BEFORE join, to verify it gets adopted');
|
||||
// Only site resources can be top-level; a host needs a parent site.
|
||||
const siteRes = await api(MASTER_URL, '/api/directory-admin/resources', {
|
||||
method: 'POST',
|
||||
token: masterToken,
|
||||
body: { name: 'E2E Site', slug: 'site_e2e', kind: 'site' }
|
||||
});
|
||||
if (siteRes.status !== 200) fail(`seeding pre-join site on master failed: ${siteRes.status} ${JSON.stringify(siteRes.body)}`);
|
||||
|
||||
const seedRes = await api(MASTER_URL, '/api/directory-admin/resources', {
|
||||
method: 'POST',
|
||||
token: masterToken,
|
||||
body: { name: 'E2E Pre-Join Host', slug: 'host_e2e_prejoin', kind: 'host', parentSlug: 'site_e2e' }
|
||||
});
|
||||
if (seedRes.status !== 200) fail(`seeding pre-join resource on master failed: ${seedRes.status} ${JSON.stringify(seedRes.body)}`);
|
||||
|
||||
step('Minting a site join key on master');
|
||||
const keyRes = await api(MASTER_URL, '/api/site/join-keys', {
|
||||
method: 'POST',
|
||||
token: masterToken,
|
||||
body: { label: 'e2e-test' }
|
||||
});
|
||||
if (keyRes.status !== 200 || !keyRes.body.key) fail(`join-key mint failed: ${keyRes.status} ${JSON.stringify(keyRes.body)}`);
|
||||
const joinKey = keyRes.body.key;
|
||||
|
||||
step('Joining spoke to master (with selfUrl, to register for live replication)');
|
||||
const joinRes = await api(SPOKE_URL, '/api/site/join', {
|
||||
method: 'POST',
|
||||
token: spokeToken,
|
||||
// master's own container-internal URL, as the spoke would reach it over the network
|
||||
body: { masterUrl: 'http://master:3001', joinKey, selfUrl: 'http://spoke:3001' }
|
||||
});
|
||||
if (joinRes.status !== 200) fail(`join failed: ${joinRes.status} ${JSON.stringify(joinRes.body)}`);
|
||||
if (!joinRes.body.replication || joinRes.body.replication.live !== true) {
|
||||
fail(`expected join to register for live replication, got ${JSON.stringify(joinRes.body.replication)}`);
|
||||
}
|
||||
|
||||
step('Verifying spoke persisted isMaster:false + masterUrl after join');
|
||||
const { body: spokeCfg } = await api(SPOKE_URL, '/api/site/config', { token: spokeToken });
|
||||
if (spokeCfg.config.isMaster !== false) fail(`spoke should be isMaster:false after join, got ${JSON.stringify(spokeCfg.config)}`);
|
||||
if (!spokeCfg.config.masterUrl) fail('spoke should have masterUrl set after join');
|
||||
|
||||
step('Verifying the spoke adopted the master\'s pre-join catalog');
|
||||
const spokeResources = await api(SPOKE_URL, '/api/directory-admin/resources', { token: spokeToken });
|
||||
const adopted = (spokeResources.body.results || spokeResources.body.resources || spokeResources.body || []);
|
||||
const found = Array.isArray(adopted) && adopted.some(r => r.slug === 'host_e2e_prejoin');
|
||||
if (!found) fail(`spoke did not adopt master's pre-join resource; got slugs=${JSON.stringify((adopted || []).map(r => r.slug))}`);
|
||||
|
||||
step('Verifying spoke is now read-only (write attempt must 403)');
|
||||
const writeAttempt = await api(SPOKE_URL, '/api/directory-admin/resources', {
|
||||
method: 'POST',
|
||||
token: spokeToken,
|
||||
body: { name: 'Should Be Rejected', slug: 'host_e2e_should_reject', kind: 'host' }
|
||||
});
|
||||
if (writeAttempt.status !== 403) fail(`expected 403 writing to spoke post-join, got ${writeAttempt.status} ${JSON.stringify(writeAttempt.body)}`);
|
||||
|
||||
step('Creating a resource on master AFTER join, to verify LIVE replication (not just the one-time join snapshot)');
|
||||
const postJoinRes = await api(MASTER_URL, '/api/directory-admin/resources', {
|
||||
method: 'POST',
|
||||
token: masterToken,
|
||||
body: { name: 'E2E Post-Join Host', slug: 'host_e2e_postjoin', kind: 'host', parentSlug: 'site_e2e' }
|
||||
});
|
||||
if (postJoinRes.status !== 200) fail(`creating post-join resource on master failed: ${postJoinRes.status} ${JSON.stringify(postJoinRes.body)}`);
|
||||
|
||||
step('Waiting for the fire-and-forget resync push to reach the spoke');
|
||||
let liveReplicated = false;
|
||||
for (let i = 0; i < 20; i++) {
|
||||
const r = await api(SPOKE_URL, '/api/directory-admin/resources', { token: spokeToken });
|
||||
const slugs = (r.body.results || r.body.resources || r.body || []).map((x) => x.slug);
|
||||
if (slugs.includes('host_e2e_postjoin')) { liveReplicated = true; break; }
|
||||
await new Promise((res) => setTimeout(res, 500));
|
||||
}
|
||||
if (!liveReplicated) fail('post-join resource never appeared on the spoke -- live replication did not fire (or resync did not apply it)');
|
||||
|
||||
step('Verifying WAN health ping from spoke to master succeeds');
|
||||
const statusRes = await api(SPOKE_URL, '/api/directory-admin/site-status', { token: spokeToken });
|
||||
if (statusRes.body.config && statusRes.body.config.wanConnected !== true) {
|
||||
fail(`expected spoke to report wanConnected:true post-join, got ${JSON.stringify(statusRes.body.config)}`);
|
||||
}
|
||||
|
||||
step('Verifying master itself is unaffected (still isMaster:true, no writes blocked)');
|
||||
const { body: masterCfg } = await api(MASTER_URL, '/api/site/config', { token: masterToken });
|
||||
if (masterCfg.config.isMaster !== true) fail('master flipped away from isMaster:true unexpectedly');
|
||||
|
||||
step('Promoting the spoke to master (coordinated handoff -- must demote the old master too)');
|
||||
const promoteRes = await api(SPOKE_URL, '/api/directory-admin/site-promote', {
|
||||
method: 'POST',
|
||||
token: spokeToken,
|
||||
body: { selfUrl: 'http://spoke:3001' }
|
||||
});
|
||||
if (promoteRes.status !== 200) fail(`promotion failed: ${promoteRes.status} ${JSON.stringify(promoteRes.body)}`);
|
||||
if (promoteRes.body.handoff !== 'previous master demoted') {
|
||||
fail(`expected the old master to be demoted as part of promotion, got handoff=${JSON.stringify(promoteRes.body.handoff)}`);
|
||||
}
|
||||
|
||||
step('Verifying the newly-promoted node is master');
|
||||
const { body: newMasterCfg } = await api(SPOKE_URL, '/api/site/config', { token: spokeToken });
|
||||
if (newMasterCfg.config.isMaster !== true) fail(`newly-promoted node should be isMaster:true, got ${JSON.stringify(newMasterCfg.config)}`);
|
||||
|
||||
step('Verifying the old master was actually demoted to a spoke of the new master');
|
||||
const { body: oldMasterCfg } = await api(MASTER_URL, '/api/site/config', { token: masterToken });
|
||||
if (oldMasterCfg.config.isMaster !== false) fail(`old master should be isMaster:false after being demoted, got ${JSON.stringify(oldMasterCfg.config)}`);
|
||||
if (oldMasterCfg.config.masterUrl !== 'http://spoke:3001') {
|
||||
fail(`old master's masterUrl should now point at the new master, got ${JSON.stringify(oldMasterCfg.config.masterUrl)}`);
|
||||
}
|
||||
|
||||
step('Verifying the (now-demoted) old master rejects writes, and the new master accepts them');
|
||||
const oldMasterWrite = await api(MASTER_URL, '/api/directory-admin/resources', {
|
||||
method: 'POST',
|
||||
token: masterToken,
|
||||
body: { name: 'Should Be Rejected Post-Demotion', slug: 'host_e2e_should_reject_2', kind: 'host' }
|
||||
});
|
||||
if (oldMasterWrite.status !== 403) fail(`expected 403 writing to the demoted old master, got ${oldMasterWrite.status} ${JSON.stringify(oldMasterWrite.body)}`);
|
||||
|
||||
const newMasterWrite = await api(SPOKE_URL, '/api/directory-admin/resources', {
|
||||
method: 'POST',
|
||||
token: spokeToken,
|
||||
body: { name: 'E2E Post-Promotion Host', slug: 'host_e2e_postpromotion', kind: 'host', parentSlug: 'site_e2e' }
|
||||
});
|
||||
if (newMasterWrite.status !== 200) fail(`expected the newly-promoted master to accept writes, got ${newMasterWrite.status} ${JSON.stringify(newMasterWrite.body)}`);
|
||||
|
||||
if (failed) {
|
||||
console.error('MULTISITE E2E: one or more checks failed (see above)');
|
||||
process.exit(1);
|
||||
}
|
||||
console.log('MULTISITE E2E PASS');
|
||||
}
|
||||
|
||||
main().catch((e) => {
|
||||
console.error('MULTISITE E2E FAIL (exception):', e.stack || e.message);
|
||||
process.exit(1);
|
||||
});
|
||||
Reference in New Issue
Block a user