Compare commits
14 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 0e59b8dcb4 | |||
| 2d862f93c9 | |||
| e0bdc0df2e | |||
| 577c264a6a | |||
| 4f09354e32 | |||
| 744c85f4bf | |||
| d8d811b0ab | |||
| 42ec5aa208 | |||
| 734c62ac83 | |||
| 1d85aa81f3 | |||
| 47f7f976ab | |||
| ace7b441c2 | |||
| 35c1a476c9 | |||
| aeed5b8723 |
@@ -9,6 +9,81 @@ orchestration code; see each submodule's own `CHANGELOG.md`
|
||||
[jump-host](https://github.com/theta42/jump-host/blob/master/CHANGELOG.md))
|
||||
for what changed inside the apps it composes.
|
||||
|
||||
## [v2.5.0] - 2026-08-10
|
||||
|
||||
Rolls up **sso-manager-node v2.5.0** and **jump-host v2.1.1** — closes the
|
||||
last gap in no-inbound relay automation (the mechanism existed at the API
|
||||
level but nothing in the real operator bring-up flow could reach it) and
|
||||
fixes two real bugs found live-testing it.
|
||||
|
||||
### theta-suite orchestration
|
||||
- **`bootstrap/site-relay-register.js`** + `CFG_SPOKE_NO_INBOUND`/
|
||||
`CFG_SPOKE_PUBLIC_HOST` (`setup.env.example`): a no-inbound spoke's
|
||||
`setup.sh` run now discovers its own jump-host's WireGuard mesh IP and
|
||||
registers it with the master on every invocation (idempotent no-op until
|
||||
the two jump-hosts are actually meshed — that peering stays a deliberate
|
||||
manual step, same as minting/pasting a site join key).
|
||||
- `docs/MULTI_SITE_SPEC.md` and the published `docs/jump-host/mesh.md` page
|
||||
updated — both still described this as "designed but not automated" after
|
||||
the API-level work had already shipped in sso-manager-node/jump-host.
|
||||
|
||||
### sso-manager-node v2.5.0
|
||||
- **No-inbound relay automation reachable from the real join flow.**
|
||||
`POST /api/site/join` now forwards `noInbound`/`meshIp`/`publicHost`
|
||||
through to `POST /api/site/spokes`, which drives `utils/proxy_client.js`
|
||||
to auto-create/update the relay route on the master's `theta-proxy` (a new
|
||||
self-service `prx_...` API token client — reuses `theta-proxy`'s existing
|
||||
token system, not a new credential type). Verified against a real running
|
||||
`theta-proxy` container.
|
||||
- **Replication traffic prefers the mesh.** `utils/site_replicate.js`'s
|
||||
fire-and-forget resync push tries a registered spoke's `meshIp` first,
|
||||
falling back to its public endpoint on failure.
|
||||
|
||||
### jump-host v2.1.1
|
||||
- **`GET /api/mesh/self`** — this gateway's own mesh IP, for local scripts
|
||||
(gated by any self-service API token, not a full admin session).
|
||||
- **Fixed: `/api/mesh/register` was unreachable via HTTP.** A route-mounting
|
||||
order bug meant every `/api/mesh/*` request hit an admin-session gate
|
||||
before `routes/mesh.js` ever ran, so a real gateway-to-gateway mesh join
|
||||
always 401'd. Found live-testing `/self` with two real containers.
|
||||
- **Fixed: the initiating side of a mesh join never recorded its own
|
||||
identity** — `GET /api/mesh/self` and the mesh UI's own-entry handling
|
||||
silently saw nothing on whichever gateway called `/join` (only the
|
||||
receiving side of `/register` persisted a self-entry). Verified with two
|
||||
real meshed containers: both sides now report their own correct mesh IP.
|
||||
- **Mesh peer removal now cleans up its kernel routes** (`wg_iface.removePeer()`)
|
||||
— verified live: routes present after `setPeer`, gone after `removePeer`.
|
||||
|
||||
## [v2.4.0] - 2026-08-10
|
||||
|
||||
Rolls up **theta-agent v2.2.0** — mDNS local-discovery is now Windows-capable,
|
||||
closing the last gap in the original multi-site design on the agent side.
|
||||
|
||||
### theta-agent v2.2.0
|
||||
- **Windows local-discovery**: the hosts override now runs on Windows
|
||||
(`%SystemRoot%\System32\drivers\etc\hosts`, CRLF-aware, `ipconfig /flushdns`
|
||||
after every change). Reachable because the agent runs as a SYSTEM service, so
|
||||
the elevation question in the spec resolved in our favor. The Windows CI leg
|
||||
now runs the real Windows write path instead of skipping.
|
||||
- **Local route pinning** (`local_route*.go`): the hosts override only fixes
|
||||
*name resolution*; the packet path is the routing table's job. If the WireGuard
|
||||
mesh tunnel is up with `AllowedIPs` covering the LAN subnet (or a full-tunnel
|
||||
`0.0.0.0/0`), the tunnel route would swallow the direct connection to the
|
||||
discovered LAN IP. Discovery now pins a `/32` host route via the owning local
|
||||
interface (`route.exe add ... metric 1` on Windows, `ip route replace` on
|
||||
Linux) and drops it on revert — closing a real gap in the shipped Linux path.
|
||||
- **Prompt reconnect**: an apply/revert signals the WebSocket loop, which
|
||||
reconnects immediately instead of waiting out its 5s backoff.
|
||||
- **Installer version fix**: the setup.exe previously hardcoded `2.1.0` in its
|
||||
file name and version resources no matter the tag; it now derives the version
|
||||
from the git tag.
|
||||
|
||||
### docs
|
||||
- `docs/MULTI_SITE_SPEC.md` status table updated: Windows local-discovery marked
|
||||
shipped; macOS remains the one unbuilt piece (hosts override compiles on
|
||||
darwin but needs `dscacheutil -flushcache` + real hardware testing, being done
|
||||
on a macOS VM).
|
||||
|
||||
## [v2.3.0] - 2026-08-10
|
||||
|
||||
Rolls up **theta-directory v2.4.0**, **jump-host v2.1.0**, **theta-agent v2.1.2**. Live catalog replication and a real gateway-to-gateway WireGuard mesh land in the same pass — multi-site directory sync stops being a one-time snapshot, and site-to-site networking becomes real infrastructure instead of a documented-but-unbuilt design. See `docs/MULTI_SITE_SPEC.md` for the full architecture and an explicit TODO list of what's still open (Windows/macOS mDNS, routing directory traffic over the mesh, `theta-proxy` no-inbound relay automation).
|
||||
|
||||
@@ -0,0 +1,128 @@
|
||||
#!/usr/bin/env node
|
||||
/*
|
||||
* theta-suite site-relay-register — runs inside the sso-manager container
|
||||
* (same pattern as site-join.js) to finish no-inbound relay automation for a
|
||||
* spoke with no public IP (MULTI_SITE_SPEC.md §5.2).
|
||||
*
|
||||
* site-join.js's initial join can't supply a mesh IP: this site's jump-host
|
||||
* isn't meshed to the master's yet at that point (mesh peering is a manual,
|
||||
* out-of-band action on both jump-hosts -- mint a join token on the master's
|
||||
* jump-host, paste it into this site's jump-host "Join a mesh" UI action --
|
||||
* the same reason the site join key itself is minted/pasted by hand rather
|
||||
* than automated). This script is the follow-up: run it (setup.sh does, on
|
||||
* every run, when CFG_SPOKE_NO_INBOUND is set) once meshing is done, and it
|
||||
* discovers this jump-host's mesh IP and registers it with the master so
|
||||
* theta-proxy there can auto-create the relay route (see sso-manager-node's
|
||||
* utils/proxy_client.js). Safe to run before meshing completes -- reports
|
||||
* "not meshed yet" and exits 0 so a re-run later just picks it up.
|
||||
*
|
||||
* docker compose exec sso-manager node /bootstrap/site-relay-register.js \
|
||||
* https://sso.this-site.example.com sso-branch2.master-domain.example.com
|
||||
*
|
||||
* Self-contained (Node built-ins + global fetch), same rule as bootstrap.js
|
||||
* and site-join.js -- it does NOT require the SSO's internal models. It
|
||||
* reads this node's own spoke role from /config/site.json (written by
|
||||
* site-join.js) and logs into the LOCAL jump-host as its bootstrap-minted
|
||||
* local admin (/config/jump-secrets.js) to call jump-host's own
|
||||
* GET /api/mesh/self.
|
||||
*
|
||||
* Output (stdout, KEY=VALUE for setup.sh): RELAY=<registered|not-meshed|not-a-spoke|skipped>.
|
||||
* Progress logs go to stderr.
|
||||
*/
|
||||
'use strict';
|
||||
|
||||
const fs = require('fs');
|
||||
|
||||
const SITE_CONFIG = '/config/site.json';
|
||||
const JUMP_SECRETS = '/config/jump-secrets.js';
|
||||
const JUMP_INTERNAL = 'http://jump-host:3002';
|
||||
|
||||
const selfUrl = process.argv[2];
|
||||
const publicHost = process.argv[3];
|
||||
|
||||
function log(msg) { console.error('[site-relay-register] ' + msg); }
|
||||
|
||||
async function main() {
|
||||
if (!selfUrl || !publicHost) {
|
||||
throw new Error('usage: node /bootstrap/site-relay-register.js <selfUrl> <publicHost>');
|
||||
}
|
||||
|
||||
if (!fs.existsSync(SITE_CONFIG)) {
|
||||
log('No /config/site.json yet — this node has not joined a master. Nothing to do.');
|
||||
console.log('RELAY=not-a-spoke');
|
||||
return;
|
||||
}
|
||||
const site = JSON.parse(fs.readFileSync(SITE_CONFIG, 'utf8'));
|
||||
if (site.isMaster || !site.masterUrl || !site.masterJoinKey) {
|
||||
log('Not a joined spoke (missing masterUrl/masterJoinKey, or this is a master). Nothing to do.');
|
||||
console.log('RELAY=not-a-spoke');
|
||||
return;
|
||||
}
|
||||
|
||||
if (!fs.existsSync(JUMP_SECRETS)) {
|
||||
log('No /config/jump-secrets.js — jump-host has not been provisioned yet. Skipping.');
|
||||
console.log('RELAY=skipped');
|
||||
return;
|
||||
}
|
||||
const jumpSecrets = require(JUMP_SECRETS);
|
||||
const jumpAdminUser = (jumpSecrets.auth && jumpSecrets.auth.adminUsers && jumpSecrets.auth.adminUsers[0]) || 'jumpadmin';
|
||||
const jumpAdminPass = (jumpSecrets.auth && jumpSecrets.auth.localAdminPass) || '';
|
||||
if (!jumpAdminPass) {
|
||||
log('jump-secrets.js has no local admin password. Skipping.');
|
||||
console.log('RELAY=skipped');
|
||||
return;
|
||||
}
|
||||
|
||||
const loginRes = await fetch(`${JUMP_INTERNAL}/api/auth/login`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
// jump-host's login route (@simpleworkjs/oidc-client's shared router)
|
||||
// expects `username`, not `uid` -- unlike sso-manager-node's own
|
||||
// /api/auth/login (see site-join.js). Confirmed against a real running
|
||||
// jump-host container; `uid` here just silently 401s.
|
||||
body: JSON.stringify({ username: jumpAdminUser, password: jumpAdminPass }),
|
||||
});
|
||||
if (!loginRes.ok) {
|
||||
throw new Error(`jump-host admin login failed (${loginRes.status}): ${await loginRes.text().catch(() => '')}`);
|
||||
}
|
||||
const { token: jumpToken } = await loginRes.json();
|
||||
if (!jumpToken) throw new Error('jump-host login returned no token');
|
||||
|
||||
const selfRes = await fetch(`${JUMP_INTERNAL}/api/mesh/self`, { headers: { 'auth-token': jumpToken } });
|
||||
if (!selfRes.ok) {
|
||||
throw new Error(`jump-host mesh self-lookup failed (${selfRes.status}): ${await selfRes.text().catch(() => '')}`);
|
||||
}
|
||||
const selfData = await selfRes.json();
|
||||
if (!selfData.meshIp) {
|
||||
log('jump-host is not meshed yet (no mesh IP assigned). Mesh-join it first (jump-host UI), then re-run setup.sh.');
|
||||
console.log('RELAY=not-meshed');
|
||||
return;
|
||||
}
|
||||
log(`Discovered mesh IP ${selfData.meshIp}. Registering with ${site.masterUrl}...`);
|
||||
|
||||
const regRes = await fetch(`${site.masterUrl.replace(/\/+$/, '')}/api/site/spokes`, {
|
||||
method: 'POST',
|
||||
headers: { Authorization: 'Bearer ' + site.masterJoinKey, 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({
|
||||
endpoint: selfUrl,
|
||||
siteSlug: site.siteSlug || '',
|
||||
noInbound: true,
|
||||
meshIp: selfData.meshIp,
|
||||
publicHost,
|
||||
}),
|
||||
});
|
||||
const text = await regRes.text().catch(() => '');
|
||||
let data = null;
|
||||
try { data = JSON.parse(text); } catch (e) { /* not JSON */ }
|
||||
if (!regRes.ok) {
|
||||
throw new Error(`relay registration failed (${regRes.status}): ${(data && data.message) || text}`);
|
||||
}
|
||||
|
||||
log(`Relay: ${(data.relay && data.relay.note) || 'registered'}`);
|
||||
console.log('RELAY=registered');
|
||||
}
|
||||
|
||||
main().catch((e) => {
|
||||
console.error('[site-relay-register] FAILED: ' + e.message);
|
||||
process.exit(1);
|
||||
});
|
||||
+12
-12
@@ -8,8 +8,8 @@
|
||||
> ## Shipped today
|
||||
> - **Join, live replication, promotion** (`sso-manager-node`): a spoke joins via a one-time export over a site join key (`POST /api/site/join-keys` / `/export` / `/join`), then registers its own endpoint so the master can push live resync pings on every catalog write — no longer a one-time snapshot. Promotion (`POST /api/directory-admin/site-promote`) coordinates a real handoff, demoting the old master as one action. Identical agent-signing keys ride the same export/resync path. Read [`sso-manager-node/docs/site-join.md`](https://github.com/theta42/theta-directory/blob/master/docs/site-join.md) and `directory_spec.md` §11 for the endpoint-level detail.
|
||||
> - **Gateway-to-gateway WireGuard mesh** (`theta-gateway`): real site-to-site tunnels via `POST /api/mesh/register`/`/join`, kernel WireGuard with a userspace `wireguard-go` fallback. Verified with an actual two-container encrypted tunnel passing traffic, not a mock.
|
||||
> - **Not yet connected to each other**: the mesh is a transport layer that exists on its own; `sso-manager-node`'s HTTPS-based join/replicate calls don't route over it yet. That wiring, plus the no-inbound relay it would enable (mechanism verified, automation not built — see status table), is the next layer.
|
||||
> - **mDNS local-discovery, Linux**: shipped and verified end-to-end — `theta-gateway` announces (`services/mdns_announce.js`), `theta-agent` discovers and applies a hosts-file override, cleanly reverts when the announcement disappears. Windows/macOS remain unbuilt — see the TODO list.
|
||||
> - **Cross-component routing + no-inbound relay automation**: `sso-manager-node`'s replication traffic now prefers a spoke's mesh IP over the open internet when one is on file (`utils/site_replicate.js`), and a no-inbound spoke's join (`POST /api/site/join` → `/api/site/spokes`) can carry `noInbound`/`meshIp`/`publicHost`, which drives `utils/proxy_client.js` to auto-create the relay route on the master's `theta-proxy` via its existing self-service API token system (reused, not a new credential type). The one piece that stays a manual, out-of-band step is the mesh peering itself (mint a join token on one jump-host, paste it into the other's "Join a mesh" UI) — `theta-suite`'s `bootstrap/site-relay-register.js` (`CFG_SPOKE_NO_INBOUND`/`CFG_SPOKE_PUBLIC_HOST`) picks up from there on the next `setup.sh` run.
|
||||
> - **mDNS local-discovery (Linux + Windows)**: shipped and verified — `theta-gateway` announces (`services/mdns_announce.js`), `theta-agent` discovers and applies a hosts-file override, cleanly reverts when the announcement disappears. Linux was verified end-to-end over real multicast; Windows shipped in `theta-agent` v2.2.0 (CRLF-aware hosts override, `ipconfig /flushdns`, and a /32 host-route pin so the WireGuard tunnel can't swallow the direct LAN path). macOS still needs real testing — see the TODO note.
|
||||
|
||||
Design scale: a handful of sites (dozen max, 254 hard ceiling — see §4), a few hundred users/hosts total. This is a deliberate, small, trusted-operator deployment, not a hyperscale/adversarial-tenant one — several decisions below (fire-and-forget replication, identical directories) trade blast-radius for simplicity *because* the scale allows it. Don't generalize these choices past that scale without re-deriving them.
|
||||
|
||||
@@ -260,18 +260,18 @@ See [`AGENT_LOCAL_DISCOVERY_SPEC.md`](./AGENT_LOCAL_DISCOVERY_SPEC.md) — split
|
||||
| Continuous/live replication (vs. one-time export-on-join) | **Shipped** (`sso-manager-node`) — a spoke registers its own endpoint at join time (`POST /api/site/spokes`), and every successful master catalog write fires a fire-and-forget push (`utils/site_replicate.js`) at every registered spoke, which re-pulls a fresh export. Verified end-to-end in `docker-compose.multisite-e2e.yml`. |
|
||||
| Identical-directory signing key | **Shipped** — `POST /api/site/export` includes the master's agent-signing key; a spoke adopts it via `agent_keys.adopt()` on join and every resync. OpenBao secret replication *beyond* this one key is still not built. |
|
||||
| Coordinated master promotion (demote the old master as one action) | **Shipped** — `POST /api/site/demote` + `site-promote`'s handoff logic. Fixed two real pre-existing bugs while wiring this in: `site-promote`'s god_admin check read a `req.user.groups` field nothing ever populated (permanently 403'd for everyone), and the read-only write-gate 403'd `site-promote` itself before the handler could run. |
|
||||
| WireGuard gateway-to-gateway mesh (`theta-gateway`) | **Shipped** — `POST /api/mesh/register`/`/join` (join-token bootstrap), `utils/wg_iface.js` (kernel WireGuard, falls back to userspace `wireguard-go`). Verified with a real two-container test: actual encrypted tunnel, real ICMP traffic across it, 0% loss. This is the mesh transport layer only — nothing in `sso-manager-node`'s replication yet routes traffic *over* it; today's site-to-site HTTPS calls (join/export/resync) still go over whatever network path already reaches the target, same as before this layer existed. |
|
||||
| No-inbound-spoke relay (master proxies a spoke with no public IP) | **Mechanism verified, automation not built.** Confirmed with a standalone test (not `theta-proxy`'s actual Lua/Redis engine, which needs its own dedicated pass to wire safely): a spoke with zero published ports, reachable only via its WG mesh IP, served a request that an external client sent to the master's public port — the master terminated the connection and relayed over the tunnel. So the underlying idea works; what's missing is `theta-proxy` automatically creating that relay route when a no-inbound spoke registers (needs a real service-to-service credential between `sso-manager-node` and `theta-proxy`/`theta-gateway` that doesn't exist yet — a new integration, not a small wiring task), and today's HTTPS-based join/replicate still requires the spoke to reach the master's API directly (and vice versa for export), so a spoke with zero inbound *and* zero outbound path still can't join at all. |
|
||||
| WireGuard gateway-to-gateway mesh (`theta-gateway`) | **Shipped** — `POST /api/mesh/register`/`/join` (join-token bootstrap), `utils/wg_iface.js` (kernel WireGuard, falls back to userspace `wireguard-go`). Verified with a real two-container test: actual encrypted tunnel, real ICMP traffic across it, 0% loss. `wg_iface.removePeer()` also cleans up the kernel routes `setPeer()` added (verified live: routes present after `setPeer`, gone after `removePeer`, own local route untouched), and `DELETE /api/mesh/gateways/:id` exposes it from the mesh UI. |
|
||||
| Cross-component routing (replication over the mesh) | **Shipped** — `utils/site_replicate.js` tries a registered spoke's `meshIp` first (falling back to its public `endpoint` on failure) when pushing resync pings; a spoke with no `meshIp` on file behaves exactly as before. |
|
||||
| No-inbound-spoke relay (master proxies a spoke with no public IP) | **Shipped at the API/automation layer, wired into the real bootstrap flow.** `POST /api/site/join`/`/api/site/spokes` accept `noInbound`/`meshIp`/`publicHost` and call `utils/proxy_client.js`, which mints/reuses a `theta-proxy` self-service API token (`prx_...`, OpenBao `secret/integrations/theta-proxy`) and calls the proxy's real Host API to create or update the relay route — verified against a real running `theta-proxy` container (`GET /api/host/:item`'s actual `{item, results: {...}}` response shape, not the flat shape first assumed). `theta-suite`'s `bootstrap/site-relay-register.js` + `CFG_SPOKE_NO_INBOUND`/`CFG_SPOKE_PUBLIC_HOST` (`setup.env.example`) drive it from the operator-facing bring-up flow, re-run automatically on every `setup.sh` invocation until the jump-host mesh IP is discoverable. What's still a manual step, deliberately: the gateway-to-gateway mesh *peering* itself (mint a join token on one jump-host, paste it into the other's UI) — same pattern as minting/pasting a site join key, not something an unattended script should do blind. A spoke with zero inbound *and* zero outbound path still can't join at all (join/export still need the spoke to reach the master's API directly). |
|
||||
| mDNS local-discovery (Linux) | **Shipped** — `theta-gateway` announces (`services/mdns_announce.js`, opt-in via `THETA_LOCAL_DISCOVERY_HOSTS`), `theta-agent` discovers and applies a hosts-file override (`local_discovery.go`, opt-in via `prefer_local_directory`). Verified end-to-end with real containers over real multicast: announce → discover → apply → clean revert on disappearance, all confirmed. Caught two real bugs along the way (`mdns.Lookup()`'s IPv6 query aborting the whole lookup even after a valid IPv4 response arrived; `rename()` failing with EBUSY over a bind-mounted `/etc/hosts`, common in every container runtime) — see the commit messages in `theta-agent`. |
|
||||
| mDNS local-discovery (Windows, macOS) | Not built — needs platform-native testing this environment can't do (hosts-file vs. stub-resolver tradeoff, elevation, DNS-cache behavior per OS — see Appendix B §3). This is now the **only unbuilt piece** of the original design. |
|
||||
| mDNS local-discovery (Windows) | **Shipped** — `theta-agent` v2.2.0: Windows hosts override (`%SystemRoot%\System32\drivers\etc\hosts`, CRLF-aware, `ipconfig /flushdns` after each change — reachable because the agent runs as a SYSTEM service, so the elevation question resolved in our favor), plus a /32 host-route pin via the owning local interface (`route.exe add ... metric 1`) so the WireGuard mesh tunnel can't swallow the direct LAN path, and a prompt WS reconnect on apply/revert. Tests run the real Windows write path on the Windows CI leg. |
|
||||
| mDNS local-discovery (macOS) | Not built — the hosts override compiles on darwin via the shared unix path, but macOS still needs `dscacheutil -flushcache` and real hardware testing (mDNSResponder behavior, hosts-file vs. native Bonjour — see Appendix B §3). Being built on a real macOS VM. |
|
||||
|
||||
### TODO — what's actually left, in rough dependency order
|
||||
### TODO — what's actually left
|
||||
|
||||
1. **mDNS local-discovery, Windows + macOS** — needs platform-native testing this Linux environment cannot do (hosts-file vs. stub-resolver tradeoff, elevation, DNS-cache quirks per OS — see Appendix B §3). Blocked on a Windows/Mac dev environment, not on design. The Linux side (announcer + agent listener) is done and verified — this is the only remaining piece of the original design with no Linux-buildable path forward.
|
||||
2. **Route `sso-manager-node`'s HTTPS traffic (join/export/resync) over the WireGuard mesh** instead of the open internet, now that the mesh exists as its own transport layer. Currently the two subsystems don't know about each other.
|
||||
3. **`theta-proxy` automation for the no-inbound relay** — mechanism is verified (see status table), but nothing creates the relay route automatically when a no-inbound spoke registers. Needs a new service-to-service credential between `sso-manager-node` and `theta-proxy`/`theta-gateway` — a real design decision (who mints it, what it authorizes), not just wiring.
|
||||
4. **OpenBao secret replication beyond the one agent-signing key** — LDAP admin creds, JWT secret, other per-deployment secrets that currently differ per site.
|
||||
5. **`theta-proxy`/`theta-gateway` service-to-service auth model in general** — items 2 and 3 both need it; worth designing once rather than inventing a credential per integration.
|
||||
6. **Mesh peer removal cleanup** — `wg_iface.removePeer()` doesn't remove the kernel routes `setPeer()` adds (flagged in code, not yet exercised because nothing removes a mesh peer today).
|
||||
1. **Full secret replication** — only the agent-signing key is replicated today. LDAP admin credentials, JWT secrets, and other per-deployment secrets still differ per site, which complicates full disaster recovery. **Paused pending a real-deployment question independent of the code**: this repo's own `conf/secrets.js` was found to contain committed real credentials during this work (LDAP bind, SMTP, VoIP.ms) — see the git-remediation note elsewhere in this repo's history. Building a feature that copies live secrets to additional sites shouldn't proceed until provider-side rotation of those specific credentials is confirmed done; the mechanism itself (generic secret sync, never touching those particular values) can still be designed without that answer.
|
||||
2. Service-to-service auth, cross-component routing, no-inbound relay automation, and mesh peer cleanup (the four items formerly listed here) are **done** — see the status table above. What remains genuinely open in that area is documented there inline (mesh peering stays a manual step by design; zero-inbound-and-zero-outbound spokes still can't join).
|
||||
|
||||
**mDNS local-discovery, macOS** is deliberately not listed above: the Linux and Windows sides are shipped and verified (`theta-agent` v2.2.0), and macOS is being built on a real macOS VM where the darwin-specific behavior (mDNSResponder/DNS-cache) can actually be tested. Check `theta-agent`'s recent history before assuming it's still open.
|
||||
|
||||
*Committed under [`docs/MULTI_SITE_SPEC.md`](file:///home/william/dev/theta42/theta-env/docs/MULTI_SITE_SPEC.md).*
|
||||
|
||||
+5
-4
@@ -10,8 +10,9 @@ Theta Suite is your one-line solution to replacing fragmented, hard-to-wire
|
||||
authentication setups with a unified security stack. It wires together OIDC
|
||||
authentication, LDAP user directories, automated host enrollment, and
|
||||
centralized secret management in a single command. It eliminates the manual
|
||||
configuration friction so you get secure access, auditability, and multi-site
|
||||
replication running in seconds.
|
||||
configuration friction so you get secure access, auditability, and
|
||||
[multi-site](sso/multi-site.html) replication when you need more than one
|
||||
location.
|
||||
|
||||
## Who This Is For
|
||||
* **Self-Hosters & Homelab Engineers:** Anyone running local bare metal,
|
||||
@@ -24,8 +25,8 @@ replication running in seconds.
|
||||
lock-in.
|
||||
* **DevOps & Systems Operators:** Engineers who value idempotent, single-command
|
||||
deployments (`./setup.sh`) and need a production-grade baseline supporting
|
||||
zero-trust proxying, SSH jump-host access control, and multi-site replication
|
||||
out of the box.
|
||||
zero-trust proxying, SSH jump-host access control, and
|
||||
[multi-site](sso/multi-site.html) replication out of the box.
|
||||
|
||||
## Screenshots
|
||||
|
||||
|
||||
@@ -84,7 +84,7 @@ Theta Gateway answers both from your directory:
|
||||
inventory, not a static list
|
||||
- **Per-user key injection** — no downstream changes, no key distribution
|
||||
- **Shell, exec, and SFTP** bridging
|
||||
- **WireGuard mesh routing** — cross-site network access alongside SSH
|
||||
- **[WireGuard mesh routing](mesh.html)** — cross-site network access alongside SSH
|
||||
- **Web UI + HTTP API** for auditing and metrics — active sessions, a searchable
|
||||
audit log, per-user/per-host counters
|
||||
- **Full audit trail** — who, target, method, result, bytes, duration, and the
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
---
|
||||
layout: default
|
||||
title: Gateway Mesh
|
||||
---
|
||||
|
||||
# Gateway Mesh
|
||||
|
||||
Theta Gateway can mesh with other Theta Gateway instances over real
|
||||
site-to-site WireGuard tunnels — separate from its [SSH jump
|
||||
host](connecting.html) role, and separate from the roaming-client/exit-node
|
||||
WireGuard feature (individual peer configs for laptops/phones). This is
|
||||
gateway-to-gateway: two sites' networks reaching each other directly.
|
||||
|
||||
## Why and when to use this
|
||||
|
||||
- **Direct site-to-site networking**, not just SSH. Once two gateways are
|
||||
meshed, hosts behind each can reach each other over the tunnel using the
|
||||
mesh addressing scheme below — not limited to jumping through SSH.
|
||||
- **No manual WireGuard config.** Meshing is a join-token exchange; both
|
||||
sides come out with a live, working peer entry for each other
|
||||
automatically.
|
||||
- **Works without a kernel WireGuard module.** Prefers in-kernel WireGuard,
|
||||
falls back to the userspace `wireguard-go` implementation automatically —
|
||||
useful for older kernels, some container/cloud images, or hosts where the
|
||||
kernel module isn't available.
|
||||
|
||||
## How it works
|
||||
|
||||
1. On the gateway you want others to join, mint a join token: **Mesh** page
|
||||
→ **Mint a Join Token**. It's single-use and expires in 15 minutes.
|
||||
2. On the new gateway, use **Join a Remote Gateway's Mesh**: paste the other
|
||||
gateway's URL and the token.
|
||||
3. Both sides now have a live WireGuard peer for each other. The **Meshed
|
||||
Gateways** table shows every peer, its assigned mesh subnet, and when it
|
||||
was last seen.
|
||||
|
||||
Each gateway is assigned a **mesh index** (an integer 1–254) the first time
|
||||
it either mints a token or is registered by another gateway. That index
|
||||
determines its subnet: `172.24.<index>.0/24` for the mesh tunnel itself, plus
|
||||
`10.<index>.0.0/16` reserved for that site's own local network — 254 sites is
|
||||
the hard ceiling this addressing scheme supports.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Both gateways need a reachable endpoint (host:port) for the WireGuard
|
||||
handshake — typically the same public host the SSH/web ports are already
|
||||
on, with UDP 51820 reachable.
|
||||
- `NET_ADMIN` capability (or equivalent) on the container/host running the
|
||||
gateway, to create the WireGuard interface.
|
||||
|
||||
## Connected to directory sync
|
||||
|
||||
[Theta Directory's multi-site join](../sso/multi-site.html) (catalog + LDAP
|
||||
replication between a master and its spokes) prefers this mesh once it's up:
|
||||
a spoke that's registered a mesh IP gets its live resync pushes routed over
|
||||
the tunnel instead of the open internet, falling back to its public endpoint
|
||||
if the mesh path fails. A spoke with no public IP at all can also register as
|
||||
no-inbound (`CFG_SPOKE_NO_INBOUND` in `theta-suite`'s `setup.env`) so the
|
||||
master auto-creates a relay route through its own `theta-proxy` — the master
|
||||
terminates TLS for that spoke's hostname and relays over this mesh. The mesh
|
||||
peering itself (this page) stays a manual step on both sides; directory join
|
||||
and relay registration pick up from there. See the [architecture
|
||||
spec](https://github.com/theta42/theta-suite/blob/master/docs/MULTI_SITE_SPEC.md)
|
||||
for the full detail.
|
||||
+2
-1
@@ -45,7 +45,8 @@ stack with one command.
|
||||
- **Direct LDAP binds** — anything that binds LDAP directly (Linux hosts
|
||||
via PAM/SSSD, Gitea, Emby, …) uses LDAPS/StartTLS against the same
|
||||
directory.
|
||||
- **Geo-Location Scaling** — built-in support for N-Way Multi-Master OpenLDAP [replication](replication.html) across physical sites.
|
||||
- **[Multi-Site](multi-site.html)** — one master site, any number of read-only spokes that join with a single key and stay live-synced, with god_admin-gated promotion if the master goes down for good.
|
||||
- **Geo-Location Scaling** — built-in support for N-Way Multi-Master OpenLDAP [replication](replication.html) across physical sites (a different, lower-level mechanism — see [Multi-Site](multi-site.html) for how the two compare).
|
||||
- **[Directory & Inventory](directory.html)** — map sites, hosts, and services as a graph with rich metadata (IP/MAC, OS/kernel, ports, git repos), auto-provisioned access groups, and automatic registration from theta-suite's agents and discovery plugins. Drives directory-aware tools like the [SSH jump host](../jump-host/).
|
||||
- **Subtype metrics & lifecycle drivers** — telemetry, log streaming, and remote control for resources tagged with a `subType` (`systemd`, `docker`, `proxmox`, `wireguard`, `postgresql`, `redis`, `k8s`, …).
|
||||
- **OpenBao-backed secrets** — per-resource and per-user secrets with explicit upward inheritance (`Resource → Host → Cluster → Site`).
|
||||
|
||||
@@ -0,0 +1,109 @@
|
||||
---
|
||||
layout: default
|
||||
title: Multi-Site (Master/Spoke Join)
|
||||
---
|
||||
|
||||
# Multi-Site (Master/Spoke Join)
|
||||
|
||||
If you run more than one physical site, Theta Directory can run one site as
|
||||
the **master** (single write authority for the shared catalog) and any
|
||||
number of **spokes** — read-only replicas that stay in sync automatically and
|
||||
run local authentication with zero WAN dependency.
|
||||
|
||||
This is a different, higher-level mechanism than [raw LDAP N-way
|
||||
replication](replication.html) — see [How this relates to LDAP
|
||||
replication](#how-this-relates-to-ldap-replication) below if you're deciding
|
||||
between the two.
|
||||
|
||||
## Why and when to use this
|
||||
|
||||
- **Zero-touch spoke setup.** One join key, one URL, and a spoke adopts the
|
||||
whole directory (users, groups, resource catalog) in one step — no manual
|
||||
`syncrepl` configuration.
|
||||
- **Single write authority, no split-brain.** Only the master accepts
|
||||
directory writes. A spoke that loses WAN connectivity keeps working for
|
||||
local reads/auth and unconditionally stays read-only — it never silently
|
||||
promotes itself. Changing which site is master always requires an explicit,
|
||||
authenticated action by a `god_admin`.
|
||||
- **Stays in sync, not just a one-time copy.** Once joined, a spoke keeps
|
||||
receiving live updates whenever the master's catalog changes — you don't
|
||||
re-run the join to pick up new hosts/apps/users.
|
||||
|
||||
## How it works
|
||||
|
||||
1. **On the master**, an admin mints a **site join key** (Directory → the
|
||||
Master Site modal → **Site Join Keys** → Mint key). It's shown once,
|
||||
stored hashed, and revocable.
|
||||
2. **On the spoke** (must be a fresh install — no users beyond the bootstrap
|
||||
admin, no enrolled agents), either:
|
||||
- Paste the master's URL and the join key into the Master Site modal's
|
||||
**Join an Existing Site** form, or
|
||||
- Set `CFG_MASTER_DIRECTORY_URL` / `CFG_MASTER_DIRECTORY_JOIN_KEY` in
|
||||
`setup.env` before the first `./setup.sh` run.
|
||||
3. The spoke pulls the master's full export (LDAP tree, resource catalog,
|
||||
agent-signing key) and adopts it, then registers its own reachable URL
|
||||
with the master so it can receive live updates going forward.
|
||||
4. From then on, every change to the master's catalog pushes to every
|
||||
registered spoke automatically. A spoke's own directory-write requests are
|
||||
rejected with a `403` pointing at the master — writes always go there.
|
||||
|
||||
### Promoting a spoke to master
|
||||
|
||||
If the master site goes down for good (or you're relocating write
|
||||
authority), a `god_admin` can promote any spoke from its own Master Site
|
||||
modal. Promotion is one coordinated action: it demotes the previous master as
|
||||
part of the same request (best-effort — an unreachable old master never
|
||||
blocks the promotion, since that's exactly the scenario this exists for), and
|
||||
every other spoke gets pointed at the new master automatically.
|
||||
|
||||
## What replicates
|
||||
|
||||
| Data | How |
|
||||
|---|---|
|
||||
| LDAP (users, groups) | Full export on join; live push on every master change |
|
||||
| Resource catalog (hosts, apps, sites) | Same |
|
||||
| Agent-signing key | Same — every site can validly sign a command for any agent enrolled at *any* site |
|
||||
|
||||
The agent-signing key being identical everywhere is a deliberate tradeoff for
|
||||
small, trusted deployments (a handful of sites, not hundreds) — it means
|
||||
compromising the least-secured spoke has the same agent-command blast radius
|
||||
as compromising the master. If that tradeoff doesn't fit your deployment,
|
||||
don't rely on this mechanism as-is.
|
||||
|
||||
Secrets *beyond* the agent-signing key (LDAP admin password, JWT secret, and
|
||||
so on) are **not** currently synced — each site still generates its own.
|
||||
|
||||
## Requirements and current limits
|
||||
|
||||
- Both sites need a network path to each other's HTTP(S) API — the master to
|
||||
pull an export from, the spoke to push replication updates back to. A site
|
||||
with no inbound path at all (e.g. behind CGNAT) can't join yet on its own;
|
||||
a relay mechanism for that case is designed but not automated (see the
|
||||
[architecture spec](https://github.com/theta42/theta-suite/blob/master/docs/MULTI_SITE_SPEC.md)
|
||||
for the current status).
|
||||
- Joining only ever happens on a **fresh install**. There's no way to merge
|
||||
an already-populated directory into a master's — re-provision the host
|
||||
first.
|
||||
|
||||
## How this relates to LDAP replication
|
||||
|
||||
[N-way LDAP replication](replication.html) is a *lower-level*, different
|
||||
mechanism: every site runs a fully independent, fully writable `slapd`, wired
|
||||
together with raw `syncrepl` environment variables, and there's no concept of
|
||||
a master or a managed join. It predates this feature and is still there for
|
||||
deployments that specifically want every site independently writable.
|
||||
|
||||
Multi-site join (this page) is the opposite design: one write authority, a
|
||||
managed onboarding flow, and automatic ongoing sync — closer to what most
|
||||
"add a second office" or "add a home-lab spoke" setups actually want. **Don't
|
||||
combine the two** — pick one per deployment.
|
||||
|
||||
## See also
|
||||
|
||||
- Full architecture and current implementation status:
|
||||
[`MULTI_SITE_SPEC.md`](https://github.com/theta42/theta-suite/blob/master/docs/MULTI_SITE_SPEC.md)
|
||||
in the `theta-suite` repo.
|
||||
- Endpoint-level detail: [`docs/site-join.md`](https://github.com/theta42/theta-directory/blob/master/docs/site-join.md)
|
||||
in the `theta-directory` repo.
|
||||
- Site-to-site networking (WireGuard mesh between gateways, independent of
|
||||
directory sync): [Theta Gateway → Mesh](../jump-host/mesh.html).
|
||||
+1
-1
Submodule jump-host updated: 29029914d2...cde45fff1d
@@ -64,6 +64,20 @@ CFG_DOMAIN=example.com
|
||||
#CFG_MASTER_DIRECTORY_URL=https://sso.master.example.com
|
||||
#CFG_MASTER_DIRECTORY_JOIN_KEY=stj_9f2e...
|
||||
|
||||
# No public IP at all (CGNAT, etc.)? The master can still reach this spoke by
|
||||
# relaying over the gateway-to-gateway WireGuard mesh instead of the open
|
||||
# internet (MULTI_SITE_SPEC.md §5.2) -- but the mesh peering itself is a
|
||||
# manual, out-of-band step on BOTH jump-hosts (mint a mesh join token on the
|
||||
# master's jump-host, paste it into this site's jump-host "Join a mesh" UI
|
||||
# action) that can't run unattended inside this script. Once that's done, set
|
||||
# these two and re-run setup.sh: it discovers this jump-host's assigned mesh
|
||||
# IP (GET /api/mesh/self) and registers it with the master, which then
|
||||
# auto-creates the relay route on its own theta-proxy. Safe to leave set
|
||||
# before meshing -- setup.sh just reports "not meshed yet" and skips until a
|
||||
# later re-run finds the mesh IP.
|
||||
#CFG_SPOKE_NO_INBOUND=true
|
||||
#CFG_SPOKE_PUBLIC_HOST=sso-branch2.master-domain.example.com
|
||||
|
||||
# ── Optional outbound HTTP(S) proxy ──────────────────────────────────────────
|
||||
# For an isolated/offline/corporate-network test host that only reaches the
|
||||
# internet through an upstream HTTP proxy — NOT the theta42 "proxy" app.
|
||||
|
||||
@@ -1279,6 +1279,25 @@ NODEEOF
|
||||
)
|
||||
echo "$JUMP_HOSTS_OUT" | sed 's/^/[setup] /'
|
||||
|
||||
# ── 7b2. No-inbound relay registration (first-run *and* every re-run) ─────────
|
||||
# CFG_SPOKE_NO_INBOUND: this site has no public IP, so the master relays to it
|
||||
# over the gateway-to-gateway WireGuard mesh (MULTI_SITE_SPEC.md §5.2). The
|
||||
# mesh peering itself is a manual, out-of-band step on both jump-hosts (mint a
|
||||
# join token on the master's jump-host, paste it into this site's jump-host
|
||||
# "Join a mesh" UI action) -- it can't run unattended here, and it commonly
|
||||
# happens AFTER this first setup.sh run finishes. So this step runs on every
|
||||
# invocation, not just first-run: it discovers this jump-host's mesh IP and
|
||||
# (re-)registers it with the master, and is a no-op until meshing is done.
|
||||
if [[ "${CFG_SPOKE_NO_INBOUND:-false}" == "true" ]]; then
|
||||
if [[ -z "${CFG_SPOKE_PUBLIC_HOST:-}" ]]; then
|
||||
warn "CFG_SPOKE_NO_INBOUND=true but CFG_SPOKE_PUBLIC_HOST is unset — skipping relay registration."
|
||||
else
|
||||
info "Checking no-inbound relay registration (CFG_SPOKE_PUBLIC_HOST=${CFG_SPOKE_PUBLIC_HOST})..."
|
||||
"${COMPOSE[@]}" exec -T sso-manager node /bootstrap/site-relay-register.js \
|
||||
"https://$CFG_SSO_HOST" "$CFG_SPOKE_PUBLIC_HOST" || warn "relay registration did not complete — check: ${COMPOSE[*]} exec sso-manager node /bootstrap/site-relay-register.js https://$CFG_SSO_HOST $CFG_SPOKE_PUBLIC_HOST"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── 7c. Install theta-agent on the host ──────────────────────────────────────
|
||||
# Controlled by CFG_THETA_AGENT_ENABLE (default: 1 = enabled)
|
||||
CFG_THETA_AGENT_ENABLE="${CFG_THETA_AGENT_ENABLE:-1}"
|
||||
|
||||
+1
-1
Submodule sso-manager-node updated: a50d1ef3c8...b6a82d58d5
+1
-1
Submodule theta-agent updated: 1837d18da8...1b332cacab
Reference in New Issue
Block a user