The published docs site had drifted from what actually shipped:
- docs/sso/multi-site.md said the no-inbound relay was "designed but
not automated" -- it's been automated since earlier this session.
Also said "don't combine" multi-site join and LDAP replication --
they're integrated now (join auto-configures LDAP MMR).
- docs/sso/replication.md (the actually-published/linked replication
page -- distinct from theta-directory's own docs/replication.md,
which isn't linked from this site's nav at all) still only described
the old fully-manual LDAP_SERVER_ID/LDAP_REPLICATION_HOSTS setup,
with zero mention of the new auto-config or CFG_LDAP_MMR_MANUAL.
- docs/index.md's feature bullet described multi-site purely as "N-Way
Multi-Master LDAP replication" -- the actual master/spoke join
feature (the more commonly-used, higher-level mechanism) wasn't
mentioned on the homepage at all.
- CFG_PUBLIC_DOMAIN (shipped, never documented anywhere an operator
would read it) now explained in multi-site.md.
- spoke.env now discoverable from quickstart.md, not just multi-site.md.
Also documents the real, load-bearing limitation from this session's
promotion/LDAP-orphan fix: the master's own replication peer list only
updates on its own next setup.sh run, not live the instant a spoke
joins or a promotion happens.
Removes LDAP_SERVER_ID/LDAP_REPLICATION_HOSTS as vars an operator has
to hand-set and keep in sync across every site. bootstrap/
site-ldap-register.js (new) asks sso-manager-node's new
GET /api/site/ldap-peers (spoke) or GET /directory-admin/
ldap-replication-config (master) for this node's assigned ServerID +
current peer list, persists it to /config/ldap-replication.env, and
restarts sso-manager only when the computed config actually changed
(OpenLDAP's static slapd.conf is only read at process start). Runs on
every setup.sh invocation -- both master (peer list grows as spokes
join) and spoke.
CFG_LDAP_MMR_MANUAL=true skips the automatic step entirely, for a
topology outside this theta-suite cluster the script can't derive on
its own -- without this escape hatch, an operator's hand-set
LDAP_SERVER_ID/LDAP_REPLICATION_HOSTS would get silently overwritten
on the next run, since every fresh install starts as a master (the
automatic step always runs by default).
Bumps sso-manager-node to pick up the new endpoints + SiteSpoke.ldapServerId.
A dedicated spoke.env for the join-a-cluster vars (CFG_MASTER_DIRECTORY_URL/
_JOIN_KEY, CFG_SPOKE_NO_INBOUND/_PUBLIC_HOST, CFG_PUBLIC_DOMAIN), split out
of setup.env purely for clarity -- setup.env still has every option and
keeps working as a single file if that's preferred. setup.sh reads both
(setup.env first, spoke.env layered on top so its values win), same
first-run-only rule as setup.env already had.
Also adds CFG_PUBLIC_DOMAIN (documented in MULTI_SITE_SPEC.md §4 but never
actually wired into setup.sh): an inbound spoke/standalone site's own public
web domain, independent of CFG_DOMAIN (the shared LDAP identity namespace,
which must stay identical across every site). Unset behaves exactly as
before -- hostnames derive from CFG_DOMAIN like any standalone install.
sso-manager-node/jump-host already had the relay-automation mechanism
(noInbound/meshIp/publicHost -> theta-proxy route via proxy_client.js,
GET /api/mesh/self on jump-host) but nothing in the actual operator
bring-up flow could ever reach it -- setup.sh, bootstrap/site-join.js,
and setup.env.example had zero wiring for it.
Add bootstrap/site-relay-register.js: reads this spoke's own role from
/config/site.json, logs into the local jump-host as its bootstrap
admin to discover its mesh IP, and registers it with the master. Mesh
peering itself stays a manual step (mint/paste a join token, same
pattern as the site join key), so this runs on every setup.sh
invocation via CFG_SPOKE_NO_INBOUND/CFG_SPOKE_PUBLIC_HOST and is a
no-op ("not meshed yet") until an operator has actually meshed the two
jump-hosts.
Also updates MULTI_SITE_SPEC.md's status table/TODO and the published
mesh.md docs page, which still described this as "designed but not
automated" after the API-level work had already shipped.
Rolls up theta-agent v2.2.0: Windows hosts override (CRLF-aware, ipconfig
/flushdns), /32 host-route pinning so the WireGuard tunnel can't swallow the
direct LAN path, and a prompt WS reconnect on apply/revert. Marks Windows
local-discovery shipped in MULTI_SITE_SPEC.md; macOS remains the one unbuilt
piece (in progress on a macOS VM).
Service-to-service auth is a prerequisite for both cross-component
routing and no-inbound relay automation (both need a real credential
between sso-manager-node and theta-proxy/theta-gateway) -- reordered so
that's not buried as item 5. Also notes that Windows/macOS mDNS is being
built by a separate session rather than silently dropping it with no
explanation.
Everything shipped this pass (live replication, master/spoke join,
gateway-to-gateway WireGuard mesh) had real spec docs in the repo
(docs/MULTI_SITE_SPEC.md, sso-manager-node's docs/site-join.md) but
nothing on the actual published docs site (theta42.github.io/theta-suite/)
-- a reader landing there would find no mention of it at all beyond a
vague, unlinked "multi-site replication" bullet on the homepage.
- New docs/sso/multi-site.md: the operator-facing master/spoke join guide
(why, how, promoting a spoke, what replicates, current limits), with an
explicit section distinguishing it from the pre-existing N-way LDAP MMR
replication page (replication.html) -- two different mechanisms that
were at real risk of being conflated with nothing to tell them apart.
- New docs/jump-host/mesh.md: the gateway-to-gateway WireGuard mesh guide,
linked from a "WireGuard mesh routing" bullet that already existed on
the jump-host homepage but pointed nowhere.
- docs/sso/index.md, docs/jump-host/index.md: link the new pages from
each component's Features list.
- docs/index.md: replaced the oversold, unlinked "multi-site replication
running in seconds" homepage copy with an accurate, linked claim.
Announce (theta-gateway) + discover/apply/revert (theta-agent) confirmed
working end-to-end over real multicast between real containers, including
two real bugs found and fixed along the way (IPv6 query abort, EBUSY on
rename over a bind-mounted /etc/hosts).
Windows/macOS mDNS is now the ONLY unbuilt piece of the original design
this session set out to implement -- and it's blocked on platform access
this environment doesn't have, not on missing design or effort.
Confirmed the core idea (master terminates a connection, relays over a
spoke's WG mesh IP to a spoke with zero published/inbound ports of its
own) with a standalone test: an external client hit the master's public
port and got a response that could only have come from the spoke,
which had no reachable port except over the tunnel.
Deliberately did NOT wire this into theta-proxy's actual Lua/Redis
routing engine -- that needs its own dedicated pass to do safely, plus a
real service-to-service credential between sso-manager-node and
theta-proxy/theta-gateway that doesn't exist yet. Recorded as verified
mechanism / unbuilt automation, not conflated with either "done" or
"unknown whether it would even work."
Updates the status table and top-of-doc callout to reflect what actually
shipped this pass: live catalog replication, identical-directory signing
key, coordinated master promotion (with two real bugs found + fixed along
the way), and a real, tested gateway-to-gateway WireGuard mesh.
Explicitly calls out what's still NOT true despite all of the above: the
mesh exists as its own transport layer but sso-manager-node's join/
replicate traffic doesn't route over it yet, so the no-inbound-spoke
relay scenario still isn't solved end-to-end. mDNS remains unbuilt.
MULTI_SITE_SPEC.md described a WireGuard-mesh + live-replication design as
if unbuilt-but-planned; meanwhile theta-directory v2.2.0-v2.3.0 (rolled up
in theta-suite v2.2.0) already shipped a simpler, real join mechanism
(one-time LDIF/catalog export over a site join key, read-only spoke
enforcement, setup.sh wiring) that this doc didn't mention at all. Added a
callout pointing at docs/site-join.md as the actual current behavior, and
corrected the status table so it no longer implies unbuilt features are
implemented.
Also adds AGENT_LOCAL_DISCOVERY_SPEC.md, a standalone handoff spec for the
mDNS "prefer local discovered directory" optimization -- confirmed not
implemented anywhere in theta-agent. Needs Windows/Mac-native investigation
this environment can't do; written so it can be picked up independently.
Unifies the GitHub Pages docs site: the SSO/Proxy/Jump Host pages, their
nav labels, and each component's own README now consistently say Theta
Directory / Theta Proxy / Theta Gateway, drop marketing sections ("Why this
over the alternatives", "Get it", "Related projects") that don't apply to a
suite component, remove every standalone/bare-metal install path, and link
to theta42.github.io/theta-suite/... instead of the old per-repo Pages sites.
Bumps submodules: theta-directory v2.0.2, proxy v2.0.1, jump-host v2.0.1.
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
seed_node_conf() (added for sso-manager-node's node-scoped secrets
engine, DESIGN.md §5) had two bugs that made it fail on every call:
- Path had an extra "data/" segment (secret/data/nodes/<id>/<name>).
bao kv put takes the mount-relative path and inserts "data/" itself
for KV v2 -- same convention seed_app_conf already uses just above it
(secret/${vault_path}, not secret/data/${vault_path}). Fixed to
secret/nodes/<id>/<name>, which resolves under the hood to the
secret/data/nodes/<id>/* path api_agent_ops.js's node-scope check
expects.
- It piped "key=value\n" lines to `bao kv put path -`, but `-` there
means "read a JSON object from stdin", not KV lines -- failed with
"invalid key/value pair \"-\"" before ever reaching OpenBao. Fixed to
pass key=value pairs as ordinary CLI args.
Verified against a real OpenBao round-trip (write via the fixed
function, read back both via the CLI and the same HTTP path the SSO's
node-scope check uses).
docs/_config.yml: "theta-suite" -> "Theta Suite" in the Jekyll site title.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
bootstrap.js no longer creates host_theta-proxy / host_theta-jump as
synthetic kind:'host' resources. Proxy and jump-host are containers running
on the one real stack host, not machines of their own -- and jump-host
resolves its SSH-reachable-hosts list from exactly kind:'host', so the
mistake wasn't just conceptual, it could offer unreachable SSH targets.
Their services now parent directly onto the stack host, like every other
component. Installs seeded between 2026-08-05 and this release self-heal on
the next ./setup.sh run: existing children are re-parented off the
synthetic hosts and the now-empty synthetic hosts are removed. Validated
live against a running instance carrying the exact bad state.
Bumps submodules to sso-manager-node v1.30.2, proxy v1.35.1,
jump-host v1.19.1.
Also: README's architecture diagram + repo layout were stale (2-service
view predating jump-host/OpenBao, 2 of 5 submodules listed); new
docs/fixtures.md + docs/screenshots.md + bootstrap/seed-demo-users.sh for
consistent, repeatable demo data and screenshot passes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0113gCdnfSCuZr6xvPDxTo3D
Rolls up sso-manager-node v1.29.0, theta-agent v1.4.0, proxy v1.34.0 and
jump-host v1.19.0.
Per-host SSO returned "400 redirect_uri is not registered for this
client". The bootstrap registered only the proxy's own management
callback, but per-host SSO calls back to
https://<protected-host>/__proxy_auth/callback -- a different URL per
proxied host, all against that one OAuth client. Now registers the
wildcard + apex patterns, and backfills them onto existing clients so
upgraded stacks are fixed too.
theta-proxy and theta-jump were seeded as hosts and then left childless
while their services hung off the stack host. Services now parent to the
host that runs them; reparent() corrects existing installs, but only when
the current parent is the one the old code set.
The proxy gets a read-only SSO API token (minted before the OpenBao
snapshot so the running proxy receives it) backing the per-host SSO group
autocomplete, and the sso-broker policy grants secret/agent/* for the
SSO's persistent theta-agent signing key.
BREAKING: theta-agents must be re-enrolled, and ./setup.sh must be re-run
for the new OpenBao grant.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- setup.sh: ldap_host defaults to localhost (the public sso.<domain> can't reach
the 389/636 LDAP ports through NAT); overridable via CFG_LDAPS_HOST
- ldap.vars access groups + ldap-client sssd filter now reference the SSO group
model (site_<loc>_hosts_access, site_<loc>_host_<host>_access, god_admin)
- GROUPS.md §5/§8 updated to the corrected naming
- gitlink: ldap-client ebaac18 (v1.24.0)
Add docs/GROUPS.md — the canonical Group & Permission Model (schema, inheritance
resolver, Directory-only management, multi-site, host-side SSSD mapping, migration)
— link it from the docs index, and note sso v1.25.0 in the changelog.
Co-Authored-By: Claude <noreply@anthropic.com>
- theta-svc token role (periodic 768h): SSO/PROXY/JUMP_VAULT_TOKEN now minted
through it; ensure_token renews periodic tokens on every setup.sh re-run and
detects/revokes/re-mints valid-but-non-periodic tokens from older installs.
- bao-renewer sidecar (docker-compose): renews the three service tokens every
12h while the stack runs.
- sso-app token role (periodic 768h) + sso-broker policy grants for
auth/token/create/sso-app and renew/revoke/lookup-accessor.
- docs/secrets.md rewritten around the new lifecycle.
- Bump sso-manager-node gitlink to v1.23.0 (real vault-403 fix + app-token
lifecycle).
Co-Authored-By: Claude <noreply@anthropic.com>
Prerequisite for the SSO Manager plugin system (shipped in sso-manager-node
v1.17.0). Adds secret/data/plugins/* (CRUD+list) + secret/metadata/plugins/*
(list/read/delete) to the sso-broker policy HCL so the SSO can store per-instance
plugin secrets in OpenBao instead of sso-secrets.js. ensure_policy is idempotent,
so re-running ./setup.sh grants the existing SSO_VAULT_TOKEN live.
Docs: secrets.md (Plugin secrets section + policy row), architecture.md.
Co-Authored-By: Claude <noreply@anthropic.com>
Rename the project to theta-suite (it is now an integrated suite of four
apps around a shared OpenBao secrets store, not a two-project env).
- theta-env -> theta-suite across the superproject: _config.yml (title +
baseurl /theta-suite + repo URLs), README, setup.sh (incl. the
THETA_SUITE_REEXECED self-update sentinel), docker-compose.yml,
bootstrap.js, lint.yml, config.example/*, docs/robots.txt, all docs,
this changelog.
- architecture.md rewritten: real 4-service + ldap-client topology, OpenBao
secrets section, OpenBao-aware config flow; removed "two containers" /
"three repos" / LDAP-"legacy" framing.
- index.md: integrated-suite framing + secrets/OpenBao + ldap-client.
- standalone.md + README: standalone reframed as advanced opt-in.
- sso-manager-node submodule -> v1.16.1 (401 fix on /conf and /vault).
Co-authored-by: Claude <noreply@anthropic.com>
Two fresh-install fixes and promote the SSH jump host from opt-in to core.
setup.sh: fix silent abort after "Minting per-app OpenBao tokens". env_get's
grep|cut pipeline returns non-zero under set -euo pipefail when .env exists
(created by the root VAULT_TOKEN env_upsert) but an app-token key is absent
(the normal first-run state); the unguarded existing assignment from env_get
then tripped set -e and killed the script before minting any token. env_get
now always returns 0 (|| true). Reproduced + verified under the exact condition.
jump host is no longer optional:
- docker-compose.yml: drop profiles jump-host from the jump-host service
(always started); rename the opt-in test fixture profile jump-host to ldap-test.
- setup.sh: SUBMODULES always includes jump-host; build/start/register/summary
no longer guarded by JUMP_ENABLED; drop the COMPOSE_PROFILES export.
- bootstrap.js: jump provisioning + directory record run unconditionally.
- setup.env.example/docs: drop optional/CFG_JUMP_HOST_ENABLED wording.
Co-authored-by: Claude <noreply@anthropic.com>
sso-dashboard.png and proxy-hosts.png still showed the pre-unification
nav; refresh with the current shared shell and add a jump-host dashboard
screenshot alongside them.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both bumps are bug-fix releases: proxy fixes a bootstrap admin lockout
bug, sso-manager-node fixes a crash on the new Sites & Replication
page. See CHANGELOG.md for the embedded submodule changelogs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- proxy -> v1.1.14
- sso-manager-node -> v1.1.14
Both bump @simpleworkjs/conf to 1.2.0 and jq-repeat to 2.2.0, and use
the new CONF_SECRETS env var instead of symlinking the mounted secrets
file into /app/conf/secrets.js. Updated theta-env's own docs/setup.sh/
docker-compose.yml comments to match -- no change to the config file
format or bind mounts.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same treatment as the proxy and sso-manager-node companion PRs. This
repo has no app UI of its own (it's a bash orchestrator), so both the
nav logo and favicon use the shared theta42.svg mark -- matching the
family look shown in its own screenshots (the SSO Manager/proxy
dashboards it stands up).
- New cross-page nav (Home/Quickstart/Architecture/Standalone/
Changelog) -- replaces index.md's old "More docs" section, now
redundant with the top nav.
- SEO: jekyll-seo-tag + jekyll-sitemap, per-page meta description,
OG/Twitter card tags, canonical URLs, JSON-LD, sitemap.xml,
robots.txt.
- Mobile: Bootstrap's responsive grid + collapsible navbar; the
screenshot pair in index.md stacks to full-width below 576px.
- Added docs/_site to .gitignore (missing entirely before -- the
other two repos already had it).
Verified with a real Jekyll build (jekyll/jekyll Docker image) +
Playwright: desktop and mobile (375px) screenshots, mobile nav
toggle, active-link highlighting, zero console/page errors, and
confirmed real SEO output (meta description, OG/Twitter tags,
canonical, JSON-LD, sitemap.xml, robots.txt) via curl against the
served site.
Adds a Keep-a-Changelog-style CHANGELOG.md, linked from README and
docs/index.md, closing the "no changelog or versioning scheme"
issue. Bumps proxy and sso-manager-node to v1.1.3 (both add their
own CHANGELOG.md, served in-app at /docs/changelog).
docs/index.md (the published site's home page) never linked to
architecture.md, quickstart.md, or standalone.md -- they were only
reachable by direct URL. Added a "More docs" section linking all
three.
Bumps proxy and sso-manager-node to v1.1.2 (air-gap fixes + in-app
/docs on both).
- Rewrite docs/index.md as a short landing page (what it is, screenshots,
why this over running the two separately, what you get, a minimal
"get it" snippet) instead of a full config/architecture reference --
that content still lives in the repo (README, docs/*.md), linked from
here.
- Cross-link to SSO Manager's and Proxy's own Pages sites.
- Screenshots are now clickable (open full size) on both the Pages site
and the README.
- Disable show_downloads in docs/_config.yml -- the Cayman theme's
"Download .zip/.tar.gz" buttons are gone; "View on GitHub" (which links
back to the repo) is the only header link now.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
theta-env's GitHub Pages site (docs/, Jekyll) was already configured and
live at https://theta42.github.io/theta-env/ but nothing in the README
linked to it, unlike proxy and sso-manager-node's READMEs -- easy to miss
entirely. Add the same top-of-README Documentation link, plus screenshots
of the composed stack (SSO dashboard + proxy host list from one
./setup.sh run). Also fixes a stale docs/index.md quickstart snippet that
said "set CFG_BASE_DN to your domain" -- CFG_DOMAIN is the actual
required variable (CFG_BASE_DN is an advanced override); everywhere else
in the docs already says CFG_DOMAIN correctly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The proxy routes every hostname it serves purely off a Host record
(ops/nginx_conf/proxy.conf has no default/self route — targetinfo.lua
does a Redis lookup per request, full stop). Nothing created these for
the SSO's own UI or the proxy's own management UI, so on a fresh
install https://<SSO_HOST> and https://<PROXY_HOST> both 404 despite
setup.sh's summary claiming they're "fronted by the proxy under TLS".
Add a step after the proxy is healthy that runs a short script inside
the proxy container calling its Host model directly (no HTTP API call,
since no authenticated session exists yet at this point in the run):
- <SSO_HOST> -> sso-manager:3001 (the Docker service)
- <PROXY_HOST> -> 127.0.0.1:3000 (the proxy's own management app)
Both created with sso_enabled: false — each app already gates its own
login, and SSO-gating the SSO's own login page would be circular.
Idempotent: skips a host that already exists.