Commit Graph

409 Commits

Author SHA1 Message Date
wmantly 358231532f Merge remote-tracking branch 'origin/master' into feat-real-gateway-count-and-service-wiring
# Conflicts:
#	nodejs/package.json
2026-08-10 22:24:55 -04:00
wmantly 0a3bdbef06 Merge pull request #201 from theta42/fix-group-dup-and-resources-perf
fix(directory): dedupe access/admin groups; stop self-healing groups on every GET
2026-08-10 19:23:17 -07:00
wmantly 611f1a3318 feat(directory): real mesh-gateway count on the Multi-Site modal; docs links
"Theta Gateways: N active gateways" was counting this app's own
unrelated WireGuard roaming-client/exit-node Resources
(metadata.subType === 'wireguard') -- a completely different subsystem
from the gateway-to-gateway mesh the modal is actually about, and it
never queried jump-host's mesh registry at all (so it couldn't show
the local self-entry either, since there was nothing mesh-related
being counted in the first place).

Added utils/jump_client.js (same pattern as utils/proxy_client.js:
reuses jump-host's existing self-service jmp_ API token system rather
than inventing a new credential) to query jump-host's real
GET /api/mesh/gateways. Reports a null count (not misleading 0) when
the integration isn't configured/reachable, surfaced distinctly in the
UI. Also added help links to the published multi-site/mesh docs on the
modal.

Includes docs links + count only -- this session also discovered that
utils/proxy_client.js's PROXY_INTERNAL_URL, and now JUMP_INTERNAL_URL,
were never actually wired into theta-suite's docker-compose.yml, so
both service-to-service integrations were unreachable in every real
deployment despite existing in code (fixed in theta-suite separately).
2026-08-10 22:17:57 -04:00
wmantly 2c3ec4e967 fix(directory): dedupe access/admin groups on repeated promotion; stop self-healing on every GET
Three independent copies of the same bug: routes/discovery.js's
POST /discovery/promote/:slug (the actual "Promote" button in the UI)
and services/discovery_reconciler.js's autoPromote path both called
ResourceGroup.create() directly with no existence check -- unlike
routes/api_directory_admin.js's own ensureResourceGroup, which already
carried a comment describing this exact "groups appear 3x" bug and
fixing it, just not everywhere it occurred. ResourceGroup has no DB
unique constraint on (resourceId, groupCn), so a resource promoted
more than once (retried UI click, or the same LXC discovered from
multiple Proxmox cluster nodes) silently accumulated duplicate
access/admin rows every time. Added ResourceGroup.ensure() (the
existing check-then-create pattern, now on the model) and switched all
three call sites to it. New regression test in tests/reconciler.test.js.

Also: GET /api/directory-admin/resources ran a full group-model
self-heal fan-out (ensureSiteGroups per site + provisionResourceGroups
per resource, each several sequential LDAP round-trips) unconditionally
on every single list -- confirmed via code read as the actual
bottleneck once a directory has more than a handful of resources, not
data volume. Moved healing to where resources actually change instead
(POST/PUT /resources, POST /discovery/promote/:slug -- PUT had none at
all before this), and added POST /resources/heal-groups as an explicit
on-demand equivalent for backfilling a directory seeded before this
change.
2026-08-10 22:08:04 -04:00
wmantly b6a82d58d5 Merge pull request #200 from theta42/release-v2.5.0
release(v2.5.0): no-inbound relay automation, mesh-preferred replication
v2.5.0
2026-08-10 18:28:27 -07:00
wmantly 18da6582ed release(v2.5.0): no-inbound relay automation, mesh-preferred replication 2026-08-10 21:25:38 -04:00
wmantly b1739ec965 Merge pull request #199 from theta42/feat-mesh-routing
feat(multi-site): route replication over the mesh when available
2026-08-10 18:09:28 -07:00
wmantly 7bc6f47070 feat(multi-site): forward noInbound/meshIp/publicHost through /api/site/join
/api/site/spokes already accepted these and drove proxy_client's relay
automation, but nothing on the real operator join path ever passed
them through -- a no-inbound spoke calling /join had no way to trigger
its own relay route. Forward them into the internal /spokes call and
surface the resulting relay note in /join's response.
2026-08-10 20:52:35 -04:00
wmantly a0964ce350 feat(multi-site): route replication traffic over the mesh when available
Cross-component routing TODO item: a spoke's resync push now prefers
its WG mesh IP (reported via the noInbound/meshIp fields added for the
relay automation) over the public endpoint, falling back to the public
endpoint if the mesh attempt fails for any reason (tunnel not actually
up between these two particular gateways yet, transient failure,
etc.) -- a mesh-routing preference must never turn into "spoke stops
getting updates."

Plain HTTP over the mesh IP, not HTTPS: the WG tunnel is already
encrypted, same reasoning already applied to the no-inbound relay
terminating at the master.

A spoke with no meshIp on file behaves exactly as before (public
endpoint only) -- this is additive, not a behavior change for spokes
that haven't opted into mesh registration.
2026-08-10 20:39:27 -04:00
wmantly 0d3ee3e2ee Merge pull request #198 from theta42/feat-no-inbound-relay
feat(multi-site): no-inbound relay automation via theta-proxy's existing API tokens
2026-08-10 17:36:03 -07:00
wmantly b95cb08c41 feat(multi-site): no-inbound relay automation via theta-proxy's existing API tokens
Implements TODO items "service-to-service auth model" and "no-inbound
relay automation" together -- the first was never really "no credential
type exists", it was that nothing wired one of the credential types that
ALREADY exist (theta-proxy, jump-host, and this app each already have
their own self-service API token system, models/api_token.js) into an
actual inter-service call. This is that wiring, not a new invented
credential type.

- utils/proxy_client.js: ensureRelayRoute({host, ip, targetPort}) calls
  theta-proxy's real Host API (GET/POST/PUT /api/host) using a `prx_...`
  token an operator mints on theta-proxy and stores in OpenBao
  (secret/integrations/theta-proxy), same pattern as agent_keys.js.
  Idempotent and best-effort -- never fails the caller if the token/URL
  isn't configured, since this is an enhancement on top of a working
  join, not a join requirement.
- POST /api/site/spokes accepts optional noInbound/meshIp/publicHost
  fields; when a spoke reports itself no-inbound, the master
  best-effort creates/updates the matching relay route automatically.
  SiteSpoke gained the fields + a relayNote for visibility.

Verified against a REAL running theta-proxy container (not mocked):
booted it standalone, logged in as the local admin, minted a real
`prx_` token via its actual API, and drove ensureRelayRoute() against
it for real. Caught a real bug doing this: GET /api/host/:item wraps
the record in `{item, results}`, not flat -- the mocked unit tests
(which I wrote first) all had the flat shape baked in and passed
cleanly, so this only surfaced against the real API. Fixed in both the
implementation and the unit tests' mocked response shape.
2026-08-10 20:31:42 -04:00
wmantly a50d1ef3c8 Merge pull request #197 from theta42/release-v2.4.0
Build OpenLDAP Base Image / build-and-push (push) Failing after 10s
release(v2.4.0): live replication, coordinated promotion, mesh UI
v2.4.0
2026-08-10 16:03:15 -07:00
wmantly daefe54ff7 release(v2.4.0): live replication, coordinated promotion, mesh UI
Rolls up this pass's multi-site work: live catalog replication (spokes
stay synced after joining, not just a one-time snapshot), identical
agent-signing keys across sites, coordinated master promotion with real
old-master demotion, the UI to actually see and use any of it, and two
real bugs found only by live two-container testing (site-promote's dead
authorization check, and masterJoinKey/replicationPushToken leaking to
the browser via GET /api/site/config).

See CHANGELOG.md for the full list.
2026-08-10 18:54:50 -04:00
wmantly dc3d760d2b feat(multi-site): UI for live replication, promotion handoff, signing key
Closes the gap where all of this session's new server-side capability
(live replication, coordinated promotion, identical signing keys) had
no UI at all -- an operator using the Master Site modal had no way to
know any of it existed or was working.

- Master Site modal: new "Live Replication" row (spoke) shows whether
  this join actually registered for live updates or is stuck on a
  one-time snapshot; new "Registered Spokes" row (master) shows how
  many spokes are receiving live pushes.
- Join form: new "this site's own reachable URL" field, prefilled from
  window.location.origin, wired to the selfUrl the join API already
  supported but the UI never sent -- a UI-driven join previously NEVER
  registered for live replication, only the setup.sh bootstrap path did.
  The success toast now reports whether live replication actually
  activated, not just "joined".
- Promote button: success toast now surfaces the handoff result (old
  master demoted / unreachable / no previous master), so the operator
  sees immediately whether the coordinated demotion actually happened.
- GET /api/site/config no longer returns masterJoinKey or
  replicationPushToken in the response -- found while wiring this up:
  live credentials were being sent straight to the browser for every
  admin session. Replaced with boolean derivatives
  (hasMasterJoinKey, liveReplication).
- GET /api/directory-admin/site-status gained liveReplication (spoke)
  and registeredSpokesCount (master) so the modal has something to render.

Verified by actually driving it in a real browser against a live
container (not just code review): logged in, opened the modal, saw the
new rows, minted a real join key end-to-end, no console errors.

docs/site-join.md rewritten to cover live replication, signing-key sync,
coordinated promotion/demote, and the new endpoints -- it previously only
described the v2.2.0-v2.3.0 one-time-snapshot behavior.
2026-08-10 18:30:48 -04:00
wmantly 9c604f0258 fix(multi-site): coordinated master promotion + a dead-on-arrival authz bug
Two real bugs, both only surfaced by the live two-container e2e test
(docker-compose.multisite-e2e.yml), not by inspection:

1. POST /site-promote's god_admin check read req.user.groups -- a field
   nothing in the codebase ever populates (Auth.checkToken returns
   User.get(), which has no .groups; every other admin gate resolves
   membership live via permission.byGroup()/Group.list(user.dn), which
   also handles nested-group membership). The check silently evaluated to
   an empty array on every request, so site-promote returned 403 for
   every user, including a real god_admin -- unusable since it shipped in
   v2.0.0. Fixed to use permission.byGroup(), the same pattern used
   elsewhere in this file and in api_site.js.

2. The read-only write-gate middleware (api_directory_admin.js) is
   registered before router.post('/site-promote', ...) later in the same
   file, so on a spoke it 403'd every promotion attempt before the
   handler ever ran -- the one mutating request a spoke must be able to
   make to itself. Exempted /site-promote from the gate.

Added coordinated demotion (MULTI_SITE_SPEC.md §3.2 -- promotion as ONE
action, never a two-step gap with two masters): site-promote now calls
the previous master's new POST /api/site/demote (Bearer the join key it
already holds, handing over a freshly-minted key for the demoted node's
own future use) before flipping itself to master. Best-effort: an
unreachable old master never blocks a god_admin's local promotion (the
WAN-outage scenario is the entire reason this control exists), it's
just reported in the response for manual reconciliation.

e2e test extended to promote the spoke, verify the old master was
actually demoted (isMaster:false, masterUrl pointing at the new master),
and verify writes now succeed on the new master and 403 on the old one.
Full chain verified passing: join -> live replication -> promotion ->
demotion -> write authority follows the promotion.
2026-08-10 16:48:33 -04:00
wmantly d27763e556 feat(multi-site): live catalog replication + identical-directory signing key
The shipped join flow (v2.2.0-v2.3.0) was a one-time snapshot: a spoke's
catalog never updated after joining. This adds the two pieces that were
explicitly designed but missing:

- Live replication: a spoke registers its own endpoint with the master
  right after joining (POST /api/site/spokes, Bearer join-key), receiving
  a pushToken. Every successful catalog write on the master now fires a
  fire-and-forget resync ping (utils/site_replicate.js) at every known
  spoke, concurrently -- one unreachable spoke never blocks or delays
  another (wired into the existing write-gate middleware in
  api_directory_admin.js). The spoke's POST /api/site/resync handler
  reuses the already-tested export+import path rather than applying a
  partial diff.

- Identical directories: POST /api/site/export now best-effort includes
  the master's agent-signing key; a spoke adopts it via agent_keys.adopt()
  on both join and every resync, so every site's sso-manager can validly
  sign a command for any agent enrolled anywhere -- the accepted tradeoff
  discussed for this deployment's scale (blast radius for simplicity).

New SiteSpoke model tracks registered spokes (endpoint + pushToken);
registered it in models/index.js (a real bug the e2e test below caught --
SiteSpoke.list() 500'd with "Cannot read properties of null (reading
'adapter')" until the model was added to initORM's model list).

Verified end-to-end against docker-compose.multisite-e2e.yml: mint join
key -> join with selfUrl -> write a NEW resource on master post-join ->
poll the spoke -> it shows up within a few seconds via the resync push,
no manual re-join needed. MULTISITE E2E PASS.

Unit tests: nodejs/tests/site_replicate.test.js (concurrent fan-out, one
failing spoke doesn't block another, empty-registry and list()-throws
edge cases).
2026-08-10 16:34:38 -04:00
wmantly e5167729a8 test: add real two-container e2e for the multi-site join flow
docker-compose.multisite-e2e.yml boots two full all-in-one instances
(master + spoke, each with bundled slapd) and a client that drives the
actual HTTP API: seeds admins, mints a join key, joins the spoke, and
verifies the spoke persisted its role across restart, adopted the
pre-join catalog, is enforced read-only, and reports WAN health.

Verified passing locally. Existing docker-compose.test.yml/e2e.yml split
ldap+redis+app into separate containers, which doesn't work here --
POST /api/site/export runs slapcat in-process, so master and spoke each
need their own bundled slapd (Dockerfile.openldap), not a shared one.
2026-08-10 16:21:53 -04:00
wmantly e915a17cbd Merge pull request #196 from theta42/release-v2.3.0
release(v2.3.0): multi-site join UI + spoke read-only + WAN health
2026-08-10 09:38:06 -07:00
wmantly e4d5e8d75c release(v2.3.0): multi-site join UI, spoke read-only, live WAN health 2026-08-10 09:35:38 -07:00
wmantly 39a9dc0282 Merge pull request #195 from theta42/feat/multi-site-join-ui
feat(site): join UI, spoke read-only, live WAN health, fresh-install guard
2026-08-10 09:15:29 -07:00
wmantly 9d266d2e4c feat(site): join UI, spoke read-only enforcement, live WAN health, fresh-install guard
Completes the multi-site join layer on top of the v2.2.0 endpoints:

- UI (Master Site modal): a fresh install (canJoin) gets a 'Join an Existing
  Site' form (master URL + stj_ key); a master gets a 'Site Join Keys' manager
  (mint/revoke/list, key shown once); WAN Sync Health now reflects a live probe.
- POST /api/site/ping (Bearer stj_ key, no admin session): lightweight master
  reachability probe for WAN health (cheap vs /export).
- Spoke read-only: directory-write routes (resources/edges/groups/secrets/
  grants/driver-action/discovered) reject with 403 pointing at the master.
- Fresh-install guard: /api/site/join refuses unless no users beyond the
  bootstrap admin and no enrolled agents (siteIsFresh), and site-status exposes
  canJoin so the UI only offers join on a genuinely fresh install. The
  bootstrap's seeded default resources are NOT the signal (they always exist).
- The spoke stores the join key (masterJoinKey) in /config/site.json so WAN
  health (and a future write-proxy) can reach the master.
- Tests: siteIsFresh cases in tests/site_join.test.js.
2026-08-10 09:12:28 -07:00
wmantly 3d18c3f0cb Merge pull request #194 from theta42/release-v2.2.0
release(v2.2.0): multi-site join server endpoints
2026-08-10 06:10:02 -07:00
wmantly 214b3a7f5a release(v2.2.0): multi-site join server endpoints, persisted site role, emoji fix 2026-08-10 06:07:07 -07:00
wmantly 08dd234710 Merge pull request #193 from theta42/feat/multi-site-join
feat(site): multi-site join server endpoints + persisted site role + emoji fix
2026-08-10 06:01:28 -07:00
wmantly c96a4b6652 feat(site): multi-site join server endpoints + persisted site role + emoji fix
Server endpoints for joining a spoke to a master directory (MULTI_SITE_SPEC.md).
This pass is server-only; setup.sh wiring and the UI are the next layer.

- Site join keys (SiteJoinKey model, stj_ prefix): mint/revoke/delete/list,
  hashed at rest, shown once — the same model as agent join keys.
- POST /api/site/export (master, Bearer stj_ key, no admin session): returns the
  local LDAP tree (slapcat LDIF) + resource catalog + siteSlug + baseDn.
- POST /api/site/join (spoke, admin): { masterUrl, joinKey } pulls the master
  export, imports resources (upsert by slug) + LDAP (ldapadd -c), and persists
  the spoke role. Refused if already a spoke.
- Persisted site role: utils/site_config.js keeps isMaster/masterUrl/siteSlug in
  /config/site.json (env seeds defaults); site-status/site-promote now use it.
- Unit tests (site_join, site_config) with in-memory stubs, wired into npm test.
- docs/site-join.md + docs router entry.
- Repairs the corrupted multi-site emojis (crown/bolt) in directory.ejs.
- .gitguardian.yml ignores the generic-password false positive on reading the
  LDAP bind credential from runtime config (never a hardcoded secret).
2026-08-10 05:58:28 -07:00
wmantly 0915043d6d Merge pull request #192 from theta42/release-v2.1.1
release(v2.1.1): site-status modal API fix, honest agent version fallback
2026-08-09 20:34:00 -07:00
wmantly ddf123d03c release(v2.1.1): site-status modal API fix, honest agent version fallback 2026-08-09 23:31:15 -07:00
wmantly 9004463311 Merge pull request #191 from theta42/fix/fresh-install
fix(ui): site-status modal API, honest agent version fallback
2026-08-09 20:29:19 -07:00
wmantly 96410fd8c4 fix(ui): site-status modal API, honest agent version fallback, Directory default org name
Fresh-install bug report fixes:

- 'Master Site' button error 'app.modal.show is not a function': the
  multi-site status modal used the legacy app.modal.show() signature; the app
  exposes app.modal.open({title, bodyHtml, size}). The site-status request
  itself worked -- only the rendering call was wrong.
- Agents with no discovery yet showed a fake 'v2.0.0' (three hardcoded
  fallbacks). Now 'unknown', so a host whose agent never connected isn't
  presented as an old version.
- The default org name (browser tab title) is set by theta-suite's setup.sh;
  that default is fixed separately there (CFG_ORG -> Theta Directory).
2026-08-09 23:24:54 -07:00
wmantly f94bb3c14c Merge pull request #190 from theta42/release-v2.1.0
release(v2.1.0): Windows install commands + drop committed binaries
2026-08-09 18:41:12 -07:00
wmantly 37e5ad0494 release(v2.1.0): Windows install commands in the Install Agent modal; drop committed binaries 2026-08-09 21:38:46 -07:00
wmantly 52b85c0d36 Merge pull request #189 from theta42/feat/windows-installer-commands
Windows install commands in the Install Agent modal; drop committed binaries
2026-08-09 18:35:20 -07:00
wmantly f98615aae2 chore(agents): label the Install Agent modal fields as Theta Directory URL 2026-08-09 20:41:50 -07:00
wmantly 72040b1851 feat(agents): Windows install commands in the Install Agent modal; drop committed binaries
The Install Agent modal now emits Windows (PowerShell) one-liners for the
join-key, pre-register and custom-config flows. Each downloads the fully-offline
setup installer and passes the same values the bash flow uses:

  join key : /SERVER_URL /JOIN_KEY
  quick    : /SERVER_URL /AUTH_TOKEN /PUBLIC_KEY
  custom   : /B64_CONFIG=<base64 agent.yml>

Binaries are no longer committed here: the setup.exe and loose agent binaries
are GitHub release artifacts (built by the theta-agent release workflow) and the
modal downloads the installer from releases/latest/download/. The SSO still
serves install.sh (the small Linux bootstrap script); the large binaries are
dropped from this repo.
2026-08-09 20:41:49 -07:00
wmantly 93a022bd66 release(v2.0.4): build Dockerfile.openldap FROM the published base image (#188)
ldapbuild now pulls ghcr.io/theta42/openldap-nestgroup:<pinned commit>
instead of compiling OpenLDAP from source on every build.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-09 17:08:25 -07:00
wmantly c30975329c feat: publish a prebuilt OpenLDAP-with-nestgroup base image (#187)
Extracts Dockerfile.openldap's `ldapbuild` stage (compile OpenLDAP from
source for the nestgroup overlay, ~5 min, dependent on git.openldap.org
being reachable) into its own Dockerfile, built and pushed to
ghcr.io/theta42/openldap-nestgroup by this workflow whenever the pinned
commit changes.

This commit only adds the new image + workflow; Dockerfile.openldap itself
still compiles from source. A follow-up change switches it to FROM the
published image once this workflow has run once and the image exists.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-09 16:54:36 -07:00
wmantly 5c0018f24a release(v2.0.3): fix Directory tab managed-filter bug and site-status 500 (#186)
- GET /api/directory-admin/resources let every kind:'host' resource through
  regardless of promotion status, so "Auto-promote to Directory" unchecked on
  a discovery plugin never kept unpromoted devices out of the Directory tab.
- GET /api/directory-admin/site-status queried the nonexistent Resource.subType
  column instead of metadata.subType, throwing SequelizeDatabaseError.
- Added a "Show ignored" toggle to Discovered Inventory (off by default).
- Untracked nodejs/config/inventory.sqlite -- the app's default runtime DB,
  not a fixture, committed by mistake across 13 prior releases.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-09 16:29:08 -07:00
wmantly 6e172bd528 release(v2.0.2): rebrand README to Theta Directory, remove standalone-install framing (#185)
Removed "Why this over the alternatives"; fixed stale links to the old
per-repo GitHub Pages site to point at the unified theta-suite docs site;
made explicit this is deployed as part of Theta Suite, not standalone; added
the agent capability/install screenshots to the gallery.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-09 16:09:30 -07:00
wmantly 3c98cd4596 Merge pull request #184 from theta42/release-v2.0.1-binaries
release(v2.0.1): bump version to 2.0.1 and package v2.0.1 agent binaries
2026-08-09 12:39:14 -07:00
wmantly d3fab004e9 release(v2.0.1): bump version to 2.0.1 and package v2.0.1 agent binaries 2026-08-09 15:35:29 -04:00
wmantly e8824cb28d Merge pull request #183 from theta42/release-v2.0.1
release(v2.0.1): fix openbao container discovery, agent version API collection, console logs, cache invalidation, site status API 500, secrets filtering, and auto-group spawning
2026-08-09 11:50:19 -07:00
wmantly 98ed99a7e9 release(v2.0.1): fix openbao container discovery, agent version API collection, console logs, cache invalidation, site status API 500, secrets filtering, and auto-group spawning 2026-08-09 14:46:33 -04:00
wmantly bb744adb5e fix(agent): update bundled v2.0.0 multi-arch binaries and install.sh in static resources (#182) 2026-08-09 01:11:55 -04:00
wmantly 420a9c9f8a docs(changelog): add v2.0.0 release notes (#181) 2026-08-09 00:18:34 -04:00
wmantly d442f1e3f9 fix(agent): preserve full telemetry (disks, cpu_details, logged_users) and format desktop_control driver-action params (#180) 2026-08-08 23:41:46 -04:00
wmantly ebd7e9e434 feat(v2.0.0): bump version to 2.0.0, add Multi-Site Master badge, site-status API, and Master promotion UI (#179)
* feat(v2.0.0): bump version to 2.0.0, add Multi-Site Master badge, site-status API, and Master promotion UI

* fix(test): update test script to run unit tests without requiring live Redis socket
2026-08-08 21:43:33 -04:00
wmantly 6ce7e67665 docs: rebrand to Theta Directory (#178) 2026-08-08 20:56:43 -04:00
wmantly 583b822e4a feat(directory): add subtype templates picker, fuzzy merge search, auto-ignore openbao, inherit parent host agent, truncate resource names, key badge positioning (#177) 2026-08-08 19:38:42 -04:00
wmantly 9025feef1b Merge pull request #176 from theta42/release-v1.33.0
Release v1.33.0 - Directory Key Badges, Discovered Inventory Merge/Ignore & Desktop Operations
2026-08-08 18:16:26 -04:00
wmantly 620f401091 release: v1.33.0 - Directory Key Badges, Discovered Inventory Merge/Ignore & Desktop Operations 2026-08-08 18:12:37 -04:00