Repository · workstreams
Workstream: HardwareControlFabric
Status: Planned · Component: Maximize developer productivity
Goal
HardwareControlFabric is the mesh that lets the platform (and its operators) reach into production hosts — list them, inspect their processes, move files, and invoke registered remote functions — without falling back to raw SSH. Today that mesh has no visual surface at all: every interaction goes through HardwareControlFabricCli or hand-written scripts against HardwareControlFabricApi, and diagnosing a production incident means manually running CLI commands or SSHing in. This workstream adds HardwareControlFabricWui, a web dashboard — backed by a new HardwareControlFabricServiceServer per the standard layered architecture — so the mesh's peers, its callable functions, and each daemon's health are visible and actionable in one place instead of reconstructed by hand during an incident. It also covers how a new daemon joins the mesh in the first place: an in-WUI "add peer" flow that either registers an existing daemon by address or walks a user through standing up a brand-new one, with the safety guarantees that a command granting a stranger's service full control of a machine demands.
Current state
| Repository | Role | State |
|---|---|---|
| HardwareControlFabricDaemon | The mesh daemon: mTLS HTTPS API (host/process introspection, file transfer, function invocation) plus a libp2p peer-discovery layer that also bootstraps UrlResolver. Runs on the production droplet (198.199.106.165). | Most actively maintained repo in the family; all PRs merged, no open issues. |
| HardwareControlFabricApi | FabricApiClient — the Kotlin client wrapping the daemon's mTLS API. |
Stable; README still describes the old SSH-based auth model rather than today's mTLS design. |
| HardwareControlFabricCli | Human/agent CLI over the Api (ping, hosts, processes, kill, file read/write, list/invoke functions, p2p peers, deploy-container-nursery). | Only existing client surface; same README staleness as the Api. |
| HardwareControlFabricCliHealthCheck | Registers a ping/status/hosts check with the org's Production Health service. | CI disabled since ~March 2026. |
| fabric-function-restart-containernursery | The only function plugin exercising the daemon's function mechanism. | CI disabled; an open, unmerged fix (PR #7) for a log-truncation bug that destroys OOM forensic logs on every restart. |
| ssh-restart-fabric-daemon / ssh-deploy-fabric-daemon | JSch-based break-glass tools to restart/redeploy the daemon itself over raw SSH, for when the daemon is down. | Byte-for-byte duplicated logic, no shared module. An open, CI-green fix (PR #10) adds the mandated -Xmx128m -XX:+ExitOnOutOfMemoryError flags after a documented 2-week zombie-daemon incident; ssh-deploy-fabric-daemon has no equivalent fix even drafted. |
No WUI exists today, and the function-plugin mechanism has exactly one consumer — the dashboard will be the first UI-facing surface for both. The family also does not yet follow the standard layered architecture: there is no ServiceServer, and the daemon itself speaks a bespoke mTLS HTTPS API rather than url:// — HardwareControlFabricApi talks to each daemon's mTLS HTTPS API directly today, which is how the CLI reaches it and how a WUI would too absent the changes below. The daemon does already run a libp2p peer for P2P discovery (it doubles as a UrlResolver bootstrap node), so the mesh is not starting from zero on the networking side — it just isn't exposing its own command API through that layer yet.
Plan / roadmap
HardwareControlFabricDaemon — migrate to url://
- [ ] Replace the bespoke mTLS HTTPS command API with a proper
url://service. The daemon already runs a libp2p peer for P2P discovery/UrlResolver bootstrap; this extends that same connection to carry the daemon's actual command surface (host/process introspection, file transfer, function invocation) as a normalurl://endpoint (e.g.url://hardware-control-fabric-daemon/<peer-id>/) reached viaopenSandboxedConnection, instead of a separate mTLS HTTPS listener with its own cert-management story. HardwareControlFabricApi'sFabricApiClientmoves onto that transport so the CLI, the ServiceServer, and anything else that talks to a daemon all use the one mesh everything else on the platform already uses.
HardwareControlFabricServiceServer
-
[ ] Stand up
url://hardware-control-fabric/. A new ServiceServer hosting aHardwareControlFabricEmbeddedthat the WUI (and, over time, the CLI) talk to exclusively overurl://— no direct daemon connections from a frontend, per the standard layered architecture's HTTPS-service rules. -
[ ] Own the peer registry, per account. The Embedded holds each user's set of registered daemons (
url://peer address, display name) and is what actually opens theFabricApiClientconnection to each one on the account's behalf. -
[ ] Gate daemon access with W3Wallet capabilities, not bare account ownership. Being in an account's peer registry is not itself authority to use a daemon: the ServiceServer grants a "may access daemon D" capability per registered peer, and every list/health/function-invoke call against that daemon requires holding it — the same capability-is-the-authority model CockroachDb and AiCliHostSupervisor already use. This is what lets a peer eventually be shared with a teammate (grant them the capability) without re-registering it or handing over an account.
-
[ ] Mint and resolve registration tokens. When a user starts the "create a new node" flow in the WUI (below), the ServiceServer generates a single-use registration token bound to that user's account and records it pending; when a freshly-launched daemon calls back with that token, the ServiceServer resolves it to the initiating user and adds the daemon's
url://peer address to that account's peer registry. -
[ ] Drive the SSH remote-install path in the background. Given a target address/port/user and a private key or password, the Embedded opens the SSH connection itself (never the browser), installs coursier if it isn't already present, and runs the same
cs launch … --register-with … <registration-token>as the manual path — reporting job progress back to the WUI (connecting, installing, launching, registered/failed) rather than blocking the request on a long-running SSH session. -
[ ] Optionally remember SSH credentials, per account, as W3Wallet Proxy capabilities — two distinct things, not one. If the user checks "remember these credentials," the ServiceServer saves, encrypted at rest:
- A reusable identity (user + private key or password, no host) the user can select for installing on other machines later, instead of retyping it.
- A per-daemon operational credential (host + port + user + private key or password), linked to this specific daemon's registry entry — kept so the ServiceServer can SSH back into that same machine later for operational recovery (restarting or redeploying the daemon if it stops responding over
url://), the same job ssh-restart-fabric-daemon / ssh-deploy-fabric-daemon do by hand today.
Either way, what the user holds afterward is a "may use this credential" W3Wallet capability, never the raw secret: invoking it drives an SSH connection from inside the ServiceServer and returns the outcome, so the private key or password never round-trips back out to the WUI or browser.
-
[ ] Bring an offline daemon back via its stored per-daemon credential. When the WUI's "Bring back online" action fires, the Embedded exercises the daemon's stored operational credential to SSH in and kill+relaunch the daemon process — folding in ssh-restart-fabric-daemon's job (including the mandated
-Xmx128m -XX:+ExitOnOutOfMemoryErrorflags from its still-unmerged PR #10) directly into the ServiceServer, rather than requiring an operator to reach for the standalone CLI.
HardwareControlFabricWui
- [ ] Peer list. List the peers registered to the current account (via the ServiceServer).
- [ ] Function invocation. Let an operator invoke registered functions on nodes from the dashboard (via the ServiceServer).
- [ ] Health dashboard. Collect health stats from the daemons by default and report them in a control panel dashboard (via the ServiceServer).
- [ ] One-button recovery for an offline daemon, when a per-daemon credential is on file. If a registered daemon shows offline/unreachable in the health dashboard and the account has a stored per-daemon operational credential for it, the dashboard offers a single "Bring back online" button next to it. Pressing it asks the ServiceServer to SSH in and restart the daemon (below); progress (connecting, restarting, back online/failed) surfaces the same way as the install-flow job status. With no stored credential for that daemon, no button is offered — the daemon just shows offline.
- [ ] Add peer. A button offering three paths:
-
Register an existing node. Enter the address of an already-running daemon to add it to the account's peer registry.
-
Create a new node. The WUI requests a fresh registration token from the ServiceServer and displays a single command the user copies onto the target machine — using coursier to fetch and launch the latest daemon straight from
kotlin.directory, with a flag carrying the ServiceServer's address and the registration token so the new daemon can call back and be associated with the right account:cs launch --repository https://kotlin.directory \ community.kotlin.hardwarecontrolfabric:daemon:latest.release -- \ --register-with url://hardware-control-fabric/register/<registration-token> -
Install remotely via SSH. A form asking only for the target's address, port (default
22), username, and either a private key or a password, plus an optional "remember these credentials" checkbox. Submitting it hands the details to the ServiceServer (never the browser), which connects out over SSH in the background and performs the same install as the manual path — no command for the user to copy at all.
-
HardwareControlFabricDaemon — peer registration mode
-
[ ]
--register-with <url>triggers an interactive confirmation, not a silent registration. Launching the daemon with this flag does not immediately register; it first prints an explicit, hard-to-miss warning and requires a typed confirmation, defaulting to No on empty input:⚠️ WARNING: This will register this machine with the HardwareControlFabric server at url://hardware-control-fabric/register/<registration-token>. The operator of that server will gain FULL, UNRESTRICTED CONTROL of this machine: they will be able to run arbitrary code as this user, read and write any file this user can access, and manage every process on this host. Only continue if you specifically intend to grant that server full control of this machine. Continue? [y/N]:Anything other than an explicit affirmative (default on Enter/EOF/any non-"y" answer) aborts without registering.
-
[ ] A deliberately shouty flag to skip the prompt, for scripted/non-interactive provisioning only. Something like
--yes-i-understand-this-grants-full-remote-control-of-this-machine— single-shot (applies to that one registration attempt only, never persisted) and named so that anyone copy-pasting a command that includes it is confronted with what they're agreeing to, rather than a terse-y/--forcethat hides the stakes. This is exactly the flag the SSH remote-install path (above) passes on the user's behalf — see Decisions for why that isn't a loophole.
Testing
- [ ] End-to-end integration tests via NetLab. Exercise the WUI and the fabric it drives across multiple simulated topologies/configurations (e.g. several daemons, peers reachable directly vs. behind a NAT) rather than validating only against the single production host.
Decisions
Resolved as the design firmed up; kept here so the rationale survives.
-
The WUI is backed by a ServiceServer, not a direct daemon connection. Per the standard layered architecture, HardwareControlFabricWui talks only to
url://hardware-control-fabric/; it never opens a connection to a daemon itself. This resolves what was previously an open question here (whether the WUI would callHardwareControlFabricApidirectly, matching the CLI's older pattern). -
The daemon itself moves onto
url://, replacing its bespoke mTLS HTTPS API. Rather than the ServiceServer (or anything else) growing a second way to reach a daemon, the daemon's own command surface becomes a normalurl://endpoint over the libp2p connection it already holds for P2P discovery — one mesh, one transport, for the CLI, the ServiceServer, and any future consumer alike. -
Joining the mesh is a token-mediated handshake, not an open door. A new daemon only ever gets associated with an account because that account's WUI session minted a single-use registration token first; the daemon calling back with a stale, guessed, or reused token is the only thing standing between "a user I invited" and "anyone who finds the command," so the token must be single-use and expire. This handshake is the fabric's realization of the ecosystem-wide Manager–Daemon Pairing pattern — the same "discoverability is not authority" negotiation, carrying a token in the launch URL and terminating in a W3Wallet capability rather than an identity, that a self-hosted daemon and its manager use everywhere on the platform.
-
Granting control is a loud, deliberate act, not a default.
--register-withstops at an interactive, default-No confirmation naming exactly what is being granted (full control of the machine, to whoever operates that server), and the only bypass is a flag whose name itself carries the warning — so even a blindly copy-pasted automation script can't silently skip informed consent. -
The SSH remote-install path is allowed to pass the bypass flag itself, because the consent already happened in the WUI. The scary confirmation exists to stop someone from unknowingly granting a stranger's server control of their machine — e.g. a copy-pasted command whose implications weren't read. When the ServiceServer runs the install itself, the user has already, from inside their own authenticated WUI session, deliberately pointed it at a specific host and asked it to install and register a daemon under their own account; there is no second party's command to blindly trust. So the ServiceServer supplying the bypass flag on the user's behalf is the same consent already given, not a way around it.
-
W3Wallet capabilities, not raw secrets or bare ownership, are the unit of authority for both stored credentials and daemon access. A remembered SSH credential is held by the ServiceServer and only ever exercised on the holder's behalf via a capability — the secret itself never leaves the boundary that stored it. Access to a registered daemon is likewise a capability, separate from the fact of being in the peer registry. This is the same "capability is the authority, secrets stay behind the proxy" model W3Wallet already establishes elsewhere in the ecosystem, applied here from v1 rather than deferred as a hardening pass.
-
Removing a peer is an instant, one-directional capability revocation, scoped to the manager that revoked it. The "may access daemon D" grant is a W3Wallet reference capability held at the manager (ServiceServer/account) that registered it, the authority home for that daemon connection; revoking it is deleting the registry entry, which rejects the grant credential on the next call. Nothing round-trips to the daemon itself, and no other manager's independent grant to the same daemon is affected — each registration is its own capability, not a shared one. Re-adding the same daemon to that manager later is a fresh registration (new token, new capability); there is no "restore" path.
-
The daemon has one authorization boundary, not two. UrlResolver's peer-identity/name-ownership verification only protects server authenticity (a client can't be routed to an impostor serving a name); it says nothing about which callers a service accepts — that's each service's own choice, layered via W3Wallet where required. So the daemon does not additionally maintain a mesh-level peer-identity allow-list alongside the capability model: its
url://command surface itself requires a W3Wallet capability to invoke, and each completed--register-withhandshake grants one more such capability, naming that manager as an authorized caller going forward. Daemon-side trust of a manager and that manager's own ownership-gating of daemon access are therefore the same mechanism, not two independent trust decisions. -
Daemons and managers are many-to-many. A single daemon can hold independent grants for, and be reachable through, more than one manager (ServiceServer/account) at once — e.g. a personal account and a team account both managing the same physical machine — and a single manager's peer registry can of course hold many daemons. Each
--register-withhandshake is additive: it adds one more independent grant to the daemon's authorized-caller set rather than replacing whatever was there before, and one manager revoking its own grant (above) has no effect on any other manager's relationship with that same daemon. -
Remembered credentials are two distinct things, not one generic store. "Remember these credentials" saves (a) a reusable, host-agnostic identity (user + secret) the account can apply to a different future target, and (b) — independently — a per-daemon operational credential bound to this host, kept specifically so the ServiceServer can SSH back in later to restart or redeploy this same daemon if it stops responding over
url://. Both are held as W3Wallet Proxy capabilities, not raw secrets; "forget this credential" deletes either one outright (the same instant, one-directional revocation shape as removing a peer), and saving happens regardless of whether the install attempt that prompted it succeeded, since the credential's validity is independent of that one connection attempt.
Open questions
- Does the one-button recovery path retire the standalone SSH tools, or sit alongside them? ssh-restart-fabric-daemon (and, for redeploys, ssh-deploy-fabric-daemon) now overlaps materially with the ServiceServer's own recovery action. Whether the standalone CLIs get retired in favor of the WUI/ServiceServer path, or deliberately kept as an independent break-glass option for when the ServiceServer itself is the thing that's down, is undecided.
Related
- ContainerNursery — the fabric's function-plugin mechanism exists mainly to restart it today.
- NetLab — supplies the simulated topologies the e2e tests run against.
- UrlResolver — the
url://fabric the daemon and ServiceServer both move onto. - W3Wallet — supplies the capability model gating both stored SSH credentials and daemon access.
- Manager–Daemon Pairing — the general architecture pattern this workstream's token-mediated, consent-gated registration handshake realizes.
Graduation
This graduates to a first-class project once the daemon speaks url:// in place of its bespoke mTLS API, HardwareControlFabricServiceServer is deployed at url://hardware-control-fabric/ with W3Wallet-gated credential and daemon-access capabilities, HardwareControlFabricWui is built and deployed with peer listing, function invocation, health dashboard, the add-peer flow (registering an existing node, the confirmation-gated coursier one-liner, and background SSH install), and one-button recovery of an offline daemon via its stored credential, the NetLab-driven e2e tests are green, and it has a page in the Documentation Repository.