← Workstreams

Repository · workstreams

Workstream: BuildTest Runner Reporting — retiring the long-held per-shard SSH connection

View on GitHub ↗

Workstream: BuildTest Runner Reporting — retiring the long-held per-shard SSH connection

Status: Planned · Component: Maximize developer productivity

2026-07-25 revision — read this first

This workstream previously aimed to eliminate SSH entirely, via an always-on url:// runner-agent daemon baked into the droplet snapshot. That is no longer the plan.

The revision came from a simple question: why must the reporting service be a separate process from the build tool? The original answer — that the coordinator needs something to answer "what is the status?" while the build runs — does not hold. The build-tool JVM has threads; it can serve queries about itself while it builds. There was no constraint, only an assumption.

Following that thread produced a sharper framing of the problem:

The expensive thing was never SSH. It was the long-held connection — a socket pinned per shard with a blocking read thread on it, for the entire build.

Short-lived SSH operations cost essentially nothing: no pinned socket, no blocking reader, no thread-per-shard. So the revised plan keeps SSH for lifecycle and pathology, and moves only the live reporting onto url://, served by the build-tool JVM itself.

What this buys, relative to the daemon design:

  • No droplet-snapshot change. No agent to bake, no systemd unit, no image rebake to ship a fix. The build tool is already fetched and pinned per run, so reporting code updates with an ordinary version bump instead of an image roll.
  • No new always-on process to supervise, version, or canary on the organization's critical CI path.
  • No version skew between an agent frozen in a slow-moving image and a coordinator that ships continuously.
  • The same measured win: the held connection and its blocking reader disappear.

What it gives up: SSH does not go away. Key/host management and the JSch dependency remain. That is a deliberate trade — SSH elimination was a purity goal, not the cost driver, and it can still be pursued later on top of this work if it ever justifies its own risk.

Three of the arguments the daemon design rested on did not survive examination:

  1. "Something must answer queries while the build runs." The build JVM can answer for itself.
  2. "Something must outlive the build process to serve its results." A short SSH session reads the on-disk spool and heap dump post-mortem — same recovery, no daemon.
  3. "Readiness by announcement replaces provider polling." The polling problem was already solved by the DigitalOcean gateway's shared cache (recorded in the buildtest-orchestration handoff's own 2026-07-12 correction). This was solving a solved problem.

Goal

Retire the persistent per-shard SSH connection. The build-tool JVM on the droplet hosts a url:// service and reports its own progress, results and build artifact over it; SSH remains, but only as short-lived sessions for machine lifecycle and for pathology (a wedged or dead JVM).

The division of labour:

url:// carries the happy path. SSH carries birth, death, and pathology — and no SSH session is held open across a build.

Which channel carries what

Operation Today Revised Why
Readiness provider poll + SSH connect-retry short SSH connect The connect is the readiness proof; the provider-poll cost is already solved by the DO gateway cache.
Stage the project archive SFTP on the held connection short SFTP One bulk upload, then done. Nothing to hold open.
Launch the build exec on the held connection short SSH exec Fire and return; the build outlives the session.
Progress, events, test results held connection, blocking read thread url:// from the build JVM The only expensive row. This is the whole point of the workstream.
Build artifact SFTP download url:// chunked pull Bulk binary, but from a live healthy JVM — output is only requested when the build succeeded.
Cancel (graceful) close the session to unblock the read url:// cancel Lets the runner flush its spool and finalize in-flight tests, rather than being shot mid-write.
Cancel (escalation) — short SSH kill See the escalation rule below.
Thread dump (wedged) jcmd over the held connection short SSH jcmd See below — a wedged process cannot be trusted to dump itself.
Heap dump (out of memory) SFTP download post-mortem SSH By definition the JVM is dead or dying.
Post-mortem after the JVM dies reattach over SSH short SSH, read the on-disk spool The evidence is on disk; nothing needs to be alive to serve it.

Cancel is graceful over url://, with an SSH escalation

Graceful beats a signal: the runner can flush its event spool, finalize the tests actually in flight, and record a clean cancelled state instead of leaving the coordinator to infer one from a truncated file.

But "wedged" is simultaneously when cancellation matters most and when the RPC will not answer. So cancellation is budgeted: ask over url://, and on timeout escalate to a hard kill over a short SSH session. This composes with the watchdog sequence — dump the threads, ask nicely, then kill.

A wedged process cannot be trusted to dump itself

A thread dump exists precisely because the process is misbehaving — deadlocked, or thrashing the collector. Asking it to dump itself through its own RPC handler is asking the patient to perform the autopsy; the service thread is plausibly the starved one. jcmd from outside, over a short SSH session, is the robust path.

The spool changes role, it does not disappear

Today the on-disk event spool is the transport: SSH reads it by line offset. Under this plan the live url:// stream is the transport, and the spool becomes the crash-recovery artifact that the post-mortem SSH path reads. The build JVM must still write it as it goes — otherwise a JVM that dies leaves nothing behind, and the heap-dump case guarantees that happens.

The line-addressed cursor semantics therefore stay meaningful; they just serve recovery rather than steady state.

test-results.xml retires — but not yet, and not for the reason first recorded

Corrected 2026-07-25 after auditing the projection rather than only the file's readers.

The first version of this section claimed the XML was "a second encoding of data the coordinator already has." That was based on checking who reads the file — no consumers in BuildTestWui, BuildTestServerService or BuildTestCli, and parseTestResultsXml confined to BuildDriver. That was the wrong question. The right one is what the XML contributes to the projection, and the answer is: real, load-bearing data.

The XML arrives as an xml_results event and RunEventProjection reconciles it against test_completed with explicit precedence — when a winning attempt's completion disagrees, the winner keeps status/passed/errorMessage and the XML is allowed to "enrich timing/memory". It is a second reconciled source, not a duplicate.

Audit result — what the XML uniquely supplied:

Field Covered by events?
Test name, status, disabled Yes — test_completed
Duration Yes — test_completed.durationMs
Console output Yes — test_console_output
Stack trace No — the events carried Throwable.message only
Peak memory on ambiguous rows No, and cannot be — see below

The stack-trace gap is closed. TestCompletedEffect.error is a Throwable; the runner rendered only .message and discarded the trace, so the trace reached consumers exclusively through the XML's <stacktrace> element. BuildTestRunner 0.0.71 now emits the full trace and causal chain on the event, bounded so one pathological failure cannot dominate the append-only event log.

The peak-memory gap is not a bug and does not close here. When two tests share a package and function name, BuildTestRunner deliberately emits peakMemoryBytes = 0 rather than guess — TestPeakMemoryTracker's own contract says the TestCompletedEffect payload "cannot identify the source file" and callers should "treat peak memory as unavailable for those ambiguous rows instead of guessing." The XML can attribute it, because it carries file-level identity. So retiring the XML costs peak-memory data on ambiguous same-named tests until kompile-core's TestCompletedEffect gains source-file attribution.

Therefore: the XML retires last, after the transport work, and either (a) TestCompletedEffect gains file attribution upstream, or (b) losing peak memory on ambiguous rows is accepted explicitly. It is no longer a cheap early deletion, and it must not be treated as one.

The general lesson is worth keeping: "nobody reads this file" is not the same as "this file carries nothing unique."

Current state

Today each shard opens one persistent JSch SSH connection, held for the shard's entire lifetime (BuildTestEmbeddedService.runShardPipeline → BuildDriver over SshExecutor). It carries, in order:

  1. Exec of setup commands — near-noop on the pre-baked snapshot image ("Dependencies already present (snapshot image), skipping installation").
  2. SFTP upload of the project archive (ssh.uploadFile), then a remote tar extract.
  3. A live stdout stream — kompile-cli runs on the droplet and its ##BTR## event stream + build log stream back in real time (ssh.execStreamingAttempt, with an idle-timeout watchdog). This is the only reason the socket is held open for the whole run.
  4. SFTP download of test-results.xml, the build artifact, and heap dumps.

Readiness is DropletManager.waitForActive polling getDroplet plus SSH connect-retries; cancellation is "close the SSH session to unblock the read." Two recent efforts changed the surroundings but not the transport: BuildTestEmbedded #304 made the per-shard orchestration threads virtual (cheap, but still one blocking read per shard), and the DO gateway (DigitalOceanDropletServiceServer #86 / BuildTestEmbedded #287) plus the admission redesign already removed the DO API-rate / droplet-quota failure modes.

The transport is already poll-based, not streaming. BuildDriver.attachToRunnerAndDrainEvents(startAtLine, …) reconnects to a durable event spool file on the droplet and seeks to the first unconsumed 1-based line; the runner never writes into a live socket. The held-open connection carries only command execution, SFTP file transfer, and line-offset file reads. Had it genuinely been streaming, this would have needed a redesign rather than a re-hosting — and the lossless re-attach property the coordinator relies on carries over to either plan unchanged.

What already exists (2026-07-25)

A four-layer url:// service set was built and published under the previous (daemon) plan, before the revision above. None of it is deployed; nothing touches the CI path; the artifacts are unused 0.0.1s.

Repository Artifact Fate under the revised plan
BuildTestRunnerAgentApi buildtest.runneragent.api:buildtest-runner-agent-api:0.0.1 Reporting half carries over (event-spool cursor, runner status, work queue, artifact pull, cancel). Staging/launch operations become SSH again and drop out.
BuildTestRunnerAgentEmbedded …:buildtest-runner-agent-embedded:0.0.1 Largely superseded — the build JVM knows its own state and need not read back a spool file it wrote. The spool writing and the work-queue semantics survive, relocated into the build tool.
BuildTestRunnerAgentServiceServer …:buildtest-runner-agent-service-server:0.0.1 (+ -client) RPC-marshaling layer carries over; the per-droplet registration and daemon main do not.
BuildTestRunnerAgentCli buildtest.runneragent.cli:buildtest-runner-agent-cli:0.0.1 Carries over — still the way to drive one droplet by hand during the cutover.

Pins across the set: foundation.url:resolver:0.0.653 with foundation.url:protocol:0.0.397 (jvm-libp2p snapshot-19) — the ContainerNursery-validated pairing. Do not move either alone; a mismatched pair presents as small calls working while large responses hang.

Two findings from building it, worth keeping regardless of plan:

  • A runner must be started at most once per run. The end-to-end test caught an implementation that refused only while a runner was still alive: a coordinator restarting after a build had finished would have re-launched the whole build instead of attaching to the recorded result. The guard must be "was a runner ever started for this run id", which is why the runner's process id belongs on disk rather than in memory.
  • Reads must never serve half a line. The spool is written while it is read, so an unterminated trailing fragment has to be withheld until its newline arrives — otherwise half an event is delivered, and its other half arrives as a separate "event" on the next read.

Why it accelerates developers

  • The pinned socket and its blocking reader go away. This is the measured cost from the original forensic — one held JSch connection per shard, with a blocking read thread, for the entire build. #304's virtual threads made those threads cheap; this makes them unnecessary.
  • Reporting ships with the build tool. No image rebake to fix a reporting bug — the build tool is fetched and pinned per run.
  • One fewer moving part in production. No always-on daemon on every droplet to supervise, version, or reason about when it is the thing that has failed.
  • The on-ramp to the runner vision is unchanged. ExecutionEnvironments §C's dispatch plane rebinds its work channel to url:// calls exactly as it would have; §C/§B own scheduling and statelessness either way.

Plan / roadmap

1. The build tool serves url://

  • The droplet-side build tool (BuildTestRunner) hosts a url:// service for the run it is executing: read events from a cursor, report status, pull the build artifact, accept dispatch work, accept a graceful cancel.
  • It continues writing its durable event spool to disk — now for crash recovery rather than as the transport.
  • It announces a per-run service identity so the coordinator addresses the specific build it launched.

2. The coordinator stops holding a connection

  • BuildTestEmbedded provisions and stages over short SSH sessions, launches the build, and then closes the connection.
  • Live reporting comes over url://. ssh.execStreamingAttempt and the per-shard blocking reader retire.
  • Pathology paths — thread dump, hard-kill escalation, post-mortem spool and heap-dump reads — open a fresh short SSH session when needed.

3. Retire test-results.xml — last, not first

  • Drop the XML round-trip once the transport work has landed. The stack-trace gap is already closed (BuildTestRunner 0.0.71); the remaining decision is whether to close the ambiguous-row peak-memory gap upstream in kompile-core or to accept losing it. Confined to BuildTestEmbedded; no downstream file consumers.

4. Cut over

  • Land the url:// reporting path alongside the held-connection path behind a per-run/per-shard switch, default off.
  • Prove one droplet by hand with BuildTestRunnerAgentCli before the coordinator is touched. This is the first genuine test of the transport: announcement, relay/NAT traversal and sandbox execution are exercised by deployment, never by unit tests.
  • Canary one shard, then one run; compare results and latency; flip the default; then delete the streaming-read path.

Decisions

  • Retire the held connection, not SSH (decided 2026-07-25, revising the original "eliminate SSH entirely via a snapshot-baked daemon"). The cost driver measured by the original forensic was the pinned socket and its blocking reader, not SSH itself. Short-lived SSH for lifecycle and pathology is cheap and needs no image change; a daemon in the snapshot is a slow-moving artifact with version skew against a continuously-shipping coordinator.
  • The build-tool JVM hosts the reporting service. There is no separate agent process. The earlier justification for one — that something must answer queries while the build runs — was an assumption, not a constraint.
  • Cancellation is graceful over url://, escalating to a hard kill over SSH on timeout. A wedged process is exactly the one that will not answer its own RPC.
  • Thread dumps are taken from outside over SSH. A misbehaving process cannot be trusted to dump itself.
  • The event log becomes the single source of truth for test results, and test-results.xml retires last — not as an early cheap deletion. An audit found the XML uniquely supplied stack traces (now fixed in BuildTestRunner 0.0.71) and peak memory for ambiguous same-named tests (which the events cannot honestly supply until kompile-core's TestCompletedEffect carries the source file). See the corrected section above.
  • Transport-only first, scheduling/statelessness later (§C/§B). Unchanged: the droplet stays a full workspace that uploads and builds, so this phase does not have to solve §B's eager-artifact-closure problem.
  • The §C dispatch plane does not wait for this (decided 2026-07-15). The lease-based pull queue with in-run re-lease and tail rescue lands inside BuildTestEmbedded over the existing SSH channel via an on-droplet queue file; the scheduler core is transport-agnostic, and only the thin queue-file adapter is later replaced by url:// calls.
  • Provisioning stays on the DO gateway. No new droplet-lifecycle machinery is invented here.
  • The streaming-read path is removed only after a production canary proves the url:// path, and stays a reversible fallback until then.

Open questions

  • Does the build tool's url:// service need per-run or per-droplet identity? Per-run is the natural fit now that the service lives inside the build process, but a droplet running successive builds then changes identity between them; whether the coordinator cares is worth settling when the contract is specified.
  • Artifact pull versus push. The build artifact is fetched from a live JVM today's plan; whether it is better pushed as it is produced (fewer round trips) or pulled (retryable, coordinator-paced) is unsettled. Pull is assumed above.
  • Heap-dump transfer shape. Multi-hundred-megabyte dumps go over post-mortem SSH under this plan, which is simple but slow. Whether they should instead land in a content-addressed store is a transfer-shape decision, not an architectural one.
  • url:// dependency surface inside the build tool. The build tool already carries Jackson, Guava, OkHttp, Moshi and zstd, so adding the resolver/protocol stack is more of the same rather than a new category of risk — but it is worth measuring its effect on build-JVM startup and footprint before committing.

Related

  • Pluggable Execution Environments — §C is the runner-vision end state this phase feeds; §B the statelessness it defers.
  • Kompile Remote Build — the CI orchestration + optimizer that sits above url://buildtest/.
  • NetLab — the manager/worker-droplets-over-url:// precedent. Note it is a worker daemon pattern, which this workstream deliberately does not copy: a NetLab worker outlives many jobs, whereas a buildtest droplet serves one build.
  • UrlResolver — the url:// transport, relay, and sandboxed-connection layer this rides on.
  • The buildtest-orchestration handoff for how the surrounding resource-cost issues were already resolved.

Graduation

This workstream is complete when no SSH session is held open across a build: the build tool reports over url://, BuildTestEmbedded's per-shard streaming read and its blocking reader are deleted, test-results.xml is retired in favour of the event log, and the remaining SSH uses are short-lived sessions for provisioning, staging, launch, cancellation escalation and post-mortem. Eliminating SSH altogether is explicitly out of scope — it may be taken up later, on top of this work, if it ever justifies its own risk.