← Handoffs

id: hf-2026-09-24-stop-buildtest-s-persistent-rpc-transport-to-containernursery-from-being-retired-mid-request-droplet-provisioning-failures-since-08-40z-2026-09-24 url: url://handoff/handoffs/hf-2026-09-24-stop-buildtest-s-persistent-rpc-transport-to-containernursery-from-being-retired-mid-request-droplet-provisioning-failures-since-08-40z-2026-09-24 title: Deploy ContainerNursery main with the bridge application-error fix (PR 634) so a failed request stops disconnecting its siblings summary: Root cause verified: deployed ContainerNursery (unmerged PR 625 head a2fc6415) disconnects a client's whole persistent RPC connection whenever one request fails, killing pending siblings. Fixed on main by PR 634, not deployed. Droplet provisioning symptom gone since 09-29 04:21Z; the same symptom is live for github-watchman and buildtest WUI callers. Remaining: settle PR 625, then a user-authorized CN deploy and log verification. created: 2026-09-24T11:26:15.171Z completed: null blocked-reason: Needs user authorization for a ContainerNursery production deploy from main (>= d3acb54, PR 634) and a decision on PR 625. dependencies:

K52 re-verification (2026-10-01 00:20 UTC)

RE-VERIFY: write-time snapshot. Re-check PR states (GraphQL), the live ContainerNursery jar and process, and /api/runs before acting. This worker merged, enqueued, deployed, restarted and published nothing.

Remaining work (only this)

  1. Decide and perform a ContainerNursery production deploy from main at or after https://github.com/CodexCoder21Organization/ContainerNursery/commit/d3acb54ea355e9428b0301cc611f6359ea486496 (PR https://github.com/CodexCoder21Organization/ContainerNursery/pull/634). Production deploys need explicit user authorization.
  2. Before that deploy, settle https://github.com/CodexCoder21Organization/ContainerNursery/pull/625. It is OPEN and DIRTY. Production CN currently runs an old head of that unmerged PR (a2fc6415). That build carries 3 commits main lacks: "Return byte arrays larger than 4 MiB through the URL facade", "Pin foundation.url:protocol 0.0.546" and a CI re-run commit. Deploying plain main would drop those. Either land a rebased 625 first, or accept the regression explicitly.
  3. After the deploy, verify: the CN process log stops showing bursts of "Persistent RPC connection to service '<svc>' was closed while requests were still pending" from local clients (github-watchman integration server, buildtest WUI). Then complete this handoff.

Mechanism (firsthand evidence)

The persistent RPC transport is closed mid-request by ContainerNursery's URL facade. UrlResolver's PersistentRpcConnection retirement logic does not do it, and neither does BuildTestEmbedded.

In the deployed CN build, any failure of one request on a push-capable persistent connection closes the shared BackendPushBridge with disconnectOutside=true. That includes an ordinary application error, such as "cancelCiBuild cannot cancel a COMPLETED build", INVALID_PARAMS or a DigitalOcean 404, and also a 30 s handler timeout. The close disconnects the client's outside connection. Every sibling request still pending on the client's PersistentRpcConnection then reads EOF ("Stream closed while reading message data (read 0 of 4 bytes)"). Each one fails as AmbiguousRpcRequestException and is not replayed.

  • OBSERVED: production CN is PID 726140, running /root/ContainerNursery/bin/container-nursery-a2fc6415-20260926T235430Z.jar (started 2026-09-27 00:18Z). Comparing PR 634's merge commit d3acb54 with a2fc6415 shows the branches diverged, with a2fc6415 12 commits behind, so it lacks the fix. javap of the deployed class UrlFacadeProvider$registerService$handler$1 (bytecode 761-796) shows the forward-failure handler calls BackendPushBridge.close$default(disconnectOutside=true, "a request on the backend bridge failed: ...") unconditionally, with no BackendRpcApplicationException check.
  • OBSERVED: CN main (dab37c7) skips that close for BackendRpcApplicationException. It maps a local 30 s deadline (LocalDeadlineExpired) to BackendRpcApplicationException, a mapping in place since PR 619. Only transport failures to the backend still close the bridge.
  • OBSERVED: PR 634's red-first real TCP/libp2p test failed on pre-fix main (0 passed, 1 failed in https://buildtest.kotlin.build/run?id=fa035671) and passes on the fix.
  • OBSERVED, still live: the CN process log covering about the last 1.5 h holds 90 "closed while requests were still pending" failures. 66 are github-watchman invalidateRepository/invalidateResource calls from the github-watchman integration server. 24 are buildtest __projection_poll/getProjectionChanges/getTestAttemptResultsPaginated calls from the buildtest WUI. They arrive in bursts of 6-15 siblings at once (23:32:59, 23:33:42, 23:39:17, 23:53:48). Application errors and 30 s TIMEOUT lines are frequent in the same log.
  • INFER: the CN log does not name the service on "[UrlFacade] RPC error" lines, and the bridge close is logged only at debug level, so the specific request that triggered each burst is not individually proven.

Droplet provisioning symptom: no longer occurring

  • OBSERVED: across /api/runs pages 1-4 (2000 runs, 09-28 08:31Z to 09-30 23:50Z), 52 runs failed with "Persistent RPC connection to service 'digitalocean-droplets' was closed while requests were still pending". The last was e05639a3 at 09-29 04:21Z, and there have been none in the roughly 43 h since.
  • OBSERVED: today's droplet failures (09-30 09:33-13:54Z: 4ca14842, 9e981129, da2842db, 982b1425, 310391a1, 723cc4d9, plus five "Could not verify droplet cleanup ... CancellationException") are a different class. The droplet service gave no reply within 30 s (DropletServiceCallStalledException, or a sandbox invalidated after a 30 s timeout), which is not a transport closed under pending requests. Today's other fleet failures are the shared-cache preparation deadline (22 runs) and ordinary test failures. Both are outside this handoff.
  • OBSERVED: production url:buildtest: runs /root/ContainerNursery-uploads/jars/buildtest-server-59e32017-20260930.jar (BuildTestServerService main 59e3201, PID 327891, started 2026-09-30 20:25Z). Its PersistentRpcConnection.class sha256 is e40537ba344adffce9a8bee8509f6f371acac88c416eeab0d511790e194572c5 (resolver 0.0.1259, sibling-drain guard). Every buildtest jar on the host since 09-28 05:10 has the same class. The earlier blocker, "deploy candidate 4012e862cde8", is obsolete: a newer main was deployed by another operator.

PR / branch state (re-verified 2026-10-01 00:05Z)

Item State Link
Embedded release MERGED 2026-09-27T08:23Z, head 74e2d183, 5/5 checks SUCCESS https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1107
Server pin MERGED 2026-09-27T09:49Z, head cb2b039a, bld-build SUCCESS https://github.com/CodexCoder21Organization/BuildTestServerService/pull/373
CN bridge application-error fix MERGED 2026-09-27 as d3acb54e; NOT deployed https://github.com/CodexCoder21Organization/ContainerNursery/pull/634
CN large-response + protocol 0.0.546 OPEN, DIRTY, last push 2026-09-27; an old head of it is what production runs https://github.com/CodexCoder21Organization/ContainerNursery/pull/625
UrlResolver sibling-drain guard MERGED 2026-09-14; deployed in buildtest https://github.com/CodexCoder21Organization/UrlResolver/pull/1088
UrlResolver 1122 OPEN, another actor pushed 2026-09-30T23:31Z, unrelated (bootstrap resync traces); not touched https://github.com/CodexCoder21Organization/UrlResolver/pull/1122
CN W26 evidence branch 69ab3053 https://github.com/CodexCoder21Organization/ContainerNursery/tree/wip/droplet-rpc-bridge-error-W26

CN main head dab37c78 (2026-09-30T22:39Z) had kotlin.build checks IN_PROGRESS at write time. Confirm it is green before building a deploy candidate from it.

No UrlResolver, BuildTestEmbedded or KompileRemoteBuild code change is needed for this handoff, and K52 opened no PR.


Earlier handoff history (superseded; retained for provenance)

W42 checkpoint (2026-09-27 04:44 UTC)

RE-VERIFY: This supersedes the earlier W42 status table below. Current CI and PR state must be checked before acting. The existing production blocked reason remains; this worker did not deploy, merge, enqueue, or alter a route.

Item Write-time state Link
Embedded release PR OPEN at 1f9fee3cc1fc8358b65cc0fc5c574b5d1f03b8c4; four fresh CI shards in progress https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1107
Server pin PR OPEN at 30646e217ab4f4b8760784bc4dcbda43466a3fdb; CLEAN and bld-build SUCCESS, 198/198 tests passed https://github.com/CodexCoder21Organization/BuildTestServerService/pull/373
Embedded release branch Pushed at 1f9fee3cc1fc8358b65cc0fc5c574b5d1f03b8c4 https://github.com/CodexCoder21Organization/BuildTestEmbedded/tree/wip/finished-shard-embedded-release-W42
Server pin branch Pushed at 30646e217ab4f4b8760784bc4dcbda43466a3fdb https://github.com/CodexCoder21Organization/BuildTestServerService/tree/wip/consume-finished-shard-release-W42

Embedded first CI had 3/4 green shards; shard 3 failed an unrelated exact-boundary test by 1ms: https://github.com/CodexCoder21Organization/BuildTestEmbedded/actions/runs/36293011316/job/108546565816 . The scheduled expiry records 1200000ms, while the test advanced a manual clock to 1200001ms and sometimes expected the later observation. The test now forces the exact boundary: old expectation failed 0/1 locally; corrected assertion passed 1/1. Production source and the published embedded JAR are unchanged. Fresh CI: https://github.com/CodexCoder21Organization/BuildTestEmbedded/actions/runs/36294972576 . Do not rerun the old failed job blindly.

The server CI run https://github.com/CodexCoder21Organization/BuildTestServerService/actions/runs/36293795824/job/108548754049 passed its build and unfiltered test suite (198/198), including the real-droplet sharding test pinned to embedded 0.0.69270330. Published embedded JAR SHA-256 is 29974203a53e8782f7c100d007cb1775e4d42ccc29f4db261e62aa53e42f3570. Local server candidate remains /tmp/claude-1000/-code/1f0f7d17-dc24-4d3a-a8ff-48753bfab06a/scratchpad/cx/lanes/W42/buildtest-server-0.0.69270330-candidate.jar, SHA-256 18441febc234bef8aadf2687e60c83f7adb9c276dd0426463fc1882a0b289012. It includes all 1674 embedded classes byte-for-byte and the expected resolver class. The local candidate was not uploaded.

Remaining: Re-verify both PRs; wait for the embedded fresh CI and fix any real failure; have the orchestrator review final heads and merge or enqueue both PRs. After main includes them, rebuild the server candidate from the merged main if source or dependencies changed, check its SHA-256 and current live route, then ask the user whether to deploy. The original extra-droplet failure no longer stands as an unexplained product risk once this server CI result is accepted; production rollout still needs the user decision.

Earlier handoff history

W42 release status (2026-09-27 04:22 UTC)

RE-VERIFY: Every state below is a write-time snapshot. Check the linked pull requests and published artifact before acting. The production blocked reason remains in force; no deployment, upload, restart, or route change is authorized by this note.

The finished-shard fix is merged in https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1105 . A retrying dynamic shard no longer provisions a replacement after every scheduled result is durable. The old real-droplet failure was a product behavior, not test isolation.

Item Write-time state Link
Embedded release PR OPEN; head b0082a5632c523bbcdd2fdcebcda34ee74efd623; four CI shards in progress https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1107
Server pin PR OPEN; head 30646e217ab4f4b8760784bc4dcbda43466a3fdb; bld-build in progress https://github.com/CodexCoder21Organization/BuildTestServerService/pull/373
Embedded release branch Pushed at b0082a5632c523bbcdd2fdcebcda34ee74efd623 https://github.com/CodexCoder21Organization/BuildTestEmbedded/tree/wip/finished-shard-embedded-release-W42
Server pin branch Pushed at 30646e217ab4f4b8760784bc4dcbda43466a3fdb https://github.com/CodexCoder21Organization/BuildTestServerService/tree/wip/consume-finished-shard-release-W42
Published embedded artifact buildtest.embedded:buildtest-embedded:0.0.69270330; JAR SHA-256 29974203a53e8782f7c100d007cb1775e4d42ccc29f4db261e62aa53e42f3570 https://kotlin.directory/buildtest/embedded/buildtest-embedded/0.0.69270330/

Local verification: two embedded regressions passed; two focused server tests passed; both Maven and server fat-JAR builds passed. Downloaded embedded JAR and POM matched the local build byte for byte. The local candidate is /tmp/claude-1000/-code/1f0f7d17-dc24-4d3a-a8ff-48753bfab06a/scratchpad/cx/lanes/W42/buildtest-server-0.0.69270330-candidate.jar, SHA-256 18441febc234bef8aadf2687e60c83f7adb9c276dd0426463fc1882a0b289012. All 1674 embedded classes in the candidate match the published JAR, and its resolver class matches the known transport fix. The candidate is local only.

Remaining: Re-verify current state; wait for required CI on both PRs, especially the server real-droplet sharding test; resolve any real failures at their cause; have the orchestrator review and merge or enqueue the PRs. Then ask the user whether to deploy the exact candidate after rechecking main and rebuilding if needed. This worker will not merge or deploy.

Prior handoff record

W26 2026-09-27 checkpoint: distinct ContainerNursery bridge defect reproduced, fix awaiting tests.

Main failed the real TCP/libp2p public-path scenario testUrlFacadeRpc_applicationErrorPreservesConcurrentBridgeCalls in https://buildtest.kotlin.build/run?id=fa035671 (0 passed, 1 failed, 613 filtered). An ordinary decoded backend application error unconditionally closes BackendPushBridge before error classification. Its held sibling then sees Stream closed and the caller receives OUTBOX_FULL instead of the application error. This is a proven source defect, not yet proof of the initiating close for incident 4d2785ee.

The minimal fix preserves the shared bridge for BackendRpcApplicationException, retaining cleanup for transport failures. Branch: https://github.com/CodexCoder21Organization/ContainerNursery/tree/wip/droplet-rpc-bridge-error-W26 at 0ab4e42aa93fdb02b5ab8f7efe057c20b5bcc7eb. No PR yet because fixed validation is incomplete. Baseline full error, invariants and two reviews are under investigations/W26. Coordinate 0.0.165 was 404 on kotlin.directory and remains unpublished.

Fixed remote submission lost buildtest transport before returning a run ID; local selected test is compiling under the shared JVM lock. Required remaining work: obtain fixed test result, run application-error and transient-stream-close neighbors, full remote suite, final rebase and main-targeted PR, required CI. No production deployment is authorized. Correlate the incident first close with bridge/request identifiers rather than assuming this source defect explains it.

Earlier hypotheses refuted: current coordinator matches published resolver 0.0.1259 classes; the provider PID did not change at 00:14:39. Evidence branch https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/droplet-rpc-close-W26 at 9c8638b94. Do not duplicate or modify https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1105 or concurrent ContainerNursery work.


W26 evidence update - 2026-09-27 00:43 UTC (RE-VERIFY)

The initial deployment premise is stale. The coordinator hotfix filename now contains PersistentRpcConnection.class e40537ba344adffce9a8bee8509f6f371acac88c416eeab0d511790e194572c5 and UrlProtocol2.class faab7edb3357ddf974e9bc4162a9807732cd9f9251ae1d8a75111c0b72aced7c, both exact matches to a freshly downloaded foundation.url:resolver:0.0.1259 JAR. File mtime 2026-09-26 22:42:12 precedes current PID 395887 start at 22:42:44. This rules out the old coordinator destructive-retirement code as an explanation of the new 00:14:39 incident.

The saved run log for https://buildtest.kotlin.build/run?id=4d2785ee identifies getDropletCreationStatus reader EOF, not a demonstrated first write refusal. The provider accepted all ten async creates; the three missing client outcomes correspond to droplets 603971810, 603971813 and 603971815. Provider PID 3641003 survived nursery replacements before and after the incident; its RESTART is recorded only at 00:24. Nursery CLI route logs omit the adopted-child interval, while per-JAR logs retain the calls. No evidence yet identifies the first connection owner that closed the stream. No production mutation was performed by W26.

Durable evidence: https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/droplet-rpc-close-W26 at first checkpoint 69a29395cc103e933a529d3f55c99e2e5aebd292 (subsequent corrections/experiment pending push). A real two-node replacement experiment is running; remote submission failed before allocation, so a single local targeted run uses the shared JVM lock. Do not duplicate existing transport fixes or claim this incident resolved from an old artifact snapshot.

This note records transport evidence only. W26 does not own or change the provisioning PR, deployment candidate, claim, or completion state of this related handoff.


W16 update: real extra droplet after completed tests

Written 2026-09-26 23:49 UTC by codex-W16. This section supersedes W9's recommendation to consider the current candidate for deployment. W9's package evidence and rollback notes below remain valid for that artifact, but the candidate includes the embedded defect found here.

The mechanism is a live dynamic shard provisioning retry after a surviving shard has already persisted every scheduled test result. CI run https://github.com/CodexCoder21Organization/BuildTestServerService/actions/runs/35975160789 created a third droplet with the same run-specific primary name after all four tests had passed; it was not an unrelated droplet counted by the test. A public-API in-process test failed on old code with three primary create attempts. A second test proved a quota-shaped failure could requeue the finished run. Both pass on the fix, as do three neighboring tests (5 passed, 0 failed). Separate test and code reviews are recorded in W16 findings.

Artifact State Link
BuildTestEmbedded fix PR open, checks pending, head 06c0d220b1ab5e1a08fc5e7036267f2c93f2208d https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1105
Server CI diagnosis Pushed evidence branch, head 3b7d56ab069f080f7c018771ebceb72dc322f4b7 https://github.com/CodexCoder21Organization/BuildTestServerService/tree/wip/real-droplet-e2e-diagnosis-W16

Remaining: get the upstream PR's required checks green and complete review; the orchestrator alone may merge. After merge, publish a new embedded artifact, update the server pin, build and test a new candidate, and return to the user's production deploy decision. Do not deploy W9's candidate, which pins embedded 0.0.69240600 and still has this retry behavior. The existing blocked reason for production authorization stays in place. No production change was made by W16.


Decide whether to deploy the verified main buildtest candidate

Written 2026-09-26T22:45:07.488944+00:00 by codex-W9.

RE-VERIFY: This is a write-time snapshot. Recheck the live route and running container with ContainerNursery CLI routes/containers, source heads through git ls-remote, and local/remote jar SHA-256 before acting. No jar was uploaded or deployed by W9. The candidate is local only by explicit user instruction.

Remaining mission and operator question

Deploy this candidate? The user narrowed W9's task to constructing and verifying a main deployment artifact, preserving the two production embedded hotfix behaviors, and stopping for the production decision. All requested package checks passed. Review the disclosed historical main integration-test failure before authorizing rollout. The broader first-write/root-close cause remains unproven and concurrent ContainerNursery work is unchanged.

W9 result

DECISION: BLOCKED-NEEDS-USER HANDOFF: hf-2026-09-24-stop-buildtest-s-persistent-rpc-transport-to-containernursery-from-being-retired-mid-request-droplet-provisioning-failures-since-08-40z-2026-09-24 (claim on exit: released) ASKED: Produce and inspect a local BuildTestServerService deployment candidate from current main, preserving both embedded hotfix behaviors and proving the packaged resolver is 0.0.1259. Record the live route, current jar, exact deploy/rollback commands, and stop before any upload or deploy. ACTUALLY TRUE: Snapshot 2026-09-26T22:43:57.502836+00:00. Production still runs hotfix jar buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar (PID 61088; url:buildtest: RUNNING on port 44233). Server main already pins the right resolver and embedded release; no dependency fix or publication is needed. Embedded main advanced through another actor's cleanup change, and open resolver/nursery work remains untouched. Both requested hotfix behaviors are present in the actual candidate. Main has mixed historical bld-build results; the failing run is disclosed below, so this is packaging verification, not an all-green regression claim.

DID:

  • Built locally under the required shared JVM lock from exact main commit https://github.com/CodexCoder21Organization/BuildTestServerService/commit/9d10b8829241c471f09afc57fbe239ad711519a5 . Source and dependency declarations unchanged. Build success: 1 passed/0 failed, 190094 ms, including the build rule's assembled method/constructor linkage gate.
  • Candidate path: /tmp/claude-1000/-code/1f0f7d17-dc24-4d3a-a8ff-48753bfab06a/scratchpad/cx/lanes/W9/buildtest-main-candidate.jar.
  • Candidate SHA-256: 4012e862cde83e9bf87ca048b8e03472b0a5ce9ae153cd22afb970e88b542aa9; size 154591475 bytes.
  • Packaged foundation/url/resolver/PersistentRpcConnection.class: e40537ba344adffce9a8bee8509f6f371acac88c416eeab0d511790e194572c5, EXACT expected match. Disassembly proves retireWriteFailedTransport closes only after getPendingRequests().isEmpty().
  • Compared every class from separately downloaded published jars against the candidate: resolver 0.0.1259 = 1244/1244 equal; embedded 0.0.69240600 = 1599/1599 equal; combined 2843 passed/0 missing/0 mismatched. ZIP integrity, unique entries, manifest, and main entry checks passed.
  • Embedded proof: ProjectionChangeJournal.class hash f7cd5f767f681a30c020c6e14f596052f239a61cc47d402ec984418f89d70ab7; BuildTestEmbeddedService.class a9be2cbffa3c50d90cc458cad57b9cd27663b22aecd2564098aa57022db670ec; BuildDriver.class 1124b52b098f162cfe85fecfd0a799ef9a069eff63bf5b5586991edffaf013a1. Actual bytecode includes retained-baseline/suffix bounds, background reconciliation start at getProjectionChanges offset 75, and the validators-only continue branch with ordinary different-byte files still rejected. ProjectionChangeJournal source at embedded publication commit 727d500bc19ea802431d62744ae531ab4acc1396 equals current Embedded main's file. This candidate follows server main's embedded pin; it does not bundle all newer Embedded main features.
  • Exact build command, after git -C ws/BuildTestServerService fetch origin and git -C ws/BuildTestServerService rebase origin/main: /tmp/claude-1000/-code/1f0f7d17-dc24-4d3a-a8ff-48753bfab06a/scratchpad/cx/jvm-slot.sh ws/BuildTestServerService/scripts/build.bash --local buildtest.server.buildFatJar /tmp/claude-1000/-code/1f0f7d17-dc24-4d3a-a8ff-48753bfab06a/scratchpad/cx/lanes/W9/buildtest-main-candidate.jar.
  • Exact offline audit command from lane: python3 evidence/verify_candidate.py buildtest-main-candidate.jar resolver-0.0.1259.jar embedded-0.0.69240600.jar. Unit tests run in W9: 0; no claim of a new unit-test pass. W7's already-documented 7-pass/0-fail resolver run was accepted as prior evidence, not rerun.
  • Nursery CLI 0.0.20 routes --json and containers --json independently confirm image /root/ContainerNursery-uploads/jars/buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar. Read-only SSH confirms PID command line, size 153679002, SHA-256 f9b80855d72b201e0f7b2e40c29ca2ea328d93edceb058c9da5f3e0a99cc5f48.
  • Exact upload/update-route/restart and rollback sequence is in /tmp/claude-1000/-code/1f0f7d17-dc24-4d3a-a8ff-48753bfab06a/scratchpad/cx/lanes/W9/evidence/deployment-commands.md, copied to the remote evidence branch. It stages /root/ContainerNursery-uploads/jars/buildtest-server-main-9d10b8829241-4012e862cde8.jar, verifies its digest, updates only the URL route image, then restarts url:buildtest:. Rollback verifies and restores the existing hotfix image and restarts the same route. Nursery upload-jar would overwrite/restart the current image, so separate staging is required. Fabric credentials are absent; the documented command uses SCP as the remaining separate-file transfer option. Only read access of the existing key is verified; no upload was attempted.
  • Test-comprehensiveness review and separate adversarial review are written in findings.md. They found no package mismatch; corrections were rollback-safe staging, full class audit, main CI qualification, and explicit limitations. No source/tests changed and no new PR is appropriate for this artifact-only task.
  • Investigated main's failed historical check via raw log download (gh run view logs were empty): https://github.com/CodexCoder21Organization/BuildTestServerService/actions/runs/35975160789 reports 195/196 successful, 1 failed: buildtest.server.e2eShardedBuildAcrossRealDroplets, expected 2 REAL droplets, observed [603259138, 603259139, 603259965]. Assertion stack is saved. Build step passed. Its cause has not been reproduced here; no production-droplet test was rerun. Earlier same-head check https://github.com/CodexCoder21Organization/BuildTestServerService/actions/runs/35973666198 succeeded.

Re-verified PRs (no PR created or changed):

  • https://github.com/CodexCoder21Organization/UrlResolver/pull/1088: MERGED, head 466c47d9e4c620bab07817cf36ff90ebb6a7d142, mergeStateStatus UNKNOWN; bld-build=SUCCESS; kotlin.build (kompile-remote-build)=SUCCESS; kotlin.build (remote)=SUCCESS.
  • https://github.com/CodexCoder21Organization/UrlResolver/pull/1122: OPEN, head d45e3ab95ecb285abef9718fcfb4cf04de74d1c5, mergeStateStatus BLOCKED; bld-build=IN_PROGRESS; kotlin.build (kompile-remote-build)=FAILURE; kotlin.build (remote)=FAILURE.
  • https://github.com/CodexCoder21Organization/BuildTestServerService/pull/369: MERGED, head 22ba5bf8981f74e2fc970b59f23a4171343388ae, mergeStateStatus UNKNOWN; bld-build=SUCCESS.
  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1084: MERGED, head e4683b5d5103c57d932556c6f535bc8be7fb780d, mergeStateStatus UNKNOWN; bld-test-shard (1 of 4)=SUCCESS; bld-test-shard (2 of 4)=SUCCESS; bld-test-shard (3 of 4)=SUCCESS; bld-test-shard (4 of 4)=SUCCESS; bld-build=SUCCESS.
  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1089: MERGED, head cbece367eb9ca795b36e7d4cf4bb26f1d81d035f, mergeStateStatus UNKNOWN; bld-test-shard (1 of 4)=SUCCESS; bld-test-shard (2 of 4)=SUCCESS; bld-test-shard (3 of 4)=SUCCESS; bld-test-shard (4 of 4)=SUCCESS; bld-build=SUCCESS.
  • https://github.com/CodexCoder21Organization/ContainerNursery/pull/625: OPEN, head a2fc6415da7c9fec508ff8337f68e59ec7eff4dd, mergeStateStatus BLOCKED; bld-build-release=SUCCESS; Build and test with Gradle=SUCCESS; kotlin.build (remote)=FAILURE; kotlin.build (kompile-remote-build)=IN_PROGRESS.

REMAINS / ORCHESTRATOR MUST DECIDE:

  • Question: deploy this candidate?
  • Options: (A) authorize an operator-controlled deployment of the exact jar above, accepting the disclosed historical integration qualification; (B) hold the artifact while the failed real-droplet check is investigated. Recommendation: review the historical test failure before approving the rollout; the requested package criteria all pass, but package identity is not a replacement for that integration result. Only the user can authorize production rollout. No worker merge/enqueue/branch protection or production mutation occurred.
  • Use evidence/deployment-commands.md for the exact sequence. Confirm exclusive coordinator process ownership during replacement and rollback. Recheck route/current jar before applying it; the jar is deliberately local only, per the explicit brief, and must be rebuilt if this lane is removed.
  • W7's first-write/root-close question remains unproven and the ContainerNursery work remains concurrent. This candidate addresses the known deployed resolver mismatch; it does not prove all causes of the original incident are resolved. Do not complete the handoff based only on candidate construction.

ARTIFACTS:

  • Evidence branch: https://github.com/CodexCoder21Organization/BuildTestServerService/tree/wip/deployment-candidate-evidence-W9 . Initial checkpoint 5ba658bbaba0145dc10ca86934a9761cc2508ba3; candidate-proof checkpoint e491249c884e3e083f391b4882607eb7d9123f07 (final remote SHA recorded below after final push). investigations/W9 contains the audit script/results, class hashes/disassembly excerpts, route/process evidence, command sequence, review findings, historical CI qualification, and report.
  • W7 evidence branch re-verified at 2226119aa6598ee31b37133874a725e5b7a6fbb8: https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/transport-deployed-evidence-W7 . Earlier invariant branch f9b5d4348c458099b599fc2fc5e8a6a53f1b6354: https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/sweep4-L1-transport-retirement-invariants .
  • Retained server hotfix branch: https://github.com/CodexCoder21Organization/BuildTestServerService/tree/hotfix/server-embedded-0.0.68290276 at 9f6bcecaef3a50ca1867950626ec77efee16583a. Embedded hotfix branch: https://github.com/CodexCoder21Organization/BuildTestEmbedded/tree/hotfix/embedded-0.0.68290276 at 739344ccd7caf546b51334787afff3f21f277856.
  • Published coordinates consumed, none published by W9: foundation.url:resolver:0.0.1259 (jar ec38e8931f58067724846cb9c02542e42fc428570c3f127c6fb9841d5dd01635), buildtest.embedded:buildtest-embedded:0.0.69240600 (jar 028e3ebacb10bb493c8710beb2feda3bd9fbb77899ab43dcffb6fe56d0eb61af).
  • Handoff RUNNING report acknowledged. Final BLOCKED report/body update and W9 claim release delivery results are recorded below. Existing W7 claim belongs to the earlier worker and was not released by W9.
  • No subagents. No candidate jar or downloaded dependency jar was uploaded anywhere. Only text evidence was pushed. No new runtime code, PR, publication, deploy, restart, route update, or remote data deletion.

Delegated work at a glance: none. Landing parallelization sweep: no PR gate exists in this narrow candidate task; build and read-only verification completed independently. Next step: the operator reviews the package evidence and disclosed CI qualification, then obtains the user's rollout decision. Lessons: compare actual packaged bytes; preserve the prior image before selecting deployment commands; distinguish artifact verification from historical or fresh integration-test results.

FINAL REMOTE EVIDENCE HEAD: 745f84aac62097844b33bf09aea744ff511d69f0, verified by git ls-remote; report snapshot and all candidate evidence at https://github.com/CodexCoder21Organization/BuildTestServerService/tree/745f84aac62097844b33bf09aea744ff511d69f0/investigations/W9 .

Relevant refs and durable artifacts

Repo Branch/ref Remote head SHA PR Contents/state
BuildTestServerService https://github.com/CodexCoder21Organization/BuildTestServerService/tree/wip/deployment-candidate-evidence-W9 745f84aac62097844b33bf09aea744ff511d69f0 no new PR Candidate verification, deployment/rollback commands, review/report; no runtime source changes
BuildTestServerService https://github.com/CodexCoder21Organization/BuildTestServerService/tree/main 9d10b8829241c471f09afc57fbe239ad711519a5 https://github.com/CodexCoder21Organization/BuildTestServerService/pull/369 Merged resolver pin; exact candidate source
BuildTestServerService https://github.com/CodexCoder21Organization/BuildTestServerService/tree/hotfix/server-embedded-0.0.68290276 9f6bcecaef3a50ca1867950626ec77efee16583a no new PR Current production hotfix jar; old resolver
BuildTestEmbedded https://github.com/CodexCoder21Organization/BuildTestEmbedded/tree/hotfix/embedded-0.0.68290276 739344ccd7caf546b51334787afff3f21f277856 no new PR Both hotfix behaviors also on main and packaged in candidate
UrlResolver https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/transport-deployed-evidence-W7 2226119aa6598ee31b37133874a725e5b7a6fbb8 no PR Prior deployment evidence and 7 passing transport regressions
UrlResolver https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/sweep4-L1-transport-retirement-invariants f9b5d4348c458099b599fc2fc5e8a6a53f1b6354 no PR Earlier invariant investigation, retained

Exact deploy and rollback instructions (not executed)

Candidate deployment and rollback commands - NOT EXECUTED

Run only after the user approves this exact candidate. This worker uploaded nothing and made no production change.

Candidate: /tmp/claude-1000/-code/1f0f7d17-dc24-4d3a-a8ff-48753bfab06a/scratchpad/cx/lanes/W9/buildtest-main-candidate.jar. SHA-256: 4012e862cde83e9bf87ca048b8e03472b0a5ce9ae153cd22afb970e88b542aa9; size: 154591475 bytes. Server main: https://github.com/CodexCoder21Organization/BuildTestServerService/commit/9d10b8829241c471f09afc57fbe239ad711519a5 .

The existing url:buildtest: route and RUNNING container both name /root/ContainerNursery-uploads/jars/buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar. PID 61088's command line confirms it. Current jar SHA-256: f9b80855d72b201e0f7b2e40c29ca2ea328d93edceb058c9da5f3e0a99cc5f48; size 153679002 bytes.

Re-check those facts immediately before deployment. If the current image or hash changed, stop and revise the rollback target. Do not overwrite the hotfix file.

The preferred Fabric upload cannot be authenticated from this lane: the documented ~/.config/hardware-control-fabric/ directory is absent and no matching Fabric credentials were found. Nursery CLI 0.0.20's upload-jar requires a route, writes its configured image path, and restarts it; it does not provide a separate staging-only upload. For this exact environment, the separate-file SCP sequence below is the remaining upload mechanism. The existing key was verified for read-only SSH; upload permission is not asserted because no upload was attempted. An upload failure stops this sequence before route mutation. An operator with Fabric credentials can instead perform the equivalent write-file upload to the same candidate path.

Upload, verify, update route, restart

export PATH=$HOME/.local/bin:$PATH
set -euo pipefail
W9_JAR=/tmp/claude-1000/-code/1f0f7d17-dc24-4d3a-a8ff-48753bfab06a/scratchpad/cx/lanes/W9/buildtest-main-candidate.jar
W9_REMOTE=/root/ContainerNursery-uploads/jars/buildtest-server-main-9d10b8829241-4012e862cde8.jar
W9_SHA=4012e862cde83e9bf87ca048b8e03472b0a5ce9ae153cd22afb970e88b542aa9
W9_CN=(coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com)

printf '%s  %s\n' "$W9_SHA" "$W9_JAR" | sha256sum --check -
"${W9_CN[@]}" routes --json
"${W9_CN[@]}" containers --json
ssh -i ~/.ssh/cn_host_diag -p 23 -o BatchMode=yes root@198.199.106.165 \
  'sha256sum /root/ContainerNursery-uploads/jars/buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar'
# Confirm the current route/hash still match the recorded rollback target before proceeding.

scp -i ~/.ssh/cn_host_diag -P 23 -o BatchMode=yes \
  "$W9_JAR" "root@198.199.106.165:$W9_REMOTE"
ssh -i ~/.ssh/cn_host_diag -p 23 -o BatchMode=yes root@198.199.106.165 \
  'printf "%s  %s\n" 4012e862cde83e9bf87ca048b8e03472b0a5ce9ae153cd22afb970e88b542aa9 /root/ContainerNursery-uploads/jars/buildtest-server-main-9d10b8829241-4012e862cde8.jar | sha256sum --check -'
"${W9_CN[@]}" update-route --domain buildtest --facade url --image "$W9_REMOTE"
"${W9_CN[@]}" restart --route-key 'url:buildtest:'
"${W9_CN[@]}" routes --json
"${W9_CN[@]}" containers --json

The update-route call specifies only the image; existing transport, environment, memory/metaspace, startup timeout and keep-warm settings are preserved. An image change can stop the previous instance; restart is a real production lifecycle action and uses its existing default, with no timeout increase. Confirm the old coordinator and its writer processes have exited before allowing a replacement to share the data directory (the Embedded README's deployment requirement). Verify the new process command line, RUNNING container image, packaged class hash, and a successful real build before declaring rollout complete. The candidate does not establish what initiated the first production write failure.

Roll back to the existing hotfix jar

export PATH=$HOME/.local/bin:$PATH
set -euo pipefail
W9_CN=(coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com)
ssh -i ~/.ssh/cn_host_diag -p 23 -o BatchMode=yes root@198.199.106.165 \
  'printf "%s  %s\n" f9b80855d72b201e0f7b2e40c29ca2ea328d93edceb058c9da5f3e0a99cc5f48 /root/ContainerNursery-uploads/jars/buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar | sha256sum --check -'
"${W9_CN[@]}" update-route --domain buildtest --facade url \
  --image /root/ContainerNursery-uploads/jars/buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar
"${W9_CN[@]}" restart --route-key 'url:buildtest:'
"${W9_CN[@]}" routes --json
"${W9_CN[@]}" containers --json

Rollback restores the old resolver 0.0.1171 along with the old service. It does not undo persisted work; both jar generations require orderly exclusive coordinator ownership. No old jar is deleted.

Next steps

  1. Re-verify source, candidate hash, and current production route/image. Keep the candidate local until approved.
  2. Operator reviews historical main CI qualification and asks the user: deploy this candidate?
  3. If approved, use the exact separate-file upload/image update/restart sequence above; verify coordinator process replacement and a successful build, and monitor the original sibling-failure symptoms. If declined, retain the evidence and jar; do not claim the handoff complete.
  4. Preserve W7's first-write/root-close question and concurrent nursery work as separate context; a package verification result alone does not establish complete incident resolution.

Prior handoff snapshot (historical context; W9 result above governs this pickup)

Verify the first write failure and reconcile the deployed old resolver with the merged drain fix

Written 2026-09-26T22:13:14.439537+00:00 by W7. RE-VERIFY: This is a snapshot. Query current PR heads/checks, git ls-remote, published artifact bytes, deployed process command lines and class hashes, and the actual run records before acting. Disk JAR identity plus matching live error text supports the deployed version finding; no JVM attach or full reproducible-build attestation was performed.

Remaining mission

The first failed-write/root-close cause remains unproven. Validate a consumer artifact carrying the already-merged sibling-drain fix while preserving the deployed hotfix behaviors, coordinate ContainerNursery's active protocol work, and leave production rollout to an explicit operator decision. Do not implement the old retirement fix again or create a redundant BuildTestServerService pin bump.

W7 terminal checkpoint and evidence

DECISION: CHECKPOINT HANDOFF: hf-2026-09-24-stop-buildtest-s-persistent-rpc-transport-to-containernursery-from-being-retired-mid-request-droplet-provisioning-failures-since-08-40z-2026-09-24 (claim on exit: held) ASKED: Establish why buildtest loses concurrent droplet-service replies, verify the deployed transport class and first failed write, and carry any real remaining fix to a reviewable state. The latest handoff already acknowledged that the sibling-drain fix was merged; its deployed use remained unverified. ACTUALLY TRUE: Verified 2026-09-26T22:13:14.439537+00:00. The same incident buildtest PID 61088 still runs buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar with the exact PersistentRpcConnection.class from published resolver 0.0.1171. ContainerNursery PID 1268157 has the exact class from 0.0.1203. Both methods unconditionally retire a transport after a write failure. Current retained logs still show one getDroplet write failure exceptionally completing sibling listDroplets and createDropletAsync calls. Published 0.0.1259 has the drain guard, and the server's bump to it is already merged. The first failed-write cause is still unproven; no new main-branch transport defect was reproduced.

DID:

  • Pushed evidence at every milestone to https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/transport-deployed-evidence-W7 (final head e1e339eb95ee091565e458ad546610b9478e48f0). No runtime source or tests were changed; no new PR was appropriate for an unreproduced defect.
  • Verified https://github.com/CodexCoder21Organization/UrlResolver/pull/1088 MERGED, head 466c47d9e4c620bab07817cf36ff90ebb6a7d142, merged 2026-09-14T07:54:09Z; bld-build, kotlin.build (remote), kotlin.build (kompile-remote-build) SUCCESS. Verified current-main guard by content, not only PR state.
  • Verified https://github.com/CodexCoder21Organization/BuildTestServerService/pull/369 MERGED, head 22ba5bf8981f74e2fc970b59f23a4171343388ae, merged 2026-09-24T04:06:22Z, bld-build SUCCESS. Main pins resolver 0.0.1259 / protocol 0.0.527 / libp2p snapshot-27. The preserved production hotfix branch was created at 04:06:21Z, one second before that merge, and explicitly pins resolver 0.0.1171; it was deployed about 27 minutes later.
  • Verified both embedded hotfix behaviors already landed: https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1084, head e4683b5d5103c57d932556c6f535bc8be7fb780d, MERGED with bld-build and all four bld-test-shard checks SUCCESS; https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1089, head cbece367eb9ca795b36e7d4cf4bb26f1d81d035f, MERGED with the same five checks SUCCESS. Hotfix/main ProjectionChangeJournal.kt and the hotfix regression files match; main retains the validator merge exception.
  • Fresh targeted/neighbour test run https://buildtest.kotlin.build/run?id=1849c65d: COMPLETED, 7 PASSED, 0 FAILED; 1838 other discovered tests FILTERED_OUT. Exact command after git fetch origin and git rebase origin/main: scripts/test.bash --remote --test testLargeRpcRequestSurvivesOrdinaryTrafficOnItsOwnConnection --test testRefusedWriteLetsInFlightSiblingDrainThenClosesTransportOnce --test testAmbiguousWriteFailureLeavesSiblingRequestDrainingOnItsTransport --test testAmbiguousWriteFailureOnDeadPeerStillFailsPendingRequestPromptly --test testPersistentRpcConnectionSerializesRetirementWithWriteAndOldReaderCleanup --test testRetiredWriteFailedTransportClosesWhenInterruptedCallerDrainsIt --test testPartialFrameOnLiveStreamLeavesSiblingToItsOwnDeadline --log /tmp/claude-1000/-code/1f0f7d17-dc24-4d3a-a8ff-48753bfab06a/scratchpad/cx/lanes/W7/transport-tests.xml. Source/test base ffaac8c79bf1e1081a2ed76a2a29348cec55bcfd; submitted evidence-branch head 24e0e35edec2c899d25b65d7f46536ee8f8dfe8a had identical runtime source/tests. Command exited 0. The real client/server test took 76.603 seconds. Server-side run.json and selected test-results.json records were captured because --log did not produce a local XML.
  • Test-comprehensiveness pass: the seven existing tests cover live client/server traffic, refused/ambiguous writes, EOF, retirement/write ordering, final caller interruption, and a partial live frame. They do not prove the initiating production root-close cause, every I5 concurrency case, or packaged consumer compatibility. No test was weakened.
  • Separate adversarial pass: verified byte equality, bytecode behavior, source/deploy divergence, and current recurrence. Corrected remaining work to avoid a redundant server pin PR or a blind three-repository upgrade. No new code finding was guessed into a fix.
  • Concurrent actors left untouched: https://github.com/CodexCoder21Organization/UrlResolver/pull/1122 OPEN at d45e3ab95ecb285abef9718fcfb4cf04de74d1c5, BLOCKED, remote FAILURE, bld-build and kompile-remote-build IN_PROGRESS; https://github.com/CodexCoder21Organization/ContainerNursery/pull/625 OPEN at a2fc6415da7c9fec508ff8337f68e59ec7eff4dd, BLOCKED, remote FAILURE, bld-build-release/Gradle SUCCESS, kompile-remote-build IN_PROGRESS. That ContainerNursery branch advances protocol to 0.0.546 while retaining resolver 0.0.1203 / snapshot-26.

REMAINS / ORCHESTRATOR MUST DECIDE:

  1. Keep this handoff open. Recover/capture the initiating writer-side exception and first root close with a common request/transport identity. The four historical run paths are absent; retained rotating logs show generated sibling failures, which omit the initiating exception. Do not call EOF or an admission-terminal message a root cause.
  2. Verify the actual main BuildTestServerService fat JAR and its published embedded dependency preserve the live hotfix behaviors and contain the intended resolver class. Main already pins 0.0.1259; add no redundant bump if its resolved packaged class is correct. The old deployed hotfix's explicit pin explains the observed class without proving a current main dependency-selection defect.
  3. If Embedded standalone/packaged resolution still requires alignment, update its complete non-transitive closure together, including protocol/libp2p and its older explicit SJVM declarations; run consumer integration and binary-linkage checks, publish an unused patch version only if needed, and update the server's embedded pin accordingly.
  4. Coordinate ContainerNursery's resolver upgrade with the active protocol 0.0.546 work above; do not downgrade or overwrite that actor's work. Validate the combined closure and packaged class, then obtain green required checks and both review passes for any new PR.
  5. Only after a concrete validated artifact is available should the user be asked to approve production rollout. This run neither built a deployment candidate nor claims deployment readiness. No merge/enqueue/publication/deploy/restart or production mutation occurred.

ARTIFACTS:

  • W7 evidence: https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/transport-deployed-evidence-W7 at e1e339eb95ee091565e458ad546610b9478e48f0; investigations/W7 contains incremental findings, deployed and published hashes, retirement disassemblies, sanitized live error stacks, consumer pins, concurrent PR state, seven selected test results, and a tooling challenge draft.
  • Prior evidence: https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/sweep4-L1-transport-retirement-invariants at f9b5d4348c458099b599fc2fc5e8a6a53f1b6354, unchanged.
  • Existing server hotfix: https://github.com/CodexCoder21Organization/BuildTestServerService/tree/hotfix/server-embedded-0.0.68290276 at 9f6bcecaef3a50ca1867950626ec77efee16583a, unchanged; one source commit is outside main.
  • Existing embedded hotfix: https://github.com/CodexCoder21Organization/BuildTestEmbedded/tree/hotfix/embedded-0.0.68290276 at 739344ccd7caf546b51334787afff3f21f277856, unchanged; equivalent behavior is now on main.
  • Published coordinates verified by fresh downloads: foundation.url:resolver:0.0.1171, 0.0.1203, 0.0.1259. 0.0.1216 returns 404. None published by W7. Class hashes respectively 86321c6db67c07d1b424507ba637d8abd2dc733e08ca9cf1dc7693dcc026e995, 4ce3ce0e3d9ad39f532a485781ee916cae3da397526cddc6133ace075ebca979, e40537ba344adffce9a8bee8509f6f371acac88c416eeab0d511790e194572c5.
  • The first RUNNING report was acknowledged. A later milestone report returned an ambiguous submitStatusReport/EOF failure; the final checkpoint report supersedes it and uses RUNNING to keep the requested claim held. Handoff title/body updated with remaining work and branch table. The stale latest-ten blocker was cleared because that old lane policy does not govern this brief.
  • No subagents. No running local test/watch processes remain. No new PR awaits merge. The challenge CLI was not invoked because it auto-merges/enqueues/deletes branches, forbidden in this run; its draft is durable for the orchestrator.

Branch and unmerged-state table

Repository Branch Remote head SHA PR State
UrlResolver https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/transport-deployed-evidence-W7 e1e339eb95ee091565e458ad546610b9478e48f0 none Evidence only; runtime source/tests unchanged; 7/7 selected tests passed
UrlResolver https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/sweep4-L1-transport-retirement-invariants f9b5d4348c458099b599fc2fc5e8a6a53f1b6354 none Prior invariant/evidence checkpoint; unchanged
BuildTestServerService https://github.com/CodexCoder21Organization/BuildTestServerService/tree/hotfix/server-embedded-0.0.68290276 9f6bcecaef3a50ca1867950626ec77efee16583a none located Existing live hotfix source; old resolver pin; unchanged by W7
BuildTestEmbedded https://github.com/CodexCoder21Organization/BuildTestEmbedded/tree/hotfix/embedded-0.0.68290276 739344ccd7caf546b51334787afff3f21f277856 equivalent behavior merged through the linked PRs Existing hotfix, unchanged by W7; behavior now on main

Deployed artifacts

Buildtest PID 61088, started 2026-09-24 04:33:56 UTC, executes /root/ContainerNursery-uploads/jars/buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar (153679002 bytes). Its sole PersistentRpcConnection.class matches resolver 0.0.1171. ContainerNursery PID 1268157, started 2026-09-22 15:32:13 UTC, executes /root/ContainerNursery/bin/container-nursery-a5a938b1-20260922T150441Z.jar (131932867 bytes); its sole corresponding class matches resolver 0.0.1203. Both file mtimes precede process starts. These old running artifacts are NOT replaced merely because main pins are newer. No artifact was deployed/published by W7.

Historical body retained for provenance; superseded observations are not current instructions

The following old lane's latest-ten gate and old version uncertainty are historical. The current worker resolved deployed class identity and successfully ran seven regressions. The first write trigger remains unknown.

Verify buildtest's deployed resolver class and first write failure before changing transport retirement

Written 2026-09-26T21:41:16.152852+00:00. RE-VERIFY: This is a write-time snapshot. Re-check published jar/class hashes, current consumer source pins, the incident deployed artifact, and the latest ten buildtest runs. Source main and published artifacts are not proof of deployed bytes.

Remaining work first

Identify the exact PersistentRpcConnection.class in the incident buildtest hotfix JAR and the cause of the first failed write. The source-level destructive-retirement fix already merged before this incident: https://github.com/CodexCoder21Organization/UrlResolver/pull/1088 (MERGED 2026-09-14T07:54:09Z). Do not write that fix again. Resume test work only after the policy's latest-ten gate allows a run.

Mission and current result

The sweep4 assignment requested a design-first, fail-first fix for the September 24 droplet-allocation outage, with I1-I5 transport invariants, no downstream workaround, and no publishing or production action in this lane. Lane outcome is REQUEUED-INFRA. Current-main destructive retirement is refuted; the historical runtime root cause remains unproven.

OBSERVED: Published resolver 0.0.1171's retireWriteFailedTransport unconditionally calls retireTransport, constructing the exact error quoted by the incident report. Published resolver 0.0.1259 and current main instead disable new checkout and close only when the generation's pending map is empty. Disassembled methods and SHA-256 hashes are in the checkpoint branch. The historical fix commit declared 0.0.1216; its Maven jar returned HTTP 404, so that coordinate is not a publication to rely on. INFERRED: The incident hotfix may contain the old class. BuildTestEmbedded source still pins 0.0.1171, but BuildTestServerService source pins 0.0.1259. Class selection in the actual incident fat JAR is the next decisive observation. This lane did not inspect any deployed JAR and did not establish why the first write failed.

OBSERVED: The latest-ten pre-run gate contained two infrastructure failures, 5b393be1 (no provisioning progress, EOF, 0/1 tests) and d4573ff3 (provisioning deadline, 0/31 tests). Policy permits at most one. Newer health checks succeeded and some current runs progressed; the gate result does not mean all current runs fail. No build or test was submitted. No new tests, production source changes, fix PR, publication, bump, merge, queue action, or production change occurred.

Durable work

Repository Branch Remote head PR State
UrlResolver https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/sweep4-L1-transport-retirement-invariants f9b5d4348c458099b599fc2fc5e8a6a53f1b6354 None Evidence-only checkpoint; no code or tests changed; no builds or tests run

Evidence directory: https://github.com/CodexCoder21Organization/UrlResolver/tree/wip/sweep4-L1-transport-retirement-invariants/investigations/sweep4-L1 . It contains findings, invariant/ownership table, published-class hashes, both retirement bytecode listings, buildtest gate evidence, blocked resume instructions, conditional consumer chain, and a challenge draft. No JARs/classes, build output, raw route output, or credentials were pushed. UrlProtocol fresh checkout was untouched; there are no unpublished source changes. Nothing was deployed or published by this lane. The pre-existing production artifact's source lineage remains unverified and is the first remaining item.

Test status

OBSERVED: Existing real client/server regression test testLargeRpcRequestSurvivesOrdinaryTrafficOnItsOwnConnection is on main. The merged PR reports 5/5 failures before, 10/10 passes after and a 26/26 targeted gate. Those are historical author-recorded results, not this lane's measurements. New invariant failure rates: N/A, zero runs. Existing deterministic sibling-drain, actual EOF, exact-once close, retirement/write ordering and partial-live-stream tests are mapped in invariants.md. Do not make an already-correct main fail artificially merely to satisfy fail-first wording.

Next steps

  1. Re-check ownership and the latest-ten run gate. Read the checkpoint's invariant table and source/artifact evidence.
  2. Through read-only Fabric commands, identify the incident hotfix file, process generation, and PersistentRpcConnection.class hash/disassembly; correlate its complete first write-failure cause chain and first-close event. Do not infer deployed code from source pins.
  3. If old destructive retirement is actually deployed, validate the existing fixed published resolver with the real-client/server regression and consumer integration tests; coordinate the supervisor's publish/bump chain below. No new resolver fix is needed for that already-fixed behavior.
  4. If deployed bytes already have the drain guard, reject destructive write retirement as the initiating explanation and reproduce the actual root/stream owner failure before changing code. Admission's error text says terminal OR physically closed; it does not prove admission initiated the close.
  5. Any new defect must still get public-API real-server fail-first invariant tests, a checkpoint, targeted and neighbour validation, independent test/code reviews and green required CI. Supervisor alone merges/enqueues; publishing and production rollout remain separate authorized work.

Publish and consumer chain (not executed)

OBSERVED at 2026-09-26: current UrlResolver source declares resolver 0.0.1259 and protocol 0.0.527. Published resolver 0.0.1259 has the drain guard; 0.0.1171 has destructive retirement. The historical fix commit declared 0.0.1216, whose artifact returned 404. Never assume merge means publication.

OBSERVED source pins:

  • BuildTestEmbedded main 1948b777514e488e7b1620ac98a0c28fcd512064: resolver 0.0.1171 with resolveTransitiveDependencies=false; protocol 0.0.503; libp2p snapshot-26. Own coordinate 0.0.69241730.
  • BuildTestServerService main 9d10b8829241c471f09afc57fbe239ad711519a5: buildtest-embedded 0.0.69240600; resolver 0.0.1259; protocol 0.0.527; accompanying libp2p pin must remain compatible.
  • ContainerNursery main 40f3f7501a153b13ac1a763464e4a0f72515253c: resolver 0.0.1203; protocol 0.0.528. No deployed class inspection was performed.

INFERRED / conditional next chain:

  1. Inspect incident and current deployed fat JARs for PersistentRpcConnection.class; compare its SHA-256/disassembly with published 0.0.1171 and 0.0.1259. Record exact artifact, process generation and class hash. This decides whether a consumer update is the right remaining work.
  2. If only the known destructive-retirement bug is present, no new resolver publication is required: 0.0.1259 is an existing candidate, subject to full compatibility/regression validation. If a new upstream defect is reproduced, merge that fix under supervisor review, allocate an unused resolver version R, publish R jar+POM with SHA verification; publish protocol P first only if changed.
  3. Update BuildTestEmbedded's non-transitive resolver edge 0.0.1171 -> verified R (possibly existing 0.0.1259), explicit protocol 0.0.503 -> paired P, and explicit runtime closure together. Validate actual merged class bytes and invariant/consumer tests; publish a new embedded version E after review.
  4. Update BuildTestServerService's embedded edge 0.0.69240600 -> E, align resolver/protocol/libp2p, and verify resulting fat JAR contains the intended class. Current direct 0.0.1259 pin alone does not establish what the historical hotfix bundled.
  5. Update ContainerNursery's resolver 0.0.1203 -> verified R; assess its existing protocol 0.0.528 compatibility instead of blindly downgrading to 0.0.527. Validate the final fat JAR's classes and its real transport integration tests.
  6. Supervisor/operator decides publication, bump PRs, and any rollout. This lane performs none of those and never merges or enqueues.

Historical incident record (2026-09-24; not re-verified live)

Remaining work first: (1) capture which process sends the first FIN or RST on each buildtest to ContainerNursery :35000 root, and the exception behind 'Persistent RPC transport ... was retired after the write ... failed' (read-only commands listed below); (2) write fail-first UrlProtocol/UrlResolver tests showing a root carrying host-initiated persistent RPC is never retired while requests are pending, and that one write failure does not retire the shared transport underneath other in-flight RPCs; (3) fix, publish, bump buildtest and ContainerNursery. Restart or redeploy is the operator's call; nothing was restarted, killed or deployed by this investigation.

inv-droplets FINAL - why buildtest cannot reach url://digitalocean-droplets/ (2026-09-24, read-only, 10:59-11:30Z)

Bottom line

  • The droplet service (DigitalOceanDropletServiceServer) is NOT down. One JVM (PID 3641003, jar droplet-service-server-0.0.82-r-derivation-fix-20260827T2205Z.jar), up since 02:41:11Z. It answers every request it gets: listDroplets median 0.6 s, max 1.9 s over 337 calls, 10:27-11:03Z. Ruled out: process dead, OOM kill, hung threads, listener not bound, ContainerNursery classifier kill, zombie generations (only one instance; ContainerNursery has been up 1d19h).
  • What actually breaks is the libp2p root connection between buildtest and ContainerNursery's libp2p host at 198.199.106.165:35000 (peer 12D3KooWLMyX...). ContainerNursery fronts every url:// route there and forwards over local TCP to the container on port 43045. These roots are closed while requests are still in flight. ContainerNursery finishes the response, then finds the root inactive with 0 response bytes written. buildtest sees EOF ("read 0 of 4 bytes"). That is the same "server wrote, client read 0 bytes" contradiction seen in earlier incidents.
  • This is not specific to the droplet service. The same response loss hits 'buildtest/*' responses, plus githubproxy, aiclisupervisormanager and kompile-remote-build. Droplet provisioning is where it shows up first because buildtest will not replay a droplet RPC whose outcome is unknown (AmbiguousRpcRequestException). It fails the run instead.
  • Not recovering. Droplet-related run failures per half hour: 08:30 1/45, 09:00 2/28, 09:30 7/50, 10:00 8/51, 10:30 17/32. Zero from 04:00 to 08:30 (about 300 runs). The newest failure is d285e7dd at 10:56:47Z, and the matching ContainerNursery log counters are still rising.

OBSERVED evidence

  1. buildtest /api/runs (500 runs): the earliest droplet-transport failure is 29f55784 at 08:41Z. Error shapes:

    • "RPC request 'listDroplets'|'getDroplet'|'getDropletCreationStatus'|'deleteDroplet' to service 'digitalocean-droplets' has an ambiguous outcome ... Persistent RPC connection ... was closed while requests were still pending. Reader transport failure: java.io.EOFException: Stream closed while reading message data (read 0 of 4 bytes)" (caa685b0, 7d1bae48, e427b7db, 67372356, ...)
    • The bytecode fetch fails the same way. d285e7dd: direct to /ip4/198.199.106.165/tcp/35000/p2p/12D3KooWLMyX... -> "io.libp2p.core.ConnectionClosedException: Channel closed [id: 0x260da919/5 ...]". Relay paths: "got 1/21 chunks", "Timeout waiting for relay response".
    • Also seen: e42aaa92 "DropletServiceCallStalledException: listDroplets stalled for 30000ms"; a0877d00 "Cannot invoke method on closed instance proxy: dropletserviceserver/DropletImpl.getName()"; 4d1abd31 / ea996e22 / 09c5c695 hit "Provisioning deadline exceeded" (600000 ms) with 18-20 droplets created per run.
  2. ContainerNursery containers: url:digitalocean-droplets: RUNNING, port 43045. health-detailed: healthy, uptime 1d19h27m.

  3. Fabric processes plus /proc (read-only):

    • droplet service 3641003: 134 threads, 97 fds, RSS 494 MB
    • buildtest 61088 (started 04:33, -Xmx4096m): 388 threads, 175 fds, RSS 2.1 GB, peak 4.1 GB
    • ContainerNursery 1268157: 470 threads, 573 fds, 153% CPU
    • jstat: garbage collection is not pathological in any of the three right now.
  4. Correlated trace, 10:42Z, using both container logs (both are stamped by ContainerNursery's clock):

    • The droplet service logged receiving getDropletCreationStatus 87f78c65 at 10:42:24.194. buildtest never got the reply: that request failed at 10:42:26.017 when the connection closed.
    • listDroplets sent at 10:42:25.916 failed 148 ms later with the same close. The service never logged receiving it.
    • deleteDroplet 603283494 was sent at 10:42:26.830 and logged by the service 9.4 s later, at 10:42:36.189.
    • deleteDroplet 603283434 attempt 1 (10:42:23.878) was never logged by the service.
  5. ContainerNursery process log (/root/ContainerNursery/data/logs/container-nursery.log and .1-.4, read with grep through fabric) shows lost replies on the server side: "Completed RPC response 'req-37993' for 'digitalocean-droplets/listDroplets' on service 'digitalocean-droplets' was not delivered on child stream '02422efffe3dedee-001359bd-0010153a-083f7fdee7e2f904-9fc055e8/7' of root generation 15945; close initiator=TRANSPORT; cause=java.io.EOFException: Root connection became inactive before any response-frame byte completed; response-frame bytes completed=0 [unhandled RpcResponseDeliveryLossEffect"

    • RpcResponseDeliveryLossEffect per file, oldest to newest: 7, 6, 7, 19, 53.
    • First 60 matches in the current file: 22 for buildtest/, 4 for digitalocean-droplets/. 23 of them are close initiator=TRANSPORT, 2 are PROTOCOL.
    • RpcResponseCloseBarrierFailureEffect per file: 226, 124, 146, 159, 715. The message is "PeerExchange response close barrier ... did not complete within 5000 milliseconds ... UnknownStreamIdMuxerException".
    • The "root generation" number in these lines climbs: about 10.8k at 08:2x, 13.3-13.7k at 09:0x-09:3x, 14.0-14.3k, 14.6-15.7k at 10:1x-10:4x, and 15.9-16.0k at 10:4x-11:1x.
  6. buildtest 10:44:30Z: "Cannot register a host-initiated substream on root connection@453ed133: expected a connection accepting new work, but inbound admission has made it terminal or the physical connection is closed." That wording comes from UrlProtocol InboundSubstreamAdmissionControl. The closed streams in buildtest's errors have small stream ids (3..30), so each root dies after only a few substreams.

  7. tcpdump of FIN/RST on tcp port 35000 (read-only, 11:19:31-45Z): 31 connections closed in 14 s, about 2.2 per second. Who sent the first FIN/RST: local client 14, remote client 11, ContainerNursery 6.

  8. ss -tnp dport :35000 (6 snapshots, 110 established rows): buildtest PID 61088 held no established connection to :35000 in any snapshot, even though every droplet RPC it makes targets 12D3KooWLMyX@:35000. buildtest has no stable root to ContainerNursery. Each burst of calls rides a newly dialled root that is torn down within seconds.

  9. Relay disconnect churn is everywhere: 1237 RelayDisconnectedPeerCleanupEffect lines in 22 min of buildtest log, 1045 of them from one peer, 12D3KooWK7DoYijp1LVpsGkH3k8gZHvPjTfPfhCZqwQS5Xi2edje (about 47 per minute, not identified). The droplet service log has about 276 in 3000 lines.

  10. ContainerNursery 30 s "RPC forward failed: TIMEOUT" lines were already there before the onset (23 in 08:4x), so they are background and do not mark the onset.

  11. buildtest's own stderr, passed through into the ContainerNursery process log ("retire" grep, 11:27Z): "Persistent RPC transport to service 'digitalocean-droplets' was retired after the write for request 'listDroplets' failed. Pending requests on that transport were not replayed because their handlers may already have run." The same text appears for getDropletCreationStatus. Also "[DropletClient] closing retired sandboxed connection after renewDroplet(603283888)". So buildtest's resolver closes ("retires") its whole shared persistent transport after a single write failure. Every other RPC in flight on it fails as ambiguous, and ContainerNursery then sees its completed replies land on a dead root (close initiator=TRANSPORT). This matches the tcpdump finding that local clients send the first FIN on most closes. The "refusing new substreams" admission message: 0 hits in the ContainerNursery logs .0-.2.

INFERRED mechanism (medium confidence)

Starting around 08:40Z and getting steadily worse, the libp2p roots from buildtest and other local clients to ContainerNursery's host on :35000 are closed while RPCs are still in flight. The evidence (tcpdump first-closer counts; ContainerNursery's close initiator=TRANSPORT; buildtest's own 'Persistent RPC transport ... was retired after the write ... failed' lines; the admission-terminal message inside buildtest) points to the closes being initiated mostly on the client side. When one write fails, buildtest's UrlResolver retires the whole shared persistent transport, so one failed write fans out into every concurrent droplet RPC failing as ambiguous. The first write failure itself is not established. The leading candidate is buildtest's own UrlProtocol host making the shared root to 12D3KooWLMyX terminal (inbound-substream admission or connection retirement: see the 10:44:30 'inbound admission has made it terminal' line), which then fails the next write on the persistent RPC stream carried by that root. ContainerNursery's own retirement (PROTOCOL) is a minority of cases.

The droplet service is the victim, not the cause. It shows up first because buildtest will not replay a droplet RPC whose outcome is unknown. The durable fix belongs in UrlProtocol/UrlResolver connection ownership: a root that carries host-initiated outbound RPC work must not be retired or closed underneath that work, and the persistent RPC connection must hold a root that stays alive. It does not belong in the droplet service, and timeouts or retries are not a fix.

Missing evidence, and the read-only command that would produce it

  • Which process sends the first FIN or RST on each buildtest<->:35000 root, and what the process logs at that moment. Command: run on the host, as one fabric launch, tcpdump -l -nn -tttt -i any 'tcp port 35000 and tcp[13]&7!=0' (SYN/FIN/RST) for 60 s. At the same time, take ss -tanpH snapshots every second from a separate shell on the host (not through two concurrent fabric calls, which collided here). Map each ephemeral port to its PID. Then line the buildtest-PID closes up against buildtest-log "Channel closed" / "closed while requests were still pending" entries and ContainerNursery "root generation N" entries.
  • Why the first write on buildtest's persistent transport fails (the trigger that retirement then amplifies). Look for the exception that caused 'retired after the write ... failed'; its cause chain is in the [container:err] lines next to those matches (fabric grep -B/-A on container-nursery.log for 'was retired after the write').
  • Whether admission retirement is involved. The admission effects ("Inbound substream admission is refusing new substreams: X of Y ...") are rate-limited and were not found in the 10k-line in-memory buffers. Command: fabric grep -c 'refusing new substreams' over /root/ContainerNursery/data/logs/container-nursery.log* (buildtest stderr passes through as [container:err]). Also grep the same files for "retire" and "terminal".
  • Which UrlProtocol / UrlResolver versions buildtest-server-hotfix-embedded-0.0.68290276 and ContainerNursery a5a938b1 bundle. Command: fabric unzip -l or jar tf on both jars (read-only).

Recommended operator action (operator's call; nothing was done)

  • Do not restart the droplet service. It is healthy and a restart would not help.
  • A ContainerNursery restart would destroy the evidence of this connection churn (root generation counters, retained state). Take a heap dump or thread dump of ContainerNursery and buildtest first; that requires attaching to the JVM, which is a signal and was not allowed here.
  • Make the durable fix in UrlProtocol/UrlResolver: add fail-first tests showing that a root used for host-initiated outbound persistent RPC is never retired while requests are pending, and that the resolver does not reuse a root that inbound admission has made terminal. Then republish and bump both buildtest and ContainerNursery.

Artifacts

In /tmp/claude-501/-code/35f7fff3-f2da-4da2-94ee-e52357bb0c25/scratchpad/sweep3/drive/inv-droplets/out/: runs*.json, droplets-80k.txt, buildtest-10k.txt, cn-disk-*.txt, cn-deliveryloss.txt, tcpdump-35000.txt, ss-35000.txt, processes.txt, jstat-samples.txt, progress.md.