Handoff: Land the reviewed resolver-pin and follow-up PRs once the buildtest executor recovers; the droplet service is already redeployed
RE-VERIFY before acting. Snapshot 2026-10-06 09:55 UTC by agent fable-20261006-upload-template (paused: blocked on the shared executor, not on code). Re-check PRs with gh pr view <n> --repo CodexCoder21Organization/<repo> --json state,mergeStateStatus,mergedAt, the fleet with https://buildtest.kotlin.build/api/health (pendingBuildRuns, startingDroplets), and the live route with container-nursery-cli health-detailed.
Why this is paused (09:55 UTC)
The kotlin.build (remote) executor has been saturated since before 05:00 UTC: pendingBuildRuns 60–72, startingDroplets 109 (= reservations whose droplets never materialized), availableCapacity 4/200. Three one-question investigations today (evidence under the session scratchpad codex-jobs/*/out/ and in the three challenges linked below) established: (1) DigitalOcean rejects s-2vcpu-4gb in sfo3 with 422 Size is not available in this region for ~10% of creates and buildtest holds each unmaterialized shard reservation for the full 600 s provisioning budget; (2) buildtest DOES retire not-found droplets (phantom-record theory refuted); (3) runs die when their shards lose SSH dispatch transport one by one over 1–3 hours (ten independent failures, aggregated as "10 of 10 shard(s) failed: SSH session is not connected" only when the last shard dies; per-droplet cause NOT established); plus the 180-minute overall deadline outpaces 1,900-test suites at 2–3 shards. The merge queue's own check timeout evicted HandoffEmbedded 23 twice. Nothing in the layers this effort owns fixes that; challenges: https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-06-0653-buildtest-executor-degraded-for-5-hours-on-2026-10-06-full.md , https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-06-0832-buildtest-fleet-saturated-on-2026-10-06-by-unmaterialized.md , and a third filed 09:52 UTC ("buildtest shards lose their SSH dispatch transport one by one…", slug 2026-10-06-095x-… in PlanRepository/challenges).
Resume checklist (in order; all PRs below are reviewed clean by the supervisor and Actions-green; only the remote check gates them)
- Confirm the executor is usable: pendingBuildRuns in the single digits and a recent 1,900-test UrlResolver run completing inside 180 minutes.
- Hold https://githubci.kotlin.build/ warm (bounded curl loop, ≤30 min) and re-request the
kotlin.build (remote) check suite once each for https://github.com/CodexCoder21Organization/UrlResolver/pull/1142 , https://github.com/CodexCoder21Organization/UrlResolver/pull/1146 , https://github.com/CodexCoder21Organization/UrlResolver/pull/1181 (gh api -X POST repos/CodexCoder21Organization/UrlResolver/check-suites/<id>/rerequest), then arm build-watchman --to-merged on each. No watcher is armed any more. https://github.com/CodexCoder21Organization/UrlResolver/pull/1140 (head 432245cd): its run 5207693e was still TESTING at 12:39 UTC with 1348/1891 passed and 6 FAILED reported mid-run, and GitHub already shows the kotlin.build (remote) check as failure (the watcher exited RED at 12:39). Before re-requesting, TRIAGE those 6 failures from the ledger (https://buildtest.kotlin.build/api/test-results?id=5207693e&page=N, api/test-attempts?id=5207693e; the view was still materializing at 12:40 and returned an error) — compare against the known pre-existing flakes (testBytecodeFetchAggregatesRelayFailureDetails, LargeSwarmGossip/DynamicFailureRecovery netlab swarm tests, the two stress tests that flake under executor load) and only re-request if every failure is one of those; a new failure is a signal about PR 1140's diff (sandbox LazyReconnectingMethodDispatcher idle-close rework) and goes to a fix-flakey-test lane first. The watcher for https://github.com/CodexCoder21Organization/UrlResolver/pull/1196 (head 901a1032, run ba31d125 queued since ~08:00 and never started) hit its 180-minute idle timeout at 10:05 UTC — on resume, check whether run ba31d125 ever started; if it is still PENDING leave it queued and just re-arm build-watchman --to-merged, otherwise re-request once like the others. https://github.com/CodexCoder21Organization/TemplateEmbedded/pull/8 (head e03aa6ac) has NO live run: its first run a2c68241 died as a zero-test upload void, the single re-request at 08:28 produced no run, and its watcher hit the 180-minute idle timeout at 10:00 UTC — re-request its kotlin.build (remote) check suite (hold the receiver warm first) and arm a watcher alongside the three UrlResolver re-requests.
- Re-enqueue https://github.com/CodexCoder21Organization/HandoffEmbedded/pull/23 via the GraphQL enqueuePullRequest mutation (merge-queue repo), watch to MERGED, then rebase and land https://github.com/CodexCoder21Organization/HandoffEmbedded/pull/18 .
- https://github.com/CodexCoder21Organization/DigitalOceanDropletServiceServer/pull/143 (81 pin lines → resolver 0.0.1303): enqueue when its remote check lands. It is ALREADY DEPLOYED (see "Deployed but NOT merged" below).
- After landings: close https://github.com/CodexCoder21Organization/UrlResolver/pull/1179 and https://github.com/CodexCoder21Organization/UrlResolver/pull/1180 as superseded by 1181; delete branch
wip/ur-rev-1147-rollback-at-every-position-test; close https://github.com/CodexCoder21Organization/TemplateEmbedded/pull/7 as superseded by 8; fix the misleading on-disk jar name on the droplet route (upload WITHOUT --route under a truthful name, then update-route --image, then restart).
- Follow-ups NOT done (record only): BuildTestServerService resolver bump to 0.0.1303 needs a coordinated BuildTestEmbedded upgrade (its binary-compat test pins protocol 0.0.527; 0.0.1303 pulls 0.0.551); the client-side "ambiguous RPC" errors on getDropletCreationStatus are NOT caused by provider-side closes (correlation refuted it) — candidate: UrlProtocol 120 s child read-idle close without in-flight check (Libp2pHostFactory.kt:5672 at 3697719), evidence in scratchpad
codex-jobs/ur-inbound-close/out/investigation.md.
- Complete this handoff when 1–5 are done.
Deployed but NOT merged — read this first
- Live: route
url:digitalocean-droplets: on 198.199.106.165 runs the jar at /root/ContainerNursery/apps/droplet-service-server-0.0.82-r-derivation-fix-20260827T2205Z.jar, whose filename is misleading: since 2026-10-06 08:52 UTC that path contains the build of DigitalOceanDropletServiceServer commit 5be3123fa57c753da7ce0268bb906cc95da8ca0b (SHA-256 30103d8d2deb18252a7da1a9c512fa8f015abfc7ae4eff242fd182f72317120b) = main 0eb32bf5 + the resolver pin bump to 0.0.1303 that is open in https://github.com/CodexCoder21Organization/DigitalOceanDropletServiceServer/pull/143 (not merged). container-nursery-cli upload-jar --route overwrote the old image in place and ignored --filename (challenge filed; memory cn-upload-jar-route-overwrites-image). The previous build (pre-142 source, resolver 0.0.1092) no longer exists on the host; a rollback must be rebuilt from main at ea337222 with resolver pins 0.0.1092, or use /root/ContainerNursery-uploads/jars/droplet-service-server-0eb32bf5-D72-20261005.jar (main 0eb32bf5 with resolver 0.0.1259 — the build D72 rolled back because of issue 1193; do not use it).
- Acceptance (08:52–08:58 UTC): provider started 08:52:10, bound 08:52:30; 0
Pending dial controller failed to construct transport dial, 0 IllegalArgumentException; serving getDroplet/renewDroplet/deleteDroplet/listDroplets/getDropletCreationStatus; the only handler errors were DigitalOcean 404s for already-deleted droplets and two benign renewDroplet … lifecycle was cancelled (already-deleted droplets). buildtest side: 0 AmbiguousRpcRequestException since the deploy.
- Reconcile by: landing PR 143 (then the on-disk name should be fixed by a clean upload WITHOUT
--route under a truthful filename + update-route --image).
- Published artifact:
foundation.url:resolver:0.0.1303 on kotlin.directory, built from UrlResolver PR 1196 head 901a1032f60d3d88da814578344347452d1a5674 (the PR branch, not main). Reconcile by merging https://github.com/CodexCoder21Organization/UrlResolver/pull/1196 unchanged.
- The droplet route's env still carries an inline
DROPLET_ADMIN_SSH_PRIVATE_KEY; do not copy it; recommend moving it out of the route config.
Remaining work (2026-10-06 09:00 UTC)
- Land https://github.com/CodexCoder21Organization/DigitalOceanDropletServiceServer/pull/143 (81 pin lines, Actions green, supervisor review clean; remote check queued on the saturated executor).
- Land the BuildTestServerService resolver pin bump (a lane is opening it; buildtest-server logs the same dial-construct error ~3/hour).
- Land UrlResolver https://github.com/CodexCoder21Organization/UrlResolver/pull/1196 (head 901a1032, enqueue-on-green watcher), https://github.com/CodexCoder21Organization/UrlResolver/pull/1142, https://github.com/CodexCoder21Organization/UrlResolver/pull/1140, https://github.com/CodexCoder21Organization/UrlResolver/pull/1146, https://github.com/CodexCoder21Organization/UrlResolver/pull/1181 — all reviewed clean, Actions green; only the remote
kotlin.build (remote) check gates them and the executor is saturated (see below). 1181's check was marked failed by the check's own timeout while its run 87636a47 is still TESTING (1493/1921, 0 failed): when the run completes green, re-request the check once (hold https://githubci.kotlin.build/ warm with a bounded curl loop first) and re-arm the watcher. Then close https://github.com/CodexCoder21Organization/UrlResolver/pull/1179 and https://github.com/CodexCoder21Organization/UrlResolver/pull/1180 as superseded by 1181, and delete branch wip/ur-rev-1147-rollback-at-every-position-test.
- Land https://github.com/CodexCoder21Organization/HandoffEmbedded/pull/23 (in the merge queue; was evicted once with
checks_timed_out and re-enqueued), then rebase and land https://github.com/CodexCoder21Organization/HandoffEmbedded/pull/18.
- Land https://github.com/CodexCoder21Organization/TemplateEmbedded/pull/8 (ready, reviewed clean, pinned to the published
com.squareup.okio:okio-fakefilesystem-jvm:3.4.0-te-cost.1; first remote run was an upload void, re-requested once), then close https://github.com/CodexCoder21Organization/TemplateEmbedded/pull/7 as superseded.
- Complete this handoff.
Executor saturation — established 2026-10-06, not a bug in the layers this effort owns
https://buildtest.kotlin.build/api/health shows startingDroplets≈109 / running≈48 / available≈4 of 200 for 8+ hours; 44–53 runs PENDING; 1,900-test suites crawl at ~5 tests/min and hit the 180-minute watchdog. Established (two codex investigations, file:line evidence in the challenge): startingDroplets = observed booting droplets + the provider-invisible remainder of granted reservations (BuildTestEmbedded 0.0.615130200605, BuildTestEmbeddedService.kt:1848–1909); 27 runs hold ~196 reservations while ~87–101 droplets exist; the provider saw 14/150 creates rejected by DigitalOcean 422 Size is not available in this region (s-2vcpu-4gb, sfo3) and holds the shard reservation until the 600 s provisioning budget expires. buildtest DOES retire not-found droplets (refuted theory); provider "cancelled renewal" errors are benign. Challenge: https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-06-0832-buildtest-fleet-saturated-on-2026-10-06-by-unmaterialized.md. Earlier challenge: https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-06-0653-buildtest-executor-degraded-for-5-hours-on-2026-10-06-full.md. No operational endpoint reconciles reservations; cleanupDroplets deletes LIVE droplets — never use it for this.
Branches and unmerged state (verified 09:00 UTC)
| Repo |
Branch |
Head |
PR |
What |
State |
| UrlResolver |
fix/recovery-dial-circuit-relay-address |
901a1032 |
https://github.com/CodexCoder21Organization/UrlResolver/pull/1196 |
circuit-relay dial exclusion + 0.0.1303 bump |
Actions green; remote check running |
| DigitalOceanDropletServiceServer |
fix/resolver-0-0-1303-circuit-relay-dial |
5be3123f |
https://github.com/CodexCoder21Organization/DigitalOceanDropletServiceServer/pull/143 |
81 resolver pins → 0.0.1303 |
Actions green; DEPLOYED (see above) |
| TemplateEmbedded |
fix/template-commit-cost |
e03aa6ac |
https://github.com/CodexCoder21Organization/TemplateEmbedded/pull/8 |
published-pin + allocation guard |
ready; remote re-requested |
| okio (fork) |
— |
merged |
https://github.com/CodexCoder21Organization/okio/pull/2 |
FakeFileSystem lookup fix + backport script |
MERGED 07:39 UTC |
| HandoffEmbedded |
fix/state-persistence-write-cost |
c5c82ea0 |
https://github.com/CodexCoder21Organization/HandoffEmbedded/pull/23 |
org.json-byte-identical state writer |
in merge queue |
| BuildTestServerService |
main |
— |
https://github.com/CodexCoder21Organization/BuildTestServerService/pull/385, https://github.com/CodexCoder21Organization/BuildTestServerService/pull/386 |
— |
MERGED 04:52 / 05:59 UTC |
Handoff: Investigate D72 rollout verification failure before another droplet-service deploy
RE-VERIFY - 2026-10-05 10:27:35 UTC. All below is a write-time snapshot. Re-check filtered CN routes and containers, facade health and host jar hashes. D72 used explicit per-lane production authorization, deployed main once, failed new-error verification, and applied the required rollback. Orchestrator controls completion; do not mark this handoff complete from this report.
D72 final report - BLOCKED after rollback
Written: 2026-10-05 10:27:35 UTC. Agent: fable-drain-20261004.
Original request: deploy DigitalOceanDropletServiceServer main containing https://github.com/CodexCoder21Organization/DigitalOceanDropletServiceServer/pull/142 to url:digitalocean-droplets: only, with baseline, unique-file upload, image-only route update, live verification, and durable evidence.
The authorized merged main was built and uploaded under a new filename, and the route was changed only in its image field. The replacement passed facade health and bytecode identity checks but produced new WebCron connection and recovery-dial errors and refused provisioning requests. The recorded rollback restored the old image, which is now exactly one RUNNING container and returns facade health OK.
Status: BLOCKED. Orchestrator action needed: assign investigation of replacement-generation WebCron EOF/recovery-dial and renewal-lifecycle errors before another authorized rollout. No deploy retry is planned. The underlying first transport-close cause is not established; this report does not claim every historical provisioning failure shares one mechanism.
Merged reference: https://github.com/CodexCoder21Organization/DigitalOceanDropletServiceServer/pull/142 - MERGED at 2026-10-02T12:02:19Z, merge commit 0eb32bf513f8e977c3538cb6e28f7916bd340353. Main was independently fetched/rebased and ancestor-checked before build. No PR created, source changes, tests altered, merge/enqueue or Maven publication.
Delegated work: no delegated work active or recently finished. Landing parallelization sweep: not applicable; this lane owns one deployment item and no landing gate.
Evidence: https://github.com/CodexCoder21Organization/PlanRepository/tree/wip/D72-droplet-deploy-20261005/handoffs/artifacts/D72. Checkpoint https://github.com/CodexCoder21Organization/PlanRepository/commit/75403505470e559456ca55983059965d26b5bbda is remote-verified; subsequent final reporting evidence remains on the same branch.
Handoff: https://www.handoff.wasmserver.com/handoffs/hf-2026-09-14-finish-landing-the-upload-ux-library-and-template-work-drive-four-green-prs-through-review-to-merge-deploy-the-two-web-apps-and-complete-the-template-service-family (authoritative url://handoff/handoffs/hf-2026-09-14-finish-landing-the-upload-ux-library-and-template-work-drive-four-green-prs-through-review-to-merge-deploy-the-two-web-apps-and-complete-the-template-service-family).
D72 deployment findings
Agent: fable-drain-20261004
Route: url:digitalocean-droplets:
Plan
- DONE: Claim handoff, read required instructions and procedure, verify main and live premises.
- DONE: Capture baseline, check concurrent activity, build and compare artifact shape.
- DONE: Record rollback, audit intent, upload new jar and change only route image.
- DONE: Verified; new-error acceptance failed, mandatory rollback applied and verified.
- IN PROGRESS: Final handoff update, BLOCKED report and process cleanup; durable evidence push complete.
Progress
2026-10-05 UTC - OBSERVED: Received explicit authorization for this route only. No production changes have been made.
2026-10-05 10:07 UTC - OBSERVED: Claimed and cleared blocked reason. PR https://github.com/CodexCoder21Organization/DigitalOceanDropletServiceServer/pull/142 is MERGED at 0eb32bf513f8e977c3538cb6e28f7916bd340353; remote main matches; merge-base ancestor check succeeds. Fresh checkout read README and ARCHITECTURE, plus Engineering Philosophy and DigitalOceanDroplets documentation. No source edits. Route snapshot matches expected image, memory 1024 MB, keepWarmSeconds -1 and seven env names. Old jar size 60145532, SHA256 552bfd3869ccf226fb3534c9745787cc0eb55b4fb5237014cdd2a2929e4fd488, mtime Sep 27 00:24:10 UTC, main class dropletserviceserver.MainKt.
OBSERVED workaround: Fabric certificate directory absent; used brief-authorized SSH key cn_host_diag on port 23 for stat/hash/manifest/audit-tail reads. CLI 0.0.20 route JSON uses nested facade.configuration.domain and container.envvars, not route_key/env. Brief's exact filter returned [] (exit 0); corrected to domain filter and removed envvars, retaining names only. Container JSON uses domain, route_type, host_port, image, state. No environment values were printed or saved.
INFER: Existing relay disconnect cleanup ERROR lines are pre-deploy background messages, not deployment results. Baseline health and container capture are still in progress. Prior evidence explicitly does not establish that this upgrade alone resolves every historical EOF; verification will report only observed scope.
2026-10-05 10:10 UTC - OBSERVED milestone: Baseline health via published digitaloceandropletcli 0.0.9 exit 0, service url://digitalocean-droplets/, 106 droplets, latency 16330ms. The CLI also emitted a bootstrap peer-exchange timeout (full output saved); this did not prevent health completion. Coordinator 400-line baseline contains accepted createDropletAsync calls (request IDs c25f4171, feb2e415, a26d1bea) during the window. Route container exactly one RUNNING old-image entry on 44207. Baseline checkpoint pushed and remote-verified at https://github.com/CodexCoder21Organization/PlanRepository/tree/wip/D72-droplet-deploy-20261005, commit https://github.com/CodexCoder21Organization/PlanRepository/commit/ba60ea4961d7bef210cad30aa6338599be9ca4c2.
OBSERVED: Build launched after fresh fetch/rebase on clean main 0eb32bf513f8e977c3538cb6e28f7916bd340353, local fat-jar mode, JDK_JAVA_OPTIONS=-Xmx512m because build.bash replaces JAVA_OPTS. PID tracked in build.pid, hard 900-second deadline. Dependency-pin warnings emitted by existing build tool are retained in build.log; no source change or test alteration made.
Rollback recorded before route change
coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com update-route --facade url --domain digitalocean-droplets --image /root/ContainerNursery/apps/droplet-service-server-0.0.82-r-derivation-fix-20260827T2205Z.jar
2026-10-05 10:12 UTC - OBSERVED: Initial handoff checkpoint update succeeded; its client printed Netty warning "An exceptionCaught() event was fired, and it reached at the tail of the pipeline. It usually means the last handler in the pipeline did not handle the exception." followed by io.netty.channel.StacklessClosedChannelException at io.netty.channel.AbstractChannel$AbstractUnsafe.write(Object, ChannelPromise)(Unknown Source). No change to deployment state. Baseline health/log checkpoint pushed at 3d0029a3f. Host audit DEPLOY format inspected read-only. Old jar retained, exact rollback above.
2026-10-05 10:14 UTC - OBSERVED milestone: Fat-jar build exit 0, 262827 ms, 61462057 bytes. No tests executed by this build rule; no source code changes made. Build warnings and exact compiler output saved in build.log. Child JVM was actively writing JarBuilder output when sampled, so the earlier silent build interval was packaging, not an unexplained stall.
OBSERVED tooling finding: Direct facade probe first sent bare health. Facade peer 12D3KooWLMyXNfwhcX1YsiNx3hnjk3GGSfsU1fydRa8bzrE6scMT answered RpcError(code=UNKNOWN_METHOD, message=No service registered for identifier: health, data=null); FacadeProbe.main(FacadeProbe.java:23) threw IllegalStateException. Corrected request to digitalocean-droplets/health (service-prefixed path). Local probe also emitted JmDNS network warnings and LAN discovery parse warnings, preserved in evidence/facade-before.txt. These are probe-client messages, not new service errors. Baseline published CLI health remains successful.
2026-10-05 10:15 UTC - OBSERVED: Corrected direct facade probe returns health=OK and bytecode keys jar,className,stdlibJar (exit 0). Artifact shape verified: main class dropletserviceserver.MainKt, 23819 classes, embedded client-impl.jar, size 61462057 vs live 60145532 (ratio 1.022). SHA256 ed1f42e91c4ed8bb8b3cbeac9e805dc1afdb9974ba6ac225d2d7251cc79ff4ec. All resolver 0.0.1259 classes 1244/1244 and protocol 0.0.527 classes 1094/1094 match published jars byte-for-byte (zero missing/different). Deploy decision: proceed after one final unchanged-route/host check. New filename /root/ContainerNursery-uploads/jars/droplet-service-server-0eb32bf5-D72-20261005.jar. Old path untouched. No Maven publication.
2026-10-05 10:17 UTC - OBSERVED milestone: Preflight route redacted JSON identical to baseline; old jar stat and hash unchanged. CN schema exposes no route-change timestamp, so concurrency check used exact snapshot equality plus old jar mtime/hash and audit inspection. DEPLOY-INTENT appended to named audit path via SSH stdin. New file uploaded with shell noclobber, no overwrite, and host stat/hash/manifest verified. New size 61462057, SHA256 ed1f42e91c4ed8bb8b3cbeac9e805dc1afdb9974ba6ac225d2d7251cc79ff4ec. Image-only update-route started; its output withheld because CLI responses may contain route values. Verification will use filtered route query. Latest checkpoint https://github.com/CodexCoder21Organization/PlanRepository/commit/eba5c88c9a05a4a4860735253369afdd33fe76c8.
2026-10-05 10:19 UTC - OBSERVED: Image-only update-route exit 0. Filtered routes query names /root/ContainerNursery-uploads/jars/droplet-service-server-0eb32bf5-D72-20261005.jar. Redacted route before/after comparison shows image as the only difference (env names, memory, keep-warm, facade, notes, absent args unchanged). New request through the facade is in progress; exactly-one RUNNING and log checks pending. No second deployment attempt, restart or other route action.
2026-10-05 10:20 UTC - OBSERVED milestone: Exactly one CN container for url:digitalocean-droplets: RUNNING on new image at host_port 34171. Direct facade health=OK, client bytecode fetched: 2045188 bytes SHA256 6382bd039b082f182831b6742b99eb171774706de95a5077a5b82c5ee6ce2ba5, byte-identical to the built client-impl.jar resource. stdlibJar 668943 bytes SHA256 cf03ecafcf482692c42f56c0367b319ff74b87d3f6f5e2b1bf4a22dc6bbafabb. Facade probe exit 0. Client-side mDNS/LAN warnings also existed before and are not classified as service warnings. First log/coordinator sample still pending.
2026-10-05 10:22 UTC - OBSERVED course-changing finding: New image passes facade health and bytecode, but verification 6d fails. Baseline log samples had 276 and 233 WARN/ERROR lines respectively, zero WebCron error lines. Post-deploy first sample has 17 WebCron error lines, new active-query timeouts, "Pending dial controller failed to construct transport dial" IllegalArgumentException messages, repeated listJobs EOF / invalidation / skipped sweep registration, and refusals for new async requests because no WebCron connection attempt completed. New-generation refusals start 10:19:08.811, including cca740ec-561b-4ca6-a380-a6ba9be79c8a / 213ece38-4d80-42f9-bff0-52593cd82d66 / 4e09e2a7-0dff-4135-9832-c177998c4e99. The observable mechanism: the replacement lacks a usable WebCron connection; its required pre-creation registration cannot run, and the existing fail-closed code refuses provider creation. Main.kt invalidation at lines 565-579 clears and closes failed WebCron handles after transport EOF. Underlying first transport-close cause is not proven by this lane; do not claim the historical bug is resolved.
OBSERVED: Coordinator after sample includes getDroplet AmbiguousRpcRequestException at 10:18:48.857, stale-provider dial failures, createDropletAsync stalls and deleteDroplet relay failures. Some are during container replacement and cannot be attributed solely to the final new generation. No provisioning absence caveat applies: provisioning was observed. No packet-capture claim is made.
Decision: apply mandatory recorded rollback for failed 6d and explainable new provisioning refusal, then verify old image RUNNING and report BLOCKED. Do not retry deploy. Full exact error lines will be pushed; no auto-merging challenge CLI used.
OBSERVED local bookkeeping correction: served-client comparison first ran from PlanRepository directory and failed FileNotFoundError for the local jar. Re-ran from lane directory at 10:20:37; exact-match assertion passed and evidence/served-bytecode-comparison.json now exists. No production effect.
2026-10-05 10:24 UTC - OBSERVED rollback milestone: rollback update-route exit 0 at 10:23:23 UTC. Route redacted JSON identical to full baseline. Exactly one restored old-image container RUNNING on host_port 36661. Facade health=OK; served old client jar SHA256 096d9dcd8dfc36749a5682b20c1f6583e9f6c89b77ce6ec55146857bb7b6c70c, matching old host jar resource. Old host jar retains size 60145532, mtime Sep 27, hash 552bfd3869ccf226fb3534c9745787cc0eb55b4fb5237014cdd2a2929e4fd488. New file remains under its unique filename with unchanged size/hash. DEPLOY-END appended with ROLLED-BACK-BLOCKED outcome.
OBSERVED warning inventory: two post-deploy 400-line samples span 10:17:42-10:19:30 and 10:20:32-10:21:04; 188 unique WARN/ERROR lines, 175 absent from baseline after timestamp removal, including 19 WebCron lines, 49 recovery-dial lines, 6 async creation failure lines. Every exact new line is in new-warning-lines.txt, with full surrounding traces in logs-after-1/2.txt. Existing relay and Netty chatter remains identified as baseline; timestamp-only comparison intentionally inventories changed peer IDs/locations too, without calling all of those new behavior. Second sample also records renewal lifecycle cancellation and unconfirmed WebCron timer reconciliation at 10:21:04.215. No pass claim for error absence.
OBSERVED local evidence-generation correction: warning inventory first ran from PlanRepository and failed FileNotFoundError for evidence/logs-before.txt; re-ran from lane directory and saved all inventory files. No production effect.
FINAL
Status: BLOCKED after mandatory rollback. No deployment retry, other route change, Maven publication, merge/enqueue, test change, source implementation change or handoff completion.
Attempted deployment: source 0eb32bf513f8e977c3538cb6e28f7916bd340353, /root/ContainerNursery-uploads/jars/droplet-service-server-0eb32bf5-D72-20261005.jar, 61462057 bytes, SHA256 ed1f42e91c4ed8bb8b3cbeac9e805dc1afdb9974ba6ac225d2d7251cc79ff4ec.
Currently live: /root/ContainerNursery/apps/droplet-service-server-0.0.82-r-derivation-fix-20260827T2205Z.jar, 60145532 bytes, SHA256 552bfd3869ccf226fb3534c9745787cc0eb55b4fb5237014cdd2a2929e4fd488. Old source identity is not inferred from filename.
| Check |
OBSERVED result |
| Main contains authorized merge |
PASS: PR MERGED; exact main SHA; ancestor verified |
| Baseline / concurrent operator check |
PASS: old jar mtime Sep 27; preflight route snapshot identical; no route timestamp field exposed |
| Build / artifact shape |
PASS: exit 0; 61.5 MB vs 60.1 MB; same main class; fat jar with 23819 classes and embedded client |
| Resolver / protocol identity |
PASS: 1244/1244 resolver 1259 and 1094/1094 protocol 527 classes byte-identical |
| New route and single RUNNING container |
PASS while deployed: new path, exactly one on 34171 |
| Facade health / served bytecode |
PASS while deployed: OK; built embedded client exactly matches served bytes |
| Coordinator provisioning symptom |
FAIL/UNRESOLVED: provisioning observed; droplet RPC EOF/stall and relay errors remain in window; cutover timing limits attribution |
| New WARN/ERROR absence |
FAIL: WebCron connection errors, recovery dial errors, provisioning refusals and renewal reconciliation errors |
| Rollback |
PASS: route fully restored; exactly one old-image container RUNNING on 36661; facade OK; old served client and file hashes match |
| Durable evidence |
WIP branch https://github.com/CodexCoder21Organization/PlanRepository/tree/wip/D72-droplet-deploy-20261005/handoffs/artifacts/D72 |
Rollback command (already executed): coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com update-route --facade url --domain digitalocean-droplets --image /root/ContainerNursery/apps/droplet-service-server-0.0.82-r-derivation-fix-20260827T2205Z.jar
Remaining: orchestrator must assign investigation of the observed replacement-generation WebCron EOF/recovery-dial and lifecycle-cancellation behavior before authorizing another rollout. This lane proves artifact identity and rollback, not a cause for every historical provisioning failure. No production implementation fix is guessed. Full error text is retained for that investigation; new image preserved for review.
Workarounds/findings: absent Fabric credentials -> authorized SSH fallback; CLI nested route/container schema differs from brief -> corrected filters; build.bash replaces JAVA_OPTS -> 512 MB JDK_JAVA_OPTIONS cap; local build for artifact >32 MiB; published CLI health emitted bootstrap exchange warning; direct facade bare method rejected -> service prefix; direct probe emitted local LAN/mDNS warnings; handoff client Netty warning despite successful update; route timestamp field absent -> snapshot/hash check. Local evidence path errors corrected as recorded. No challenge auto-merge CLI run.
Documentation read: DigitalOceanDropletServiceServer README.md and ARCHITECTURE.md; DocumentationRepository PHILOSOPHY.md and projects/DigitalOceanDroplets.md; PlanRepository README.md and handoffs/README.md; prior L6b investigation.md; DigitalOceanDropletCli README; applicable GitHub, deployment, Fabric and status-report skills. No unit tests were written or modified.
2026-10-05 10:26 UTC - OBSERVED: Final verification evidence checkpoint pushed and remote-verified at https://github.com/CodexCoder21Organization/PlanRepository/commit/75403505470e559456ca55983059965d26b5bbda. Both repository working trees clean. No build jars or credentials committed.
Next steps: orchestrator reviews the rollback evidence and assigns the observed WebCron/recovery-dial investigation before considering another deployment.
Lessons: health OK does not prove provisioning or timer registration works; compare application errors with the pre-change baseline. Keep old files unchanged and record rollback before a route mutation so a failed verification can be reversed directly. Confirm the running image and served bytecode independently; filenames and dependency pins alone do not establish live identity.
Prior handoff body retained for history
Handoff: Deploy DigitalOceanDropletServiceServer main 0eb32bf5 after explicit user authorization
RE-VERIFY — 2026-10-04 11:52 UTC. All state below is a write-time snapshot. Before acting, re-check the PRs with gh pr view <n> --repo <owner/repo> --json state,mergeStateStatus,statusCheckRollup,mergedAt,headRefOid, branch heads with git ls-remote, and the route with ContainerNursery CLI routes --json and containers --json. The live checks below supersede older snapshots preserved after this section.
Mission and remaining scope
The original upload UX, Library view, and Template family work is complete. This handoff now has one remaining action: deploy the merged DigitalOceanDropletServiceServer URL-stack upgrade after the user explicitly authorizes the production change. This lane did not deploy, publish, restart, merge, enqueue, or change any repository.
Current verified state
- DigitalOceanDropletServiceServer PR 142 is MERGED at
307c9bd5f3b7fdacbd970417f2a0f28edf4c6282 (2026-10-02 12:02:19 UTC). Its bld-all-tests, bld-build, kotlin.build (kompile-remote-build), and required kotlin.build (remote) checks report SUCCESS.
- Current
main is 0eb32bf513f8e977c3538cb6e28f7916bd340353, and its commit subject is Upgrade DigitalOcean droplet service URL stack (#142). The merged PR branch remains at 307c9bd5f3b7fdacbd970417f2a0f28edf4c6282: upgrade/url-stack-1259-527-sjvm-0.0.51.
- PR 142 pins resolver
0.0.1259, protocol 0.0.527, and libSJVM 0.0.51, confirmed in the checked-in build.kts at current main.
- ContainerNursery currently reports route
url:digitalocean-droplets: as RUNNING on host port 44207, using /root/ContainerNursery/apps/droplet-service-server-0.0.82-r-derivation-fix-20260827T2205Z.jar; route memory is 1024 MB and keep-warm is -1.
- Carried forward from the other lane's evidence, as requested: the running jar's classes match resolver
0.0.1092 and protocol 0.0.470. L11 did not repeat the class-by-class comparison. The source evidence branch is wip/L6b-oct2-rpc-evidence-20261004 at c497f854782d9b0fad35483f8099bc6652e87412.
- Route inspection displayed only environment-variable names, not values. The route has an inline
DROPLET_ADMIN_SSH_PRIVATE_KEY; keep every existing environment variable and the 1024 MB memory limit unchanged. Do not copy or print secret values.
UrlResolver PR tracking note
The open broad queue is handoff hf-2026-10-02-drive-the-remaining-open-urlresolver-prs-to-completion. Its latest checkpoint is working other PRs and explicitly lists 1140/1142 as other-agent work and 1146 as other-session work; the live claim was held by lane-hfURrest with heartbeat at 11:46 UTC. This is the routing pointer only. L11 did not edit that handoff or change PR ownership. These three PRs are recorded here so the orchestrator can reconcile them without keeping their work in this handoff.
Exact rollout after authorization
No Maven publication is part of this rollout: build the service fat JAR from the current merged main and deploy that file. Re-check that main still contains PR 142 before building.
-
From a clean DigitalOceanDropletServiceServer checkout, update main to the current origin/main and build outside the repository:
git fetch origin
git switch main
git pull --ff-only origin main
./scripts/build.bash --remote dropletserviceserver.buildFatJar() /tmp/droplet-service-server-0eb32bf5-20261004.jar
sha256sum /tmp/droplet-service-server-0eb32bf5-20261004.jar
At this verification, origin/main is 0eb32bf513f8e977c3538cb6e28f7916bd340353. If main advances before authorization, use its new SHA and update the JAR filename accordingly.
-
Before changing the route, append the established DEPLOY-INTENT entry to /root/cn-watchdog.audit using HardwareControlFabric write-file --append. Include the source SHA and old/new image paths. Upload the JAR under a new filename so the rollback file remains untouched:
coursier launch community.kotlin.hardwarecontrolfabric:cli:0.2.6 -r https://kotlin.directory -- \
-H 198.199.106.165 \
-c ~/.config/hardware-control-fabric/fabric-client-cert.pem \
-k ~/.config/hardware-control-fabric/fabric-client-key.pem \
--ca-cert ~/.config/hardware-control-fabric/fabric-server-cert.pem \
write-file \
--remote-output-path /root/ContainerNursery-uploads/jars/droplet-service-server-0eb32bf5-20261004.jar \
--local-input-path /tmp/droplet-service-server-0eb32bf5-20261004.jar \
--create-dirs
-
Change only the route image with ContainerNursery CLI 0.0.20 or newer:
coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- \
update-route \
--facade url \
--domain digitalocean-droplets \
--image /root/ContainerNursery-uploads/jars/droplet-service-server-0eb32bf5-20261004.jar \
--url https://api.nursery.wasmserver.com
URL routes omit --port. The CLI changes only fields provided. ContainerNursery's PUT /routes stops the active container for that route; the replacement starts on the next request, so no separate restart command is needed. The existing environment variables, memory limit, and other route fields remain in place.
-
Verify the route now names the new image. After the next normal request starts the service, verify containers --json reports the route RUNNING on the new image. Confirm the live service serves the PR 142 resolver/protocol pair, BuildTest no longer logs AmbiguousRpcRequestException for droplet RPCs, and dependent services reconnect. For the packet check, rediscover the current BuildTest libp2p port at that time; do not reuse historical port 46411. The acceptance condition from the prior capture is that the droplet-side child-stream open is no longer followed by the reset and FIN at approximately 1.000 second.
-
Append the established DEPLOY-END entry to /root/cn-watchdog.audit with the outcome and artifact hash. The audit file is too large for HardwareControlFabric read-file; preserve the existing audit format rather than copying or truncating it.
Rollback: change only the route image back to the currently configured jar, then let the next request start it:
coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- \
update-route \
--facade url \
--domain digitalocean-droplets \
--image /root/ContainerNursery/apps/droplet-service-server-0.0.82-r-derivation-fix-20260827T2205Z.jar \
--url https://api.nursery.wasmserver.com
The previous kotlin.build (remote) check for PR 142 reports SUCCESS in GitHub. A separate attempt to recover exact ledger counts from run 07248a1b was incomplete: page 1 showed 100 PASSED of a reported total 275 with complete=false, pages 2 and 3 returned HTTP 503, and a later attempt query returned an empty list. No additional tests or CI reruns were started, and this handoff does not claim an exact pass/fail count from that ledger.
Remaining work
Only the production rollout above remains, pending the user's explicit authorization. No deployment, restart, publication, or UrlResolver PR work was performed by L11.
Prior write-time snapshot (preserved for history)
The following body is retained from the earlier handoff snapshot. Its dates, statuses, ownership notes, and next-step lists are historical; the current state and commands at the top of this handoff take precedence.
Deploy the droplet service from merged PR 142 (resolver 0.0.1259) and settle ownership of UrlResolver PRs 1140, 1142 and 1146 — the Template family is complete
Re-verified 2026-10-03 15:35–15:50 UTC by a Claude session (agent fable-code-20260914's successor, running as Opus) on the user's request "Create/update handoff".
RE-VERIFY: everything below is a write-time snapshot. PRs: gh pr view <n> --repo CodexCoder21Organization/<repo> --json state,mergeStateStatus,headRefOid,statusCheckRollup,mergedAt. Live routes: CN CLI ≥0.0.20 routes --json (coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com, coursier = /home/u/bin/cs). Gate ledgers: read every page of https://buildtest.kotlin.build/api/test-results?id=<run>&page=N&pageSize=300&attempts=true AND api/test-attempts?id=<run> (a ledger whose final statuses are all PASSED can hide retried FAILED attempts); a zero-test FAILURE (provisioning deadline, SSH session, run.json persistence) is infrastructure — re-request, never a verdict.
Mission summary
Original request: land the upload-UX work (FiledropWui + PhotoGenerationManagerWui per the GOOD_DESIGN.md File Uploads section, "fixed/upgraded merged and deployed"), the PhotoGenerationManagerWui Library view, and the Template family — TemplateApi, TemplateEmbedded, then TemplateServiceServer at url://templates/ and TemplateCli. Later directives: "codex quota reset, drive this to completion"; "you are fable, delegate to codex"; 2026-09-28 05:25 UTC "Drive this to completion delegating to opus instead of codex/gpt"; 2026-09-28 19:14 UTC "stop and update the handoff"; 2026-10-03 "Create/update handoff".
The original request is DONE. Upload-UX, Library view, TemplateApi, TemplateEmbedded, TemplateServiceServer and TemplateCli are all merged, and community.kotlin.templates:template-cli:0.0.2 is published (POM 200; the last Template-family step).
What remains in this handoff (two items):
- Deploy the droplet service from merged https://github.com/CodexCoder21Organization/DigitalOceanDropletServiceServer/pull/142 (merged 2026-10-02 12:02Z at 307c9bd5f3; resolver 0.0.1259 / protocol 0.0.527 / SJVM 0.0.51; artifact coordinate 0.0.85). Not deployed: the
url:digitalocean-droplets: route still runs /root/ContainerNursery/apps/droplet-service-server-0.0.82-r-derivation-fix-20260827T2205Z.jar (resolver 0.0.1092), and 0.0.85 is not published either (latest on kotlin.directory is 0.0.84; a fat-jar deploy does not need the publish). This is the established mechanism-level fix for buildtest's ambiguous-EOF droplet RPC failures — see "Mechanism" below. Steps: build dropletserviceserver.buildFatJar from main ≥307c9bd5f3; HCF write-file it to /root/ContainerNursery-uploads/jars/droplet-service-server-<sha>-<date>.jar (never upload-jar); update-route image only (keep memoryLimitMb 1024 and every envvar — the route env holds an inline DROPLET_ADMIN_SSH_PRIVATE_KEY: do not copy it; recommend moving it out); DEPLOY-INTENT/END lines in /root/cn-watchdog.audit; rollback = the 0.0.82 jar path above. Acceptance: capture SYN/FIN/RST on buildtest's libp2p port (it was :46411 on 2026-09-28; re-read with ss -tnp) — the droplet-side "50 B + 57 B child-stream open, only ACKs back, reset + FIN at exactly 1.000 s" pattern must be gone; AmbiguousRpcRequestException against digitalocean-droplets must stop appearing in buildtest-server stdout; dependants (droplet WUI, netlab, screenshottest) reconnect. Coordinate first with https://handoff.wasmserver.com handoff hf-2026-10-02-establish-why-coordinator-to-droplet-service-rpcs-fail-through-the-containernursery-url-facade-and-fix-the-fleet-wide-provisioning-failures: it is investigating the same EOF symptom (run 866f31b1, 2026-10-02 18:40–19:35Z) and its body does not mention resolver 0.0.1092 on the droplet side, the 1 s whole-connection close, or PR 142 — whoever picks either handoff up should merge the two lines of investigation.
- Settle ownership of the three still-open UrlResolver PRs from this effort (none should have two owners). The 2026-10-02 handoff
hf-2026-10-02-drive-the-remaining-open-urlresolver-prs-to-completion-land-1157-1167-1119-1153-fix-1154-s-failing-test-resolve-1155-and-1149-then-salvage-the-seven-stale-prs lists 1146 in its "stale — salvage" queue and says 1140 and 1142 belong to "other agents — do not touch" (i.e. this handoff). Recommended: hand all three to that UrlResolver handoff (one owner per repository's PR queue) and drop them from this one; until then:
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1146 (OPEN, head b7a24f48cb, 9 commits; idle-reaper head-of-line fix +
markClosed() joins in-flight idle closes; round-2 review never re-run). bld-build FAILURE at 04:05Z on 10-03 (failing test not yet named — Actions run 37092442431, read gh run view 37092442431 --log-failed); kotlin.build (remote) run abfe50fd was a zero-test provisioning-deadline void. A related newer PR https://github.com/CodexCoder21Organization/UrlResolver/pull/1179 ("Keep early sandbox idle callbacks from losing cleanup") touches the same dispatcher — check overlap before salvaging 1146.
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1142 (OPEN, head f9171707a5, 9 commits; returning provider supersedes a withdrawal tombstone by strict timestamp; discriminating test proven 11/11 against the mutant).
bld-build SUCCESS; kotlin.build (remote) run 02b14432 FAILED 1890/1895 — FAILED attempts on 13 tests: stressTestActiveQueryMalformedResponseAlwaysReachesParsing, stressTestWithdrawalLossDuringInitialBootstrapHandshake, testActiveQueryHonoursNoTimeoutBudgetOnSkewedResolverClock, testActiveQueryRetriesAfterRecognizedWrongResponseType, testConcurrentColdSandboxFactoriesAcrossDispatchersStayWithinHeapBudget, testDeeplyNestedMapReturnIsRejectedNotStackOverflow, testOnPeerAnnouncementReturnsQuicklyWithUnreachablePeers, testOnPeersReceivedReturnsQuicklyWithUnreachablePeers, testOnServiceAnnouncementReturnsQuicklyWithUnreachablePeers, testOversizedGossipAdmissionPreservesAlreadyQueuedPeerTraffic, testRelayBackedLongMaxRequestUsesFallbackAfterClockAdvances, testSandboxedProxyRedialsAfterSilentPersistentRpcTimeout, testUrlResolverServiceRegistrationAndDiscoveryViaRelayIsolated. Broad, mostly unrelated to the diff (a fleet-wide degraded run is likely) — compare against a main run of the same window before attributing anything; each confirmed flake goes to fix-flakey-test.
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1140 (OPEN, head 89acb1a200; large-RPC own-connection harness fix, reviewed READY earlier).
bld-build SUCCESS; kotlin.build (remote) FAILED 1884/1888 with "10 of 10 shard(s) failed: SSH session is not connected" — infrastructure; re-request.
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1141 was closed on 2026-10-01 by a Fable decision in favour of https://github.com/CodexCoder21Organization/UrlResolver/pull/1156 (same flake, fixed in all three affected tests). Nothing to do.
Landed since the 2026-09-28 snapshot (verified 2026-10-03)
| PR |
Merged |
Note |
| https://github.com/CodexCoder21Organization/TemplateCli/pull/2 |
10-02 15:43Z |
pins to server/client 0.0.2 + embedded 0.0.4; template-cli:0.0.2 published |
| https://github.com/CodexCoder21Organization/UrlResolver/pull/1143 |
09-30 00:58Z |
buffered PeerExchange harness fix |
| https://github.com/CodexCoder21Organization/UrlResolver/pull/1144 |
09-30 10:04Z |
replacement dial honours the opener's budget |
| https://github.com/CodexCoder21Organization/UrlResolver/pull/1147 |
09-29 09:42Z |
projection partial-seed fix |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1129 |
09-30 01:36Z |
ambiguous droplet reads delivered as plain RuntimeException are reissued |
| https://github.com/CodexCoder21Organization/BuildTestServerService/pull/377 |
09-30 00:43Z |
detachedForCaller keeps the SandboxException type (the owning-layer fix) |
| https://github.com/CodexCoder21Organization/DigitalOceanDropletServiceServer/pull/142 |
10-02 12:02Z |
resolver stack upgrade — merged, not deployed |
Deployments now live (not this session's; recorded so nobody reverts them): url:buildtest: → /root/ContainerNursery-uploads/jars/buildtest-server-2caebf69-20261003-deploy4b.jar (BuildTestServerService #390, embedded 0.0.69274746 — includes #1129 and #377); url:handoff: → /root/ContainerNursery-uploads/jars/handoff-service-server-d2ab35f-20261002.jar (HandoffServiceServer #26 / Embedded 0.0.23 — supersedes this effort's 2026-09-28 deploy of 99ab7220, which carried the streamed-persistence OOM fix that remains in the newer jar).
Follow-ups from this effort now owned elsewhere (do not duplicate): the unguarded cause-chain walk that hangs (BuildTestServerService https://github.com/CodexCoder21Organization/BuildTestServerService/pull/385) and discarded-generation failures (https://github.com/CodexCoder21Organization/BuildTestServerService/pull/386); handoff startup full-string read (https://github.com/CodexCoder21Organization/HandoffEmbedded/pull/18); projection subscription/pad holes found in the 1147 review (UrlResolver https://github.com/CodexCoder21Organization/UrlResolver/pull/1180, https://github.com/CodexCoder21Organization/UrlResolver/pull/1181); idle-callback cleanup (https://github.com/CodexCoder21Organization/UrlResolver/pull/1179); TemplateEmbedded near-budget point-cost tests (https://github.com/CodexCoder21Organization/TemplateEmbedded/pull/7). All open at write time and authored by other sessions.
Mechanism for item 1 (established 2026-09-28 with two packet captures; VERIFY before acting)
The droplet service (resolver 0.0.1092, deployed UrlProtocol2.class byte-identical to the published artifact) opens a gossip child stream on the connection it shares with buildtest-server and gives negotiation 1 s (UrlResolver.trySendOverExistingConnection, withTimeout(1000) ~l.18447 at commit 1318ecb6d); on timeout the catch at l.18462-18473 calls conn.close() on the WHOLE shared connection. buildtest-server's PersistentRpcConnection to url://digitalocean-droplets/ rides that connection, so every in-flight listDroplets / getDropletCreationStatus / create reply gets EOF at a frame boundary ("read 0 of 4 bytes") → AmbiguousRpcRequestException. buildtest-server misses the 1 s deadline because its route runs -XX:+UseSerialGC (15.7 % of wall time in 0.3–1.0 s pauses on 09-28). Every droplet-side close in a 120 s capture followed: 50 B + 57 B open → only ACKs → reset + FIN at exactly 1.000 s ±0.04 s. Refuted then: DigitalOcean limits, a connection-manager prune, the 15 s relay keepalive, buildtest-specific targeting. Fixed upstream in UrlResolver PR 1105 + PR 1080 (resolver ≥0.0.1259, DiscoveryConnectionRecovery.retireWhenIdle, reproducers tests/testGossipNegotiationFailurePreservesPendingRpc.kts, tests/testBootstrapExchangeEofKeepsAcceptedRpc.kts); PR 142 moves the droplet service onto it. Restarting does not help (same code, same churn). Note the 10-02 handoff's facade-hop (CN :35000 → 44207) observations are consistent with this: the droplet side is the one closing. Separate follow-up: buildtest-server's old generation grew 1.48 → 1.81 GB in 2 minutes on 09-28 — needs a heap histogram, not GC tuning. Challenges: PlanRepository challenges/2026-09-28-1734-buildtest-droplet-allocation-failures-droplet-service.md (+ PR 9940 extension).
Branches and unmerged state (verified 2026-10-03 against GitHub)
| Repo |
Branch |
Head |
PR |
State |
| UrlResolver |
fix/idle-reaper-head-of-line |
b7a24f48cb |
1146 |
open; bld-build red (test unnamed); remote void |
| UrlResolver |
fix/relay-teardown-close-flake |
f9171707a5 |
1142 |
open; remote 1890/1895 (13 tests with FAILED attempts, unattributed) |
| UrlResolver |
fix/large-rpc-own-connection-flake |
89acb1a200 |
1140 |
open; remote failed on SSH shards (infra) |
| UrlResolver |
wip/ur-rev-1147-rollback-at-every-position-test |
db85fff980 |
none |
review-written position-sweep test for 1147 (merged); never run — fold into 1180/1181's owner or delete |
| BuildTestServerService |
wip/btss-fix2-detach-droplet-rpc-values-handoff-2026-09-28 |
9e81e1d5fe |
none |
2026-09-27 pre-history of #371; superseded by #371/#377 — safe to delete after a glance |
| BuildTestServerService |
fix/detached-failure-keeps-sandbox-type |
f6fc93f447 |
none |
this effort's WIP; superseded by merged #377 (same fix) and #385 (loop guard) — safe to delete |
| UrlResolver |
wip/ur-fix-1141-review1 |
f4ad93d374 |
none |
staging for closed 1141 — safe to delete |
Nothing exists only locally: the 2026-09-28 sweep pushed every clone; the scratchpad (/tmp/claude-1000/-code/45e7c653-2e5b-4771-ac13-37a8d108396d/scratchpad/) holds lane briefs, findings and reports only and is not durable.
Next steps
- Read the 10-02 droplet-RPC handoff; agree one owner for the droplet-side fix. Then deploy PR 142's main (steps and acceptance above), re-capture to confirm the 1.000 s close pattern is gone, and record the result in both handoffs.
- Move 1140 / 1142 / 1146 to the 10-02 UrlResolver handoff (update its "other agents' PRs" note), or, if this handoff keeps them: re-request 1140's remote gate; attribute 1142's 13 FAILED tests against a contemporaneous main run; name 1146's bld-build failure, reconcile it with #1179, re-review round 2, and re-gate.
- Once both are done, complete this handoff with
handoff-cli complete — after first checking each item firsthand (the snapshot above is not evidence).
Reusable / operational knowledge
- Merge-queue enqueue: GraphQL
enqueuePullRequest; eviction reasons via timelineItems(itemTypes:[REMOVED_FROM_MERGE_QUEUE_EVENT]); merge-group checks via the entry's headCommit.oid.
- CN deploys: HCF
write-file (new filename), CN CLI update-route --facade url --domain <name> --image <path>, audit lines via HCF write-file --append; HCF read-file cannot serve the ~9 MB /root/cn-watchdog.audit (tail it over read-only SSH).
- Local box: gate JVM launches on cgroup headroom (
memory.max − memory.current ≥ 3 GB), JAVA_OPTS=-Xmx512m for remote kompile clients, no JAVA_TOOL_OPTIONS for child-JVM tests; one full scripts/test.bash --test . at a time, but a low-priority full suite must not starve targeted runs.
History — the 2026-09-28 19:35 UTC snapshot (kept for the mechanism chain and evidence; superseded where the section above says so)
RE-VERIFY: everything below is a write-time snapshot. Before acting, re-run for each PR: gh pr view <n> --repo CodexCoder21Organization/<repo> --json state,mergeStateStatus,headRefOid,statusCheckRollup,mergedAt; for merge queues: GraphQL repository.mergeQueue.entries and the PR's RemovedFromMergeQueueEvent timeline items; for buildtest runs: read EVERY ledger page of https://buildtest.kotlin.build/api/test-results?id=<run>&page=N&pageSize=300&attempts=true (100-row cap; re-fetch a page that comes back empty; the gate is zero FAILED attempts, and a green check can hide retried FAILED attempts while a red check can carry 0 executed tests). Live services: handoff-cli health, the CN CLI (coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com, coursier = /home/u/bin/cs).
Mission summary
The original request: land the upload-UX work (FiledropWui + PhotoGenerationManagerWui per the GOOD_DESIGN.md File Uploads section, "fixed/upgraded merged and deployed"), the PhotoGenerationManagerWui Library view, and the Template family — TemplateApi, TemplateEmbedded, then TemplateServiceServer at url://templates/ and TemplateCli. Later directives in force: "codex quota reset, drive this to completion", "you are fable, delegate to codex", and since 05:25 UTC on 2026-09-28 "Drive this to completion delegating to opus instead of codex/gpt".
What is DONE (verified 19:20 UTC): upload-UX, Library view, TemplateApi, TemplateEmbedded (0.0.4 published), TemplateServiceServer (PR 1 and the pin-bump PR 2 merged; community.kotlin.templates:template-service-server:0.0.2 and templates-client:0.0.2 published, POMs 200) — 43 PRs merged this session (list below). The handoff service was fixed upstream and deployed (HandoffEmbedded 0.0.21 / HandoffServiceServer 0.0.22, pid 3077541 since 18:01:31Z; report mirrors no longer rewrite state.json).
What REMAINS: (1) TemplateCli PR 2 → closure of its review findings → re-gate → enqueue → publish template-cli:0.0.2 (the last Template-family step; template-cli/0.0.2 POM currently 404). (2) Land the gate-green, final-reviewed UrlResolver PRs 1143 and 1144 and DigitalOceanDropletServiceServer 142; re-review and land 1146, 1147, 1141, 1142; re-gate 1140. (3) BuildTestEmbedded 1129 (buildtest voids runs on ambiguous droplet reads) → closure → publish → BuildTestServerService pin → redeploy url:buildtest:; BuildTestServerService typed-failure fix (branch pushed). (4) Deploy the droplet service after 142 merges (the infrastructure root cause of today's voided runs). (5) The UrlResolver flake queue.
What was found and done (the chain)
Template family
- TemplateServiceServer PR 1 (https://github.com/CodexCoder21Organization/TemplateServiceServer/pull/1, MERGED 17:51:15Z): the
url://templates/ service with a SERVING→CLOSING→CLOSED request lifecycle (exactly one closer owns catalog close; later closers provably wait; callback-thread closes are rejected by the catalog but teardown still completes; the delegated start is the single owner of failure cleanup so an interrupted failed start reports its interrupt once), lexical bind-domain validation url://localhost.tcp:<port>/, borrowed resources closed only after catalog quiescence, fixtures moved to a build rule so runner wall-clock measures the contract. Seven review rounds, every finding closed with a mutant-killing test (MB/ME/MG/MH/MF/MN3/MJ). Gate run 17824f6c 47/47 zero FAILED.
- TemplateServiceServer PR 2 (https://github.com/CodexCoder21Organization/TemplateServiceServer/pull/2, MERGED 18:54:10Z): pins template-embedded 0.0.4 (SFSE test pins 0.0.25 — the version 0.0.4 was verified against; 0.0.26 exists but 0.0.4's POM does not depend on SFSE so nothing resolves it transitively); gate 6b6bd3a5 47/47 zero FAILED; published server/client 0.0.2 from a5ae295c (JAR SHA-256
a7a1dfffd5bd94c832ac849e7cb4707b23bd24d0766ec527403a796cadf25391 / f7b4d360a861444176b3acec6b81a622a05bcde7323982719bdcbcfdf917aaa6). Fact worth knowing: only the server's buildFatJar bundles /templates-client-impl.jar and /stdlib.jar; the published server/client jars do not (same as 0.0.1); embedders relying on loadServerResource defaults need the fat jar; TemplateCli passes explicit jars.
- TemplateCli PR 2 (https://github.com/CodexCoder21Organization/TemplateCli/pull/2, OPEN, head e2cafefae324de8e43d1897f6621e8c2b9988133 = 337ffba + the WIP closure commit "close round-8 review items": S1
tests/testInvalidBindUrlNeverDials.kts (real run() against a counting listener; three invalid spellings → exit 2, full message, zero connections; the correct form connects and exits 4; passed 1/1 locally; its mutation check was NOT run), N3 stray-line note in the child-JVM failure message, N4 README 0.0.2 note + the jar remark moved to a code comment; still open: N1 (body says 13 cases, the test sends 14), N2 (tied-timestamps wording), G1 (fill in the remote result + ledger counts) — see scratchpad/opus-jobs/tcli-bump/out/findings.md "PAUSED STATE"): pins client/server 0.0.2, embedded 0.0.4, SFSE 0.0.25 (test), resolver 0.0.1262 / protocol 0.0.527 / libp2p snapshot-27 (aligned with the client's POM), CLI 0.0.2, test-support 0.0.7; bind-mode .tcp URLs open through the client's openBindModeTemplateCatalog (which owns and closes the connection) instead of a hand-rolled bridge; lexical validation of url://localhost.tcp:<1..65535>/ as a usage error (exit 2) before any dial; the test server's counting decorator implements QuiescentTemplateCatalog (server 0.0.2 accepts only that or FileStoredTemplateCatalog) and tears down close() → awaitQuiescence() → manager close → checkNoOpenFiles(); round-5/7 carry-overs closed (claimkey-fail-keep provider mode with mutant proof; README nits; stdout JSON asserted in both JSON cases; the tied-timestamps terminal-state poll removed on the 0.0.4 awaitQuiescence contract; testHardLinkCapabilityFallback runs 14 probe cases in one child JVM — 17.5 s → 2.2 s). Review (tcli-rev8, scratchpad/opus-jobs/tcli-rev8/out/final.md): READY on the code; before enqueue: S1 should-fix — commit a non-injected run() test proving invalid bind URLs never dial (a counting ServerSocket; invalid forms → exit 2 with 0 accepts; valid form dials); N1 body says 13 cases, the test sends 14; N2 the "control without the barrier" was not barrier-free (the handle's own teardown runs awaitQuiescence(); the fail-first proof is upstream in TemplateEmbedded tests/testAwaitQuiescenceDrainsResultSink.kts); N4 README 0.0.2 note for the exit-2 rule; then fill the PR body's promised remote-suite result + per-attempt ledger counts. The supervisor's read of the source hunks: no findings. First gate run 4a9ccf62 (58 tests) was all-PENDING at 19:20 (buildtest provisioning). Publish template-cli:0.0.2 only after a zero-FAILED ledger on the final head (probe https://kotlin.directory/community/kotlin/templates/template-cli/0.0.2/template-cli-0.0.2.pom first — 404 at write time; publish-maven-artifact skill; --api-url https://api.kotlin.directory/upload).
Handoff service OOM (done, deployed)
- Mechanism (
scratchpad/opus-jobs/hs-oom2/out/final.md): the OOM thread was an RPC request thread in HandoffEmbedded.persistLocked(:2115) ← submitStatusReport(:1131): the entire durable state (840 handoffs, 9.4 MB on disk) serialised into one ~21 MB UTF-16 String on every status-report mirror although reports/claims are runtime-only; git-sync thread idle (not PR 23's path); "metaspace 99 %" was a used/committed ratio, refuted.
- Fix: https://github.com/CodexCoder21Organization/HandoffEmbedded/pull/15 (MERGED; streamed
root.write(bufferedWriter, 2, 0) — byte-identical output proven on a 3.6 MB adversarial document; reports persist only when they attach a conversation or complete; three fail-first tests) published as handoff:embedded:0.0.21; https://github.com/CodexCoder21Organization/HandoffServiceServer/pull/24 (MERGED; pins 0.0.21, server 0.0.22). https://github.com/CodexCoder21Organization/HandoffEmbedded/pull/14 closed as superseded.
- Deployed 18:01Z (
scratchpad/opus-jobs/hs-deploy2/out/final.md): jar /root/ContainerNursery-uploads/jars/handoff-service-server-99ab7220-20260928.jar (SHA-256 e05c1ba1a7499486cb3749380738d1306145dcd424525914f02218a929f9f1cc, built from main 99ab7220ef); route url:handoff image was the only field changed; verified: four real report mirrors succeeded without rewriting state.json, git-sync clean, WUI 200, no heap dump, 0 OOM lines, GC rate halved. Rollback: update-route --facade url --domain handoff --image /root/ContainerNursery-uploads/jars/handoff-service-server-fa8d8756-20260928.jar. Follow-ups: startup still reads state.json into one String (~40–50 MB transient at 9.5 MB and growing); the whole-file state needs an append-only/per-handoff store; -Xmx128m remains tight (old-gen 45–75 of 87 MB); HCF read-file cannot serve the 9 MB audit log (tail via read-only SSH). ServiceAtlas 53 (deployment history) MERGED.
Buildtest voided runs — the day's infrastructure defect, two layers + the root cause
- Symptom: merge-group entries and gate runs marked FAILED with zero tests executed (85143156, 6358e195, 151161ae, 0ed8aaf1, PR 142's 6706985d/c223c170, 1144's 7f8162ff, …): "Could not verify droplet cleanup before requeueing …
AmbiguousRpcRequestException: RPC 'listDroplets' to 'digitalocean-droplets' … transport failed after the request write began". Four evictions of UrlResolver 1143 and one of TemplateEmbedded 6 today were this. Challenges: 2026-09-28-1301, -1619, -1734 (bt-alloc mechanism), PlanRepository PRs 9936/9940 merged.
- Layer A — consumer (bt-alloc → bt-fix): BuildTestEmbedded's
isSandboxedAmbiguousRpcTransportFailure (ProviderRpcFailureClassification.kt:61-62) required a top-level SandboxException but production delivers RuntimeException("Sandboxed code threw an exception: …"), so PR 1087's bounded replacement read never ran. Fix: https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1129 (OPEN, head 42260b6ce): classifier accepts either shape by exact message matching; production-shape fail-first test through a real sandboxed service (main: production's exact error text; fix: PASS); guard test that non-ambiguous failures still fail the run; sibling flake twoAmbiguousCleanupListingsThenComplete root-caused from the run record (4.4 s analysis + a 5 s reconciler floor exceeded its 10 s latch) and fixed with the siblings' dropletReconcilerIntervalMs = 100L (no bound changed). Deviation: the tests replicate PR 371's detachedForCaller verbatim because no published server artifact contains it — switch to the real artifact once published. Review (bt-rev-1129, interim, scratchpad/opus-jobs/bt-rev-1129/out/final.md): NOT-READY with no must-fix — should-fix: a production-shape test for the provisioning getDroplet path (detachedAmbiguousGetDropletKeepsRecordedDroplet, must fail on main); state the publish/deploy chain in the PR body; file the upstream follow-up; nits. The 10× flake-fix run and the sibling re-run were not completed. Note: getDropletCreationStatus does NOT share the classifier (it has its own text-based isAmbiguousCreateResponseFailure, DropletManager.kt:1143) — the brief's premise was refuted; the PR body is correct.
- Layer B — the type-loss site (bts-detach): BuildTestServerService
src/buildtest/server/BoundedDropletService.kt:234 Throwable.detachedForCaller() (PR https://github.com/CodexCoder21Organization/BuildTestServerService/pull/371, 2026-09-27) re-creates a SandboxException as a plain RuntimeException, applied at :594 — live in production (buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar, pid 3126087). Branch https://github.com/CodexCoder21Organization/BuildTestServerService/tree/fix/detached-failure-keeps-sandbox-type (head f6fc93f): two reproducers (tests/sandboxFailureKeepsItsTypeAcrossSandboxGenerationRetirement.kts, tests/sandboxFailureInsideCauseGraphKeepsItsType.kts) and a WIP fix, not yet compiled or run (re-create SandboxException via its public (String, Throwable) constructor after detaching the cause graph; drop a cause edge that loops back). Invariant: a detached failure holds no delegate references AND keeps the type, message, cause messages, suppressed failures and stack trace. Main-side verdict obtained after the pause (run c38ab61f): sandboxFailureKeepsItsTypeAcrossSandboxGenerationRetirement FAILS on main as intended (caller receives java.lang.RuntimeException where a SandboxException is required — the fail-first evidence); sandboxFailureInsideCauseGraphKeepsItsType hung (30 s timeout inside isSandboxConnectionInvalidation, BoundedDropletService.kt:780: the cause-chain walk has no loop guard, so a failure whose cause chain loops back hangs the caller thread forever — a second defect; the WIP fix only breaks loops through a SandboxException). Remaining: split that test into nested-type and loop files, guard the walk (a loop of plain exceptions must not hang), run the new tests + #371's tests on the fix, full suite, PR, publish recommendation. Follow-up: SandboxConnection{Open,Generation}DiscardedException are flattened by the same function, making the type check at :648 unreachable on that path (reachability unverified).
- Root cause — the droplet service's connection churn (ds-churn,
scratchpad/opus-jobs/ds-churn/out/final.md): the droplet service runs resolver 0.0.1092, whose gossip path closes the WHOLE shared connection when one child stream is not negotiated within 1 s (UrlResolver.kt:18462-18473 at 1318ecb6d; deployed UrlProtocol2.class byte-identical to the published artifact); buildtest-server trips it because its route runs -XX:+UseSerialGC (15.7 % of wall time in 0.3–1.0 s stop-the-world pauses; 7,388 s GC in 44,911 s uptime). Two packet captures: every droplet-side close follows "50 B + 57 B child-stream open → only ACKs → reset + FIN at exactly 1.000 s". Not a prune, not the relay keepalive, not buildtest-specific (third most-churned peer). Fixed upstream in UrlResolver PR 1105 + 1080 (resolver ≥0.0.1259, DiscoveryConnectionRecovery.retireWhenIdle, existing reproducers). Restarting does not help. Deploy vehicle: https://github.com/CodexCoder21Organization/DigitalOceanDropletServiceServer/pull/142 (OPEN, rebased head b14991d128, resolver 0.0.1259/protocol 0.0.527/SJVM 0.0.51/observable 0.3.22/libp2p snapshot-27, version 0.0.85, no src/; full local suite 274/274; Actions green; kotlin.build (remote) run 9d3ebb38 executing at 19:20 with 102 PASSED / 0 FAILED so far). Supervisor's read: clean. Follow-up: buildtest-server's old generation grew 1.48→1.81 GB in 2 minutes — needs a heap histogram, not GC tuning.
- Newly observed 19:13 (unfiled): 1143's fourth merge-group run 646ef02e passed 1845/1845 yet the run was FAILED because buildtest "Could not persist build run … run.json after one quick retry … Attempt history checkpoint …" — a buildtest server-side persistence defect distinct from the droplet churn; file a challenge and investigate (BuildTestServerService/BuildTestEmbedded run persistence).
UrlResolver flake campaign (each item = one lane, one PR)
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1143 (OPEN, head bcbe83d685, harness-only: the buffered PeerExchange test's gate holds the response until the responder's half-close request; reviewed READY; supervisor's read clean; all branch checks green). Evicted four times — 16:31 and 17:17 zero-test provisioning voids; 18:27 merge-group
bld-build (no retries) failed on testLiveProjectionResilienceHighRevisionLineageRecovers (a flake in the live-projection family, unrelated); 19:13 the run.json persistence failure above with 1845/1845 passed. Re-enqueue (GraphQL enqueuePullRequest) after reading the eviction reason each time. Residual 1-in-25 cold failure ("exact catalog peer not applied") has a follow-up spec in scratchpad/opus-jobs/ur-rev-1143/out/final.md item 4.
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1144 (OPEN, head 17faeda051, product: the retired-parent replacement dial honours the opener's own declared budget instead of a fixed internal 1 s window; no constant changed — reviewer's table in
scratchpad/opus-jobs/ur-rev-1144/out/final.md; round-1 tests added: cancel-during-held-handshake and upper-bound, both mutant-proven). The supervisor disclosed the "budget pass-through is a defect fix, not a timeout increase" judgement to the user in every report from 14:45 on; no objection was received. Gate: bld-build SUCCESS; kotlin.build (remote) run 0cecfa74 SUCCESS (1849 rows; 100 PASSED on page 1 at 19:20 — read all 19 pages before enqueue). Supervisor's read of the delta: clean. Follow-up: a losing/timed-out attempt's recovery dial runs to its declared window holding one of 24 admission permits because connectWithRelayFallback never cancels its promise.
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1146 (OPEN, round-2 head 2561552fc8, idle-reaper head-of-line fix: each blocking idle close runs on its own thread;
markClosed() joins in-flight idle closes; long calls re-arm for a full debounce; failure-time diagnostic in the flaky idle-teardown test). Round-1 review NOT-READY (scratchpad/opus-jobs/ur-rev-1146/out/final.md) → round 2 pushed with the decided invariant (closed means closed; idle closes never delay each other; a hung close pins only its own generation and one named thread), the #1106 test rewritten (20/20 consecutive), two fail-first join tests. Needs a re-review of round 2. Its full local suite was killed unfinished by the pause (the lane was stopped at 19:30 without writing a paused-state section; partial log at scratchpad/opus-jobs/ur-reaper/out/runs/full.log) — rerun one full suite on the final head; gate run ea8b0675 had not provisioned at 19:19. The underlying flake (testIdleSandboxTeardownDoesNotDelaySharedClockCallbacks, Actions run 36422654305) has NO established mechanism (ur-flake6 stopped honestly: JVM stall vs clock step vs in-JVM stall undiscriminated; the new diagnostic discriminates them next time).
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1147 (OPEN, head d7165e277, live-projection partial seed: listeners marked attached only after every property is seeded, rollback + rethrow on failure; fail-first in run 44a4730d, pass in 6388f74f, 112/112 projection tests in 51ecd7e2; gate ccb9286d 1846/1846 zero FAILED attempts — supervisor-read, all 19 pages;
bld-build in progress). Supervisor's read of the source: clean. Review (ur-rev-1147, INTERIM, scratchpad/opus-jobs/ur-rev-1147/out/final.md): NOT-READY with small bounded fixes — #2 move the addChangeListener loop inside the try (a throw at index k otherwise leaves listeners 0..k-1 attached forever with listenersAttached=false); #4 test gap: only the last property is failed, once — add the position-sweep test (draft pushed to wip/ur-rev-1147-rollback-at-every-position-test, db85fff9, never run: fail each of title/heartbeat/todos twice, mutate between, assert the retry snapshot, a POLL delta, and a fresh SUBSCRIBE after UNSUBSCRIBE); #5 correction to the supervisor's gate read: the ledger's final statuses were all PASSED, but /api/test-attempts?id=ccb9286d shows one retried FAILED attempt (stressTestConcurrentRegistrationsReliablyPropagateToRelay, 63 s, unrelated → flake queue), so "zero FAILED attempts" was wrong — read api/test-attempts as well as api/test-results; #7 on attach failure handleSubscribe has already registered the token (push tokens are not lease-managed and get no close listener → live forever; listeners never detach; fix: freeSubscriptionLocked(token) under monitor before rethrowing, with a fail-first push test — fix here or in a tracked lane); #8 retired-token hole confirmed at ProjectionServiceHandler.kt:830-835 (own lane; fail-first spec in the verdict); #9 a latent, permanent lock-order deadlock exists (SUBSCRIBE holds monitor while calculateCurrentValue awaits another thread's publication drain, whose finalizePublication → markStale → onInvalidated blocks on monitor) but is refuted as the CI trigger — own lane with the invariant "no observable read that can wait on another thread's drain may run under monitor"; #10 a third trigger candidate: UrlResolver.kt:15173-15190 catches only Exception, so a java.lang.Error also escapes with no response; #11 PR body must list #7/#8/#10. Required before READY: #4 run green, #11, a decision on #7; #2 is a one-line move.
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1142 (OPEN, head 7d09d9c4b0, product: a returning provider supersedes a withdrawal tombstone by strict timestamp; review NOT-READY on a test gap; the strengthened scenario (a) — relay
registrationTimeoutMs = 10_000 plus periodic committed-catalog refreshes during the hold — passes on head 15/15 + 5/5 with zero FAILED across 19 pages; mutant evidence complete: run 4baba76a fails 11/11 on the mutant (the literal review scenario could not discriminate because onServiceWithdrawal also records the withdrawal in RelayService, which then omits the service from the returning registration — the strengthened scenario waits for the relay's 10 s withdrawal memory to expire while the catalog's 60 s tombstone is live and refreshes the committed catalog every 500 ms during the hold; scenarios (b)/(c) included). Gate: kotlin.build (remote) run 1258645b reported 1850 passed / 0 failed by the runs API — its ledger pages are not yet read; bld-build SUCCESS. See "PAUSED STATE 2" in scratchpad/opus-jobs/ur-fix-1142-1141/out/findings.md for the exact ledger commands. Needs a re-review of the new test, then enqueue.
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1141 (OPEN, head 431f87b7a4, harness:
addPeer after the live connection, zero-relay-connections assertion, latches sized for two attempts; pre-fix 0/10 copies pass, fixed 10/10). Checks pending at 19:20. Needs a re-review, then enqueue.
- https://github.com/CodexCoder21Organization/UrlResolver/pull/1140 (OPEN, head 03471cce85, reviewed READY;
bld-build red on unrelated flakes) — re-gate after the flake PRs land.
- Flake queue (each needs a fix-flakey-test lane):
testLiveProjectionResilienceHighRevisionLineageRecovers, stressTestSynchronousPeerExchangeServiceDiscovery (<500 ms average bound, load-sensitive), stressTestActiveQueryMalformedResponseAlwaysReachesParsing, testPruneRetainsPeerWhenTransportCompletesAfterCallerTimeout, stressTestFreshJoinWithdrawalDeliveryUnderConcurrentReannounce, stressTestWithdrawalLossDuringInitialBootstrapHandshake, testStressPeerExchangeRaceCondition, testEagerJoinHostCreationCompletesWithOccupiedCarrier, testQuietParameterObservableReconcilesAfterRequestRedials, the idle-teardown flake (mechanism unknown), the closed-resolver retention item (scratchpad/opus-jobs/ur-retain/p.md), plus the 1143 residual. UrlResolver's required bld-build runs 1845 tests with no retries, so every landing there is a draw against this population until these land.
TemplateEmbedded
- https://github.com/CodexCoder21Organization/TemplateEmbedded/pull/6 MERGED 18:41Z (admission-test hardening; bounded teardown naming unfinished iterations; atomic counters; quiescence-gate arming proven by two mutants). Follow-ups: three pre-existing near-budget tests (
testTemplatePaginationReadsOnlyRequestedPage, testRealLambdaStackStringFixture, testTemplatePointStorageCostIsConstant, 22–26 s of a 30 s budget) and the two deadline tests' per-iteration off-CPU time (unattributed) need their own fix-flakey-test passes; kompile-cli 0.0.105 --remote fails the buildtest health RPC in this repo (challenge 2026-09-28-1615; 0.0.113 works).
Relevant PRs / refs
Merged this session (43): TemplateServiceServer 2 (18:54), TemplateEmbedded 6 (18:41), ServiceAtlas 53 (18:16), HandoffServiceServer 24 (17:53), TemplateServiceServer 1 (17:51), HandoffEmbedded 15 (17:51), SimpleFileSystemServiceEmbedded 26 (16:16), ServiceAtlas 52 (14:29), TemplateEmbedded 5 (14:00), GithubProxyServerService 48 (13:48), SimpleFileSystemServiceEmbedded 25 (11:39), ServiceAtlas 51 (11:23), HandoffServiceServer 23 (10:40), UrlResolver 1132 (08:47), TemplateCli 1 (08:25), ServiceAtlas 50 (07:54), TemplateEmbedded 4/3/2/1, SimpleFileSystemServiceEmbedded 21/20/19, kotlin-build-ci 290/288, ServiceAtlas 49/48/47/46/45, FiledropWui 27/26, BuildTestEmbedded 1106, PhotoGenerationManagerWui 12/13, DummyModelServiceServer 1, BuildTestServerService 371/355, TemplateApi 1, DocumentationRepository 225, FiledropApi 4, FiledropEmbedded 12, FiledropServiceServer 24 — all under https://github.com/CodexCoder21Organization/<repo>/pull/<n>. Plus PlanRepository challenge PRs 9936, 9940 and the challenge files named above.
Open PRs (state at 19:20 UTC): see the per-PR bullets above for TemplateCli 2, UrlResolver 1143/1144/1146/1147/1142/1141/1140, DigitalOceanDropletServiceServer 142, BuildTestEmbedded 1129.
Branches and unmerged state (every one verified against git ls-remote; the local scratchpad is disposable):
| Repo |
Branch |
Remote head |
PR |
What is on it |
State |
| CodexCoder21Organization/TemplateCli |
chore/pins-server-0-0-2 |
e2cafefae324de8e43d1897f6621e8c2b9988133 (WIP closure on top of 337ffba) |
https://github.com/CodexCoder21Organization/TemplateCli/pull/2 |
pins + bind-mode via client + carry-overs + S1/N3/N4 closure |
reviewed READY on code at 337ffba; N1/N2/G1 body edits pending; S1 mutation check not run; gate on e2cafef unread (a new run will have started) |
| CodexCoder21Organization/UrlResolver |
fix/bootstrap-buffered-response-flake |
bcbe83d685 |
1143 |
harness fix |
final-reviewed; evicted 4×; re-enqueue |
| CodexCoder21Organization/UrlResolver |
fix/cancelled-discovery-late-replacement-flake |
17faeda051 |
1144 |
budget pass-through + 2 review tests |
gates green at 19:20; read 0cecfa74's 19 pages; enqueue |
| CodexCoder21Organization/UrlResolver |
fix/idle-reaper-head-of-line |
2561552fc8 |
1146 |
round 2 |
needs re-review + full-suite result |
| CodexCoder21Organization/UrlResolver |
fix/live-projection-poll-fallback-npe |
d7165e277d |
1147 |
partial-seed fix |
gate 1846/1846 zero FAILED; review interim; enqueue after review |
| CodexCoder21Organization/UrlResolver |
wip/ur-rev-1147-rollback-at-every-position-test |
db85fff980dcfef990619b45d9c6fefe35049fdf |
no PR |
review lane's rollback-at-every-position test, on top of d7165e277 |
written, never run |
| CodexCoder21Organization/UrlResolver |
fix/relay-teardown-close-flake |
7d09d9c4b0 |
1142 |
strengthened timestamp test + comment rewording |
needs round-4 mutant read + re-review |
| CodexCoder21Organization/UrlResolver |
fix/active-query-read-failure-flake |
431f87b7a4 |
1141 |
harness fix round 1 |
needs re-review; wip/ur-fix-1141-review1 (f4ad93d374) is its superseded staging branch |
| CodexCoder21Organization/UrlResolver |
fix/large-rpc-own-connection-flake |
03471cce85 |
1140 |
harness fix |
reviewed READY; re-gate |
| CodexCoder21Organization/DigitalOceanDropletServiceServer |
upgrade/url-stack-1259-527-sjvm-0.0.51 |
b14991d128 |
142 |
resolver stack upgrade, 0.0.85 |
274/274 local; gate 9d3ebb38 executing; enqueue on a clean ledger; then deploy |
| CodexCoder21Organization/BuildTestEmbedded |
fix/ambiguous-droplet-read-classifier |
42260b6ce9 |
1129 |
classifier fix + tests + flake fix |
review interim NOT-READY (should-fixes); shards running |
| CodexCoder21Organization/BuildTestServerService |
fix/detached-failure-keeps-sandbox-type |
f6fc93f447 |
no PR |
2 reproducers + WIP fix |
fix not compiled/run |
| CodexCoder21Organization/BuildTestServerService |
wip/btss-fix2-detach-droplet-rpc-values-handoff-2026-09-28 |
9e81e1d5fe2a2ada500ce36fad9b6e0c56845044 |
no PR |
7 commits from an earlier lane (2026-09-27) found unpushed in codex-jobs/workspace/btss-fix2: "Detach droplet RPC values from sandbox lifetime" + tests — likely the pre-history of PR 371 |
pushed for safety; may be superseded by #371 — compare before using |
| CodexCoder21Organization/TemplateServiceServer |
chore/template-embedded-0-0-4 |
a5ae295c93 |
2 (MERGED) |
— |
merged |
Deliberately not pushed (throwaway experiments, not repository material): mutant copies and diagnostic probes under scratchpad/codex-jobs/workspace/{sfse-fix2-m1..m5, sfse-rev-tests-m1..m3, tss-flake-diag, tss-flake2, tss-flake3-diag, tss-rev5-adv, ur-fix-1142-mut} — each is a mutated src/ or a testDiag*/testExp* file used to prove a review finding; the proofs are recorded in the corresponding scratchpad/opus-jobs/<lane>/out/final.md. The scratchpad itself (/tmp/claude-1000/-code/45e7c653-2e5b-4771-ac13-37a8d108396d/scratchpad/) holds every lane's brief (opus-jobs/<lane>/p.md), findings and final reports; it is not durable.
Deployed / published but not merged: none at write time — every deployed jar (handoff-service-server 99ab7220; github-proxy-server 091c8183; buildtest-wui e4b472be, earlier today) was built from a merged main, and every published artifact (SFSE 0.0.25/0.0.26, TemplateEmbedded 0.0.4, handoff:embedded 0.0.21, template-service-server/templates-client 0.0.2) comes from a merged head. The droplet service still runs droplet-service-server-0.0.82-r-derivation-fix-20260827T2205Z.jar (resolver 0.0.1092 — the churn) and buildtest-server runs buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar (filename date stale; contents = server main ≥ f0fb90a with embedded ~0.0.69273972; no audit record of its 05:10Z replacement — provenance should be recorded by whoever deployed it).
Next steps (in order; each is one lane per item, work-queue style, 6 workers max)
- TemplateCli 2: read the live head; if the tcli-bump closure (S1 no-dial
run() test, N1/N2/N4 text) is not pushed, do it; gate → read every ledger page → supervisor read → enqueue → after merge, publish template-cli:0.0.2 (probe the POM first) → the Template family is complete.
- Re-enqueue 1143 after reading the last eviction reason (
RemovedFromMergeQueueEvent + the merge-group commit's check-runs); enqueue 1144 on a clean read of run 0cecfa74's 19 pages (my final review is done); enqueue 142 on a clean read of 9d3ebb38's 3 pages (my read is done), then a deploy lane for the droplet service (route url:digitalocean-droplets:, jar via dropletserviceserver.buildFatJar, upload under a new filename with HCF write-file, update-route --image, keep memoryLimitMb 1024, rollback jar above; post-deploy: the 1.000 s droplet-side FIN pattern must be gone and AmbiguousRpcRequestException must stop appearing in buildtest-server stdout).
- 1147: finish the review (interim in
opus-jobs/ur-rev-1147/out/), then enqueue on its clean gate. 1146: re-review round 2 (invariant, markClosed join, 20/20) + read its full-suite result. 1142: read the round-4 mutant ledger (opus-jobs/ur-fix-1142-1141/out/findings.md "PAUSED STATE") — the test must fail the mutant near 100 % — then re-review. 1141: re-review. 1140: re-gate after the flake PRs land.
- BuildTestEmbedded 1129: close bt-rev-1129's should-fixes (provisioning-path production-shape test; publish/deploy chain in the body; file the upstream follow-up), 10× the flake fix, re-review → enqueue → publish
buildtest.embedded:buildtest-embedded (next version above 0.0.69273972; probe the POM) → bump BuildTestServerService's pin (build.kts:338) → build buildtest.server.buildFatJar from server main (contains #371) → deploy url:buildtest: under a new filename, keep the hotfix jar for rollback, audit lines → re-request the voided runs. bts-detach: finish (branch f6fc93f) → PR → publish → include in the same server deploy. Also file the new buildtest run.json persistence failure (run 646ef02e) as a challenge and investigate.
- Work the UrlResolver flake queue and the TemplateEmbedded near-budget follow-ups, one lane per item.
Reusable / operational knowledge (learned today)
- Buildtest ledger:
api/test-results caps at 100 rows/page regardless of pageSize; pages sometimes return empty while "still building" — re-fetch; complete lags the GitHub check; the gate is zero FAILED attempts on every page. A zero-test FAILURE (provisioning void, run.json persistence failure) is infrastructure: re-request the check (gh api -X POST repos/<o>/<r>/check-runs/<id>/rerequest) and say so; never treat it as a verdict. Merge-queue enqueue: GraphQL enqueuePullRequest; eviction reasons via timelineItems(itemTypes:[REMOVED_FROM_MERGE_QUEUE_EVENT]); the merge-group commit's checks via its oid from mergeQueue.entries.
- Local machine: cgroup headroom =
memory.max − memory.current (≥3 GB before a JVM); JAVA_OPTS=-Xmx512m for remote kompile clients; do NOT export JAVA_TOOL_OPTIONS for child-JVM tests (the child prints a banner into its merged output); only ONE full scripts/test.bash --test . machine-wide — but targeted tests in a second bounded JVM are fine, and a low-priority full suite must never starve the critical path. jars/coursier → /home/u/bin/cs if a repo's downloaded launcher crashes (never commit). kompile-cli 0.0.105 --remote fails the buildtest health RPC; 0.0.112+ compile local src with --local.
- Deploy playbook (CN): HCF
write-file the jar to /root/ContainerNursery-uploads/jars/<name>-<sha>-<date>.jar (never upload-jar, it overwrites in place), CN CLI ≥0.0.20 update-route --facade url --domain <name> --image <path> (only the image field), append DEPLOY-INTENT/END to /root/cn-watchdog.audit (HCF write-file --append; read-file cannot serve the 9 MB log — tail it over read-only SSH), verify the absence of undesired behaviour (first WUI frame, the exact operation that used to fail, no new heap dumps). Route env for the droplet service contains an inline private key — do not copy it; recommend moving it out.
- Opus delegate lanes hit a shared session limit twice today (08:10, 14:52; ~2 h spacing, reset within ~30–70 min); resume killed lanes in place with
SendMessage and a named resume point ("your tree is at X with Y uncommitted; continue from step Z") — context survives.
- Review discipline that paid off: every review round that named a surviving mutant closed in one pass; "no timeout increased" is shown as a table of every constant and every caller against its declared budget; a control that still contains the barrier proves nothing; read the arithmetic in a run record before reproducing.