← Handoffs

Repository · handoffs

Make the UrlProtocol main branch reliably green: land the reviewed fix PRs once the build service provisions again, then finish the flake queue

View on GitHub ↗

id: hf-2026-09-20-make-urlprotocol-main-reliably-green-land-the-launcher-pin-then-merge-the-six-reviewed-prs-and-fix-the-last-flake url: url://handoff/handoffs/hf-2026-09-20-make-urlprotocol-main-reliably-green-land-the-launcher-pin-then-merge-the-six-reviewed-prs-and-fix-the-last-flake title: Make the UrlProtocol main branch reliably green: land the ten reviewed fix PRs once the build service provisions again, then finish the flake queue summary: Once the shared build service provisions droplets again, no-bump-rebase and re-enqueue 626, then 629, 634, 598, 603, 611, 621, re-request and enqueue test-only 630 and 633, fix 636's Actions failures and re-review 636/634/621, then work the flake queue (root-close fix from a 1/1 diagnostic, finite-backpressure, relay-withdrawal, Netty frame deadline, listener-failure fix) and get the operator's two decisions (correct the main-branch refusal test; environment-bound timeouts via the kompile-core compiler cap). Main is unchanged at 78deea8c; every required remote check is red on provisioning since 17:50 UTC; all partial work is on wip branches and the PlanRepository artifacts branch listed in the body. created: 2026-09-20T22:32:03.445Z completed: null blocked-reason: BLOCKED-EXCLUDED: cold shared-cache preparation for UrlProtocol launcher pin misses its unchanged deadline; owning BuildTest/kompile preparation investigation is excluded. dependencies:

  • url://handoff/handoffs/hf-2026-09-20-resolve-the-fatal-admission-test-contract-before-finishing-urlprotocol-drain-reconciliation

Make the UrlProtocol main branch reliably green: land the reviewed fix PRs once the build service provisions again, then finish the flake queue

Written 2026-10-02 ~21:45 UTC by session fable-code6d-2026-10-02 (explicit stop by the operator). RE-VERIFY before acting: everything below is a write-time snapshot (PR states and checks verified at 2026-10-02 21:40 UTC). Re-check with gh pr list --repo CodexCoder21Organization/UrlProtocol --state open --json number,headRefOid,mergeStateStatus,statusCheckRollup, gh api repos/CodexCoder21Organization/UrlProtocol/commits/main, the merge queue GraphQL query, and https://buildtest.kotlin.build/api/health. Previous campaign state (2026-09-20 to 2026-10-02 morning) is in the earlier versions of this handoff's reports (handoff-cli reports <id>) and in the artifacts branch below.

Mission (operator's /goal, verbatim)

"ensure the main branch of https://github.com/CodexCoder21Organization/UrlProtocol is reliably green. Also triage the open pull requests, decide which are bad/obsolsete and close them or which are good/salvagable and fix/merge/salvage the good parts and drive to completion. Parallelize to the extent possible and race PRs."

Where it stands (verified)

  • main = 78deea8c61311e47596dad9c99b72763292e70e8 (unchanged since 2026-10-01; nothing landed today). Merge queue: PR622:AWAITING_CHECKS. Two entries (PR 622 by a sibling session and PR 626 by this one) were removed by the merge-queue bot at 20:10 UTC after 622's queue run exceeded the build coordinator's 180-minute deadline.
  • The shared build service (kotlin.build (remote), the required check) cannot currently finish or even start runs. Since ~17:50 UTC every UrlProtocol head fails before tests run with "Provisioning failed before tests ran: Provisioning deadline exceeded for runId=... currentAttempt=4" (runs 3d6c934b/629, cbbb7a50/633, f550dd4d/636, 3a562f7b/618); run 98bf14af (634) lost all shards to "Provisioning stopped ... method=createBuildDroplet" after 2241 passes; run d9c01208 (621) "Upload of build failed after 5 attempts". Earlier (14:49-16:30 admissions) runs ran but exceeded the 180-minute overall deadline while TESTING at 94-97 percent (fleet oversubscribed: 38 TESTING runs at 17:10; 120 active droplets at 21:24). Challenge filed: https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-02-1812-remote-check-runs-on-urlprotocol-exceed-the-ci-coordinator.md (and the morning one: https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-02-1040-build-service-infrastructure-not-test-flakes-is-the-main.md). A read-only investigation lane was started at 21:27 and stopped 15 minutes later with no conclusion (findings in the artifacts branch: findings/invprov-findings.md). Nothing can land until provisioning works; do not re-request checks into a saturated or non-provisioning fleet (each re-request burns ten droplets for three hours).
  • Decision: no per-PR coordinate bumps. Every PR keeps main's coordinates = "foundation.url:protocol:0.0.569" line (see REBASE-TEMPLATE.md in the artifacts branch); one publish PR bumps the coordinate after the batch lands (precedent https://github.com/CodexCoder21Organization/UrlProtocol/pull/625). PRs that still carry a bump (636 at 0.0.602, 634 at 0.0.595, 621 at 0.0.592, 629 at 0.0.598, 598 at 0.0.596, 616 at 0.0.597) get it dropped at their no-bump rebase.
  • Two decisions are open with the operator (asked in the status reports, unanswered): (1) may the main-branch test testForwardedOutputRefusalPreservesAtomicBoundary be corrected? Its final clean-write assertion contradicts the refusal behaviour (a refusal closes the incomplete backend source, so the backend's remaining write gets Broken pipe); reproducer testForwardedOutputRefusalBeforeBackendTailWrite fails 3/3 on PR 616 and 3/3 on main (lane f616b). Recommendation A: fix the fixture in its own PR, then 616 lands behind it. (2) The remaining 30-second timeout families (candidate-prepare stress, materialization budget, peer-exchange admission, TCP-server-runs-every-request, large-relay-response, aggregate-refusal-throttle) are environment-bound: droplet CPU sharing after kompile-core PR 319 dropped the C1 compiler cap for test JVMs (memory note; partly mitigated by kompile-core PR 350). Recommendation A: restore the test-JVM compiler cap / thresholds in kompile-core; then B: partition the two heaviest tests. Investigations that refuted code-level hypotheses with numbers: lanes fdial, fdial2, f4, f5 (findings in the artifacts branch).

Open PRs (verified 2026-10-02 21:40 UTC)

PR branch head author merge state required checks
637 wip/rpc-stream-reader-ownership 760676789f753cc4278ab5482c76ae13f471a751 codexcoder21 BLOCKED bld-all-tests=SUCCESS
636 fix/relay-keepalive-on-injected-clock c7dde8b8f607da7a77b0cd22269ad0b9461220f9 codexcoder21 BLOCKED bld-all-tests=FAILURE, kotlin.build (remote)=FAILURE
635 fix/connection-closed-record-without-owner-scope 1c78aeabd156785603a297d59baf09ba83a4fb42 codexcoder21 BLOCKED bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE
634 fix/pending-relay-registration-recovery-flake e2eb21efa71425d1542a9eef8ebf76268c1572cf codexcoder21 BLOCKED bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE
633 fix/peer-registry-removal-batch-fixture-race bc74f12424061d784046f34aa4fbbc07871d0478 codexcoder21 BLOCKED bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE
632 fix/Lexp2-conditional-evidence-removal 50a754f2ef186f6adfb3d33379a56c88783fe2f7 codexcoder21 BLOCKED bld-all-tests=, kotlin.build (remote)=
631 fix/request-write-final-chunk-phase de7007d161465fc859af823f7da7093eaf689961 codexcoder21 BLOCKED bld-all-tests=FAILURE, kotlin.build (remote)=FAILURE
630 fix/root-close-reconciliation-tick-flake 465049c14b3214d9f2296eca64717a7976cacc84 codexcoder21 BLOCKED bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE
629 fix/mdns-hostinfo-startup-stall 449425ed4ede1ae4943d918d8055921448fb82fb codexcoder21 BLOCKED bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE
628 fix/loopback-dial-timeout-flake afc84f6460a0f694694af4207267246da6253f3a codexcoder21 BLOCKED (draft) bld-all-tests=FAILURE, kotlin.build (remote)=FAILURE
626 fix/tcp-listener-release-and-effect-delivery 452ecde9196e2f0bc19645b97d111139463b5c55 codexcoder21 UNSTABLE bld-all-tests=SUCCESS, kotlin.build (remote)=SUCCESS
622 wip/sweep4-P3c-preadmission-close-signal 8e83a9f943a2c8c528ca03ad94f17ec9b29e8995 codexcoder21 UNSTABLE bld-all-tests=SUCCESS, kotlin.build (remote)=SUCCESS
621 fix/libp2p-snapshot-28-pin dc832c02c8f8b3a07fa0fb8c425e2120557c23ea codexcoder21 BLOCKED bld-all-tests=, kotlin.build (remote)=FAILURE
618 fix/relay-gossip-large-frame-child-flow-control 9d2057d6058884c9ddce68158a20210e09c61256 codexcoder21 BLOCKED bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE
617 fix/peer-registry-eviction-off-caller-thread b33bbb5b6d2a66d9881f94eee33510609c79e980 codexcoder21 DIRTY (draft) kotlin.build (remote)=FAILURE
616 fix/parse-budget-not-held-while-receiving 32bda7b35c3998e246e017da14f60f22fb928f01 codexcoder21 BLOCKED bld-all-tests=FAILURE, kotlin.build (remote)=FAILURE
611 fix/stall-retirement-requires-no-progress b30978b8ec5cd97486872c9f4cce2f7f0366a01d codexcoder21 BLOCKED bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE
604 fix/close-drains-active-sends ca45f8b4cd4b2c29aaa3d3d59c318fc74551dd24 codexcoder21 BLOCKED bld-all-tests=SUCCESS, kotlin.build (remote)=
603 wip/flake-admission-manual-read-2026-09-20 581ee4c28bdddbe7d59dba7facf86743fbbb3d6c codexcoder21 BLOCKED bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE
598 wip/FL3c-flake-2026-09-14 5e59b3df2564f1c4dd27d10e22093e4b73275d51 codexcoder21 BLOCKED bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE
594 wip/FL4-flake-2026-09-14 d53ef33f58ad7f54e59e3dba136855e89cb41908 codexcoder21 UNSTABLE bld-all-tests=SUCCESS, kotlin.build (remote)=SUCCESS
574 wip/r71-tcp-serving-lifecycle 2b8e8c305c5fb1f6b3fda990757350ef62b42a35 codexcoder21 BLOCKED bld-all-tests=, kotlin.build (remote)=

Ownership and next action per PR of this campaign (all authored by codexcoder21 = this session's delegates unless noted):

  • 626 (TcpServer lifecycle rewrite + effect-delivery deferral): final-reviewed by me, two delegate review rounds clean; was enqueued, evicted 20:10. Next: no-bump rebase (REBASE-TEMPLATE.md), re-enqueue when the service provisions.
  • 636 (relay keepalive ticks and pong deadlines on the injected Clock; re-proposal of closed 585): opened 18:56 from branch fix/relay-keepalive-on-injected-clock at c7dde8b8; local proof 34/34; Actions suite red on this head: testCurrentGenerationFailedRelayRegistrationSchedulesNextReconnect and testCurrentGenerationThrowingRelayRegistrationSchedulesNextReconnect ("The initial registration must schedule exactly one reconnect check") - plausibly caused by the keepalive tick now sharing the Clock those tests count - plus testPendingRelayRegistrationPreservesSameRelayFailuresUntilAcknowledged (main's flake, fixed by 634). Full log: artifacts branch actions-logs/actions-636-110981530691.log. A fix lane (brief p-f636.md) ran 10 minutes and was stopped; nothing pushed. My own read of the source diff found nothing blocking; two points for the delegate review (brief p-rv636.md): a possible idle-tick leak if a connection is replaced without detaching, and the pong-timeout expression's readability. Next: fix the Actions failures (f636 brief), then the delegate review, then my final review.
  • 634 (pending relay registration recovery): review MUST-FIX closed by e2eb21ef (failure ownership captured under the registration monitor; proof 28/28, 10/10). Delta re-review (brief p-rv634b.md) was running 15 minutes when stopped; no comment posted. Next: re-review, then no-bump rebase (drop 0.0.595), re-request, enqueue.
  • 621 (jvm-libp2p snapshot pin; shared with the sibling "drain" session which once enqueued it): eaa993a3 (seven pause fixtures observe the pause) and dc832c02 (review finding 8: deadline guards closed at 63 s despite a 4,096 s declaration budget; fixed; mutations 7/7; validation 9/9). Actions red on this head was a curl: (22) ... 504 downloading tooling; I re-ran the failed job at 21:31 (check its result). Next: delta re-review (brief p-rv621c.md), then rebase/land; coordinate 0.0.592 to drop.
  • 629 (mDNS startup stall isolation; upstream jvm-libp2p issue 43): final-reviewed; should-fix verified and pushed 449425ed (selection proof discriminates). Next: no-bump rebase (drop 0.0.598), re-request, enqueue.
  • 630 (root-close reconciliation test barriers) and 633 (synchronised FakeFileSystem fixture): test-only, final-reviewed; remote red on infrastructure only (630's run: three environment timeouts, two assertion flakes listed in the inventory addendum, then SSH loss on all shards; 633: interrupted submission twice, then provisioning). Next: re-request when provisioning works; enqueue on green.
  • 603 (admission probe cleanup), 598 (announcement snapshot + replay defect; three review rounds), 611 (awaited-response outbox grace; nine findings fixed): final-reviewed; remote red on the coordinator deadline / deleted-run upload. Next: no-bump rebases after the first landing, re-request, enqueue.
  • 616 (receive lanes): fixed and final-reviewed but CI red on the pre-existing main test defect = decision 1.
  • 594 (connection-listener isolation diagnostics): held on decision 2; review rv594 found five MUST-FIX items (see finals/rv594-final.md); staged fixes on wip/594-listener-isolation-fix-handoff-2026-10-02 (unverified).
  • 628 (draft reproducer, this campaign), 622 (sibling session; rv622 review posted with a MUST-FIX on description references), 604, 574, 618, 617, 631, 632, 635, 637: other sessions'. 585, 580, 578 closed by triage (585 with nine re-proposals: M12 = 636, M2 = listener-failure tests below, M3 and six more pending; assessment in findings/s585-findings.md).

Flake inventory and what was fixed

  • Fixed in open PRs: root-close reconciliation ticks (630), pending relay registration (634), FakeFileSystem fixture race (633), mDNS startup stall (629), admission probe (603), relay cache/replay (598), keepalive on injected Clock (636).
  • New, unfixed, with evidence (see inventory-addendum.md in the artifacts branch; counts are over today's five full remote runs): testFiniteBackpressureWaitDoesNotRestartForOtherTrafficProgress 3/5 ("A finite 200 ms admission budget must not restart when unrelated channel writes drain. Expected <0>, actual <1>"); testRelayEmitsWithdrawalAtExpiryAfterProviderTransportClose 2/5 ("The observer must register with the real relay before the provider announces its service"); testNettyFrameDeadlineBeforeBodyIsTerminal 2/5 ("Expected ExecutionException ... but was completed successfully", testNettyFrameDeadlineBeforeBodyIsTerminal.kt:127). Briefs p-fbackpressure.md, p-fwithdrawal.md, p-fnettyframe.md are written; the first two ran 15 minutes and were stopped (scaffolding only on their wip branches).
  • Root close racing event-loop shutdown (testProtocolRootCloseCallerRacingEventLoopShutdown, from PR 605, failed on 621's Actions run at iteration 42 "the root is closed"): a deterministic diagnostic on branch wip/root-close-racing-event-loop-shutdown-handoff-2026-10-02 reproduces 1/1 on main: the Netty event loop terminates without closing the root, the physical close is rejected (RejectedExecutionException wrapped in IOException), every caller shares that failed result, and the root stays open. Invariant to implement: after any caller's close completes, success or failure, the root is closed, and the shared failed result still carries the rejection. Brief p-frootclose2.md. Ten amplified copies all timed out at 30 s (not a usable reproducer).
  • Listener failure closes its own transport (585 re-proposal M2): seven public-API tests fail 7/7 on unchanged main (three invariants: close transport before the listener exits; preserve the listener cause and suppress cleanup failures; fatal Error still propagates); on wip/relay-listener-failure-closes-transport-tests-handoff-2026-10-02; no fix yet. Sub-handoff: https://www.handoff.wasmserver.com/handoffs/hf-2026-10-02-finish-the-relay-listener-startup-and-fatal-failure-fix.
  • Also pending from reviews: 585 M3 (listener settles before close observer), a stale-selector race noted by lane f2, fdial3 instrumentation (low priority).

Relevant PRs / refs (branches pushed at stop; all verified with git ls-remote)

Repo Branch Remote head SHA PR What is on it State
CodexCoder21Organization/PlanRepository handoff-artifacts/urlprotocol-main-green-2026-09-20 4a8219f (short; see branch) no PR handoffs/artifacts/urlprotocol-main-green-2026-09-20/resume-2026-10-02/: PLAN.md (campaign log), COMMON.md (lane rules), every brief (p-*.md, REVIEW/FLAKE/REBASE templates), every lane's findings and final message, inventory-addendum.md, the 21:30 report, Actions logs, patches (594 fix, 629 proof, 634 draft, 585 listener tests) documents only
CodexCoder21Organization/UrlProtocol wip/594-listener-isolation-fix-handoff-2026-10-02 c95a9e46c (short) https://github.com/CodexCoder21Organization/UrlProtocol/pull/594 (not pushed to it) staged review fixes for 594 compiles unknown, never run; held on decision 2
CodexCoder21Organization/UrlProtocol wip/root-close-racing-event-loop-shutdown-handoff-2026-10-02 6b3f0595 (short) no PR deterministic diagnostic test (reproduces 1/1 on main) + ten amplified copies (time out; delete) reproducer only, no fix
CodexCoder21Organization/UrlProtocol wip/relay-listener-failure-closes-transport-tests-handoff-2026-10-02 817a7a7e (short) no PR seven listener-failure tests fail-first 7/7 on main, no fix
CodexCoder21Organization/UrlProtocol wip/finite-backpressure-wait-restart-flake-handoff-2026-10-02 3a9efea4 (short) no PR 15 minutes of investigation scaffolding no findings yet; safe to discard
CodexCoder21Organization/UrlProtocol wip/relay-withdrawal-observer-order-flake-handoff-2026-10-02 5515c793 (short) no PR 15 minutes of investigation scaffolding no findings yet; safe to discard
CodexCoder21Organization/UrlProtocol PR branches above (fix/..., see table) per PR table per PR table the campaign's fix PRs see per-PR notes

Deployed or published but not merged: nothing. This campaign published no artifact and deployed nothing (checked: no publish lane ran; kotlin.directory was not written to; the coordinate stays 0.0.569 on main).

Next steps

  1. Confirm the build service provisions again (https://buildtest.kotlin.build/api/health, recent runs reaching TESTING and COMPLETED); if not, that is the blocker and belongs to the service owner (invprov findings give the starting evidence; see memory notes on droplet-service RPC stalls and the droplet limit).
  2. Re-run Actions on 621 if still red; fix 636's Actions failures (brief p-f636.md) and re-review 636 (p-rv636.md) and 634 (p-rv634b.md) and 621 (p-rv621c.md); delegates never merge - the final review and the enqueue stay with the orchestrator.
  3. No-bump rebase (REBASE-TEMPLATE.md) and re-enqueue 626 first; then 629, 634, 598, 603, 611, 621 in any order (independent); re-request 630 and 633 and enqueue on green; after the batch, one publish PR bumping the coordinate.
  4. Flake queue, ordered by recurrence: root-close fix (p-frootclose2.md), fbackpressure, fwithdrawal, fnettyframe, the listener-failure fix (M2), M3, stale-selector race.
  5. Get the operator's answers to decisions 1 and 2; act on them (616 and 594 depend on them).

Reusable / operational knowledge

  • Lane rules that held: at most three test-running lanes share the host's single kompile lock (flock SCRATCH/local-build.lock scripts/test.bash --local --test ...; batch selectors into one invocation; never hold the lock idle); 120-minute budgets; checkpoint unverified work as patches or wip branches; one push per lane; reviews run against the final head and are void if it moves.
  • gh pr edit fails on this repo (deprecated Projects field): edit descriptions with gh api -X PATCH repos/CodexCoder21Organization/UrlProtocol/pulls/N -F body=@file.
  • Build-service ground truth is on the host: ssh -p 23 root@198.199.106.165, /root/buildtest-data/runs/<id>*/{run.json,test-events.jsonl,build.log}; the API hides runs older than the newest window. Classify a red remote check from its summary first: "Provisioning deadline exceeded", "Provisioning stopped", "Upload ... failed", "interrupted before its run ID was recorded", "exceeded its overall deadline of 180 minutes", "SSH session is not connected" are all infrastructure; only named FAILED tests are the PR's.
  • Re-request a check suite with gh api -X POST repos/.../check-suites/<id>/rerequest after warming https://github-webhooks.wasmserver.com/; when the coordinator stays bound to a deleted run, push an empty commit instead.
  • build-watchman 0.0.21 as a tracked background task per PR (--pr plain, --until-merged for the queue); its raw-log probe against the host fails today, which blocks its stall verdict but not polling.
  • Do not pkill -f a string that appears in your own shell command (it kills your shell); kill delegate lanes by the pid that owns their -o <lane>-final.md flag, then sweep their detached test loops.