Repository · handoffs
id: hf-2026-09-20-make-urlprotocol-main-reliably-green-land-the-launcher-pin-then-merge-the-six-reviewed-prs-and-fix-the-last-flake url: url://handoff/handoffs/hf-2026-09-20-make-urlprotocol-main-reliably-green-land-the-launcher-pin-then-merge-the-six-reviewed-prs-and-fix-the-last-flake title: Make the UrlProtocol main branch reliably green: land the ten reviewed fix PRs once the build service provisions again, then finish the flake queue summary: Once the shared build service provisions droplets again, no-bump-rebase and re-enqueue 626, then 629, 634, 598, 603, 611, 621, re-request and enqueue test-only 630 and 633, fix 636's Actions failures and re-review 636/634/621, then work the flake queue (root-close fix from a 1/1 diagnostic, finite-backpressure, relay-withdrawal, Netty frame deadline, listener-failure fix) and get the operator's two decisions (correct the main-branch refusal test; environment-bound timeouts via the kompile-core compiler cap). Main is unchanged at 78deea8c; every required remote check is red on provisioning since 17:50 UTC; all partial work is on wip branches and the PlanRepository artifacts branch listed in the body. created: 2026-09-20T22:32:03.445Z completed: null blocked-reason: BLOCKED-EXCLUDED: cold shared-cache preparation for UrlProtocol launcher pin misses its unchanged deadline; owning BuildTest/kompile preparation investigation is excluded. dependencies:
- url://handoff/handoffs/hf-2026-09-20-resolve-the-fatal-admission-test-contract-before-finishing-urlprotocol-drain-reconciliation
Make the UrlProtocol main branch reliably green: land the reviewed fix PRs once the build service provisions again, then finish the flake queue
Written 2026-10-02 ~21:45 UTC by session fable-code6d-2026-10-02 (explicit stop by the operator). RE-VERIFY before acting: everything below is a write-time snapshot (PR states and checks verified at 2026-10-02 21:40 UTC). Re-check with gh pr list --repo CodexCoder21Organization/UrlProtocol --state open --json number,headRefOid,mergeStateStatus,statusCheckRollup, gh api repos/CodexCoder21Organization/UrlProtocol/commits/main, the merge queue GraphQL query, and https://buildtest.kotlin.build/api/health. Previous campaign state (2026-09-20 to 2026-10-02 morning) is in the earlier versions of this handoff's reports (handoff-cli reports <id>) and in the artifacts branch below.
Mission (operator's /goal, verbatim)
"ensure the main branch of https://github.com/CodexCoder21Organization/UrlProtocol is reliably green. Also triage the open pull requests, decide which are bad/obsolsete and close them or which are good/salvagable and fix/merge/salvage the good parts and drive to completion. Parallelize to the extent possible and race PRs."
Where it stands (verified)
- main =
78deea8c61311e47596dad9c99b72763292e70e8(unchanged since 2026-10-01; nothing landed today). Merge queue: PR622:AWAITING_CHECKS. Two entries (PR 622 by a sibling session and PR 626 by this one) were removed by the merge-queue bot at 20:10 UTC after 622's queue run exceeded the build coordinator's 180-minute deadline. - The shared build service (kotlin.build (remote), the required check) cannot currently finish or even start runs. Since ~17:50 UTC every UrlProtocol head fails before tests run with "Provisioning failed before tests ran: Provisioning deadline exceeded for runId=... currentAttempt=4" (runs 3d6c934b/629, cbbb7a50/633, f550dd4d/636, 3a562f7b/618); run 98bf14af (634) lost all shards to "Provisioning stopped ... method=createBuildDroplet" after 2241 passes; run d9c01208 (621) "Upload of build failed after 5 attempts". Earlier (14:49-16:30 admissions) runs ran but exceeded the 180-minute overall deadline while TESTING at 94-97 percent (fleet oversubscribed: 38 TESTING runs at 17:10; 120 active droplets at 21:24). Challenge filed: https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-02-1812-remote-check-runs-on-urlprotocol-exceed-the-ci-coordinator.md (and the morning one: https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-02-1040-build-service-infrastructure-not-test-flakes-is-the-main.md). A read-only investigation lane was started at 21:27 and stopped 15 minutes later with no conclusion (findings in the artifacts branch:
findings/invprov-findings.md). Nothing can land until provisioning works; do not re-request checks into a saturated or non-provisioning fleet (each re-request burns ten droplets for three hours). - Decision: no per-PR coordinate bumps. Every PR keeps main's
coordinates = "foundation.url:protocol:0.0.569"line (seeREBASE-TEMPLATE.mdin the artifacts branch); one publish PR bumps the coordinate after the batch lands (precedent https://github.com/CodexCoder21Organization/UrlProtocol/pull/625). PRs that still carry a bump (636 at 0.0.602, 634 at 0.0.595, 621 at 0.0.592, 629 at 0.0.598, 598 at 0.0.596, 616 at 0.0.597) get it dropped at their no-bump rebase. - Two decisions are open with the operator (asked in the status reports, unanswered): (1) may the main-branch test
testForwardedOutputRefusalPreservesAtomicBoundarybe corrected? Its final clean-write assertion contradicts the refusal behaviour (a refusal closes the incomplete backend source, so the backend's remaining write gets Broken pipe); reproducertestForwardedOutputRefusalBeforeBackendTailWritefails 3/3 on PR 616 and 3/3 on main (lane f616b). Recommendation A: fix the fixture in its own PR, then 616 lands behind it. (2) The remaining 30-second timeout families (candidate-prepare stress, materialization budget, peer-exchange admission, TCP-server-runs-every-request, large-relay-response, aggregate-refusal-throttle) are environment-bound: droplet CPU sharing after kompile-core PR 319 dropped the C1 compiler cap for test JVMs (memory note; partly mitigated by kompile-core PR 350). Recommendation A: restore the test-JVM compiler cap / thresholds in kompile-core; then B: partition the two heaviest tests. Investigations that refuted code-level hypotheses with numbers: lanes fdial, fdial2, f4, f5 (findings in the artifacts branch).
Open PRs (verified 2026-10-02 21:40 UTC)
| PR | branch | head | author | merge state | required checks |
|---|---|---|---|---|---|
| 637 | wip/rpc-stream-reader-ownership |
760676789f753cc4278ab5482c76ae13f471a751 |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS |
| 636 | fix/relay-keepalive-on-injected-clock |
c7dde8b8f607da7a77b0cd22269ad0b9461220f9 |
codexcoder21 | BLOCKED | bld-all-tests=FAILURE, kotlin.build (remote)=FAILURE |
| 635 | fix/connection-closed-record-without-owner-scope |
1c78aeabd156785603a297d59baf09ba83a4fb42 |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE |
| 634 | fix/pending-relay-registration-recovery-flake |
e2eb21efa71425d1542a9eef8ebf76268c1572cf |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE |
| 633 | fix/peer-registry-removal-batch-fixture-race |
bc74f12424061d784046f34aa4fbbc07871d0478 |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE |
| 632 | fix/Lexp2-conditional-evidence-removal |
50a754f2ef186f6adfb3d33379a56c88783fe2f7 |
codexcoder21 | BLOCKED | bld-all-tests=, kotlin.build (remote)= |
| 631 | fix/request-write-final-chunk-phase |
de7007d161465fc859af823f7da7093eaf689961 |
codexcoder21 | BLOCKED | bld-all-tests=FAILURE, kotlin.build (remote)=FAILURE |
| 630 | fix/root-close-reconciliation-tick-flake |
465049c14b3214d9f2296eca64717a7976cacc84 |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE |
| 629 | fix/mdns-hostinfo-startup-stall |
449425ed4ede1ae4943d918d8055921448fb82fb |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE |
| 628 | fix/loopback-dial-timeout-flake |
afc84f6460a0f694694af4207267246da6253f3a |
codexcoder21 | BLOCKED (draft) | bld-all-tests=FAILURE, kotlin.build (remote)=FAILURE |
| 626 | fix/tcp-listener-release-and-effect-delivery |
452ecde9196e2f0bc19645b97d111139463b5c55 |
codexcoder21 | UNSTABLE | bld-all-tests=SUCCESS, kotlin.build (remote)=SUCCESS |
| 622 | wip/sweep4-P3c-preadmission-close-signal |
8e83a9f943a2c8c528ca03ad94f17ec9b29e8995 |
codexcoder21 | UNSTABLE | bld-all-tests=SUCCESS, kotlin.build (remote)=SUCCESS |
| 621 | fix/libp2p-snapshot-28-pin |
dc832c02c8f8b3a07fa0fb8c425e2120557c23ea |
codexcoder21 | BLOCKED | bld-all-tests=, kotlin.build (remote)=FAILURE |
| 618 | fix/relay-gossip-large-frame-child-flow-control |
9d2057d6058884c9ddce68158a20210e09c61256 |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE |
| 617 | fix/peer-registry-eviction-off-caller-thread |
b33bbb5b6d2a66d9881f94eee33510609c79e980 |
codexcoder21 | DIRTY (draft) | kotlin.build (remote)=FAILURE |
| 616 | fix/parse-budget-not-held-while-receiving |
32bda7b35c3998e246e017da14f60f22fb928f01 |
codexcoder21 | BLOCKED | bld-all-tests=FAILURE, kotlin.build (remote)=FAILURE |
| 611 | fix/stall-retirement-requires-no-progress |
b30978b8ec5cd97486872c9f4cce2f7f0366a01d |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE |
| 604 | fix/close-drains-active-sends |
ca45f8b4cd4b2c29aaa3d3d59c318fc74551dd24 |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS, kotlin.build (remote)= |
| 603 | wip/flake-admission-manual-read-2026-09-20 |
581ee4c28bdddbe7d59dba7facf86743fbbb3d6c |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE |
| 598 | wip/FL3c-flake-2026-09-14 |
5e59b3df2564f1c4dd27d10e22093e4b73275d51 |
codexcoder21 | BLOCKED | bld-all-tests=SUCCESS, kotlin.build (remote)=FAILURE |
| 594 | wip/FL4-flake-2026-09-14 |
d53ef33f58ad7f54e59e3dba136855e89cb41908 |
codexcoder21 | UNSTABLE | bld-all-tests=SUCCESS, kotlin.build (remote)=SUCCESS |
| 574 | wip/r71-tcp-serving-lifecycle |
2b8e8c305c5fb1f6b3fda990757350ef62b42a35 |
codexcoder21 | BLOCKED | bld-all-tests=, kotlin.build (remote)= |
Ownership and next action per PR of this campaign (all authored by codexcoder21 = this session's delegates unless noted):
- 626 (TcpServer lifecycle rewrite + effect-delivery deferral): final-reviewed by me, two delegate review rounds clean; was enqueued, evicted 20:10. Next: no-bump rebase (REBASE-TEMPLATE.md), re-enqueue when the service provisions.
- 636 (relay keepalive ticks and pong deadlines on the injected Clock; re-proposal of closed 585): opened 18:56 from branch
fix/relay-keepalive-on-injected-clockat c7dde8b8; local proof 34/34; Actions suite red on this head:testCurrentGenerationFailedRelayRegistrationSchedulesNextReconnectandtestCurrentGenerationThrowingRelayRegistrationSchedulesNextReconnect("The initial registration must schedule exactly one reconnect check") - plausibly caused by the keepalive tick now sharing the Clock those tests count - plustestPendingRelayRegistrationPreservesSameRelayFailuresUntilAcknowledged(main's flake, fixed by 634). Full log: artifacts branchactions-logs/actions-636-110981530691.log. A fix lane (briefp-f636.md) ran 10 minutes and was stopped; nothing pushed. My own read of the source diff found nothing blocking; two points for the delegate review (briefp-rv636.md): a possible idle-tick leak if a connection is replaced without detaching, and the pong-timeout expression's readability. Next: fix the Actions failures (f636 brief), then the delegate review, then my final review. - 634 (pending relay registration recovery): review MUST-FIX closed by e2eb21ef (failure ownership captured under the registration monitor; proof 28/28, 10/10). Delta re-review (brief
p-rv634b.md) was running 15 minutes when stopped; no comment posted. Next: re-review, then no-bump rebase (drop 0.0.595), re-request, enqueue. - 621 (jvm-libp2p snapshot pin; shared with the sibling "drain" session which once enqueued it): eaa993a3 (seven pause fixtures observe the pause) and dc832c02 (review finding 8: deadline guards closed at 63 s despite a 4,096 s declaration budget; fixed; mutations 7/7; validation 9/9). Actions red on this head was a
curl: (22) ... 504downloading tooling; I re-ran the failed job at 21:31 (check its result). Next: delta re-review (briefp-rv621c.md), then rebase/land; coordinate 0.0.592 to drop. - 629 (mDNS startup stall isolation; upstream jvm-libp2p issue 43): final-reviewed; should-fix verified and pushed 449425ed (selection proof discriminates). Next: no-bump rebase (drop 0.0.598), re-request, enqueue.
- 630 (root-close reconciliation test barriers) and 633 (synchronised FakeFileSystem fixture): test-only, final-reviewed; remote red on infrastructure only (630's run: three environment timeouts, two assertion flakes listed in the inventory addendum, then SSH loss on all shards; 633: interrupted submission twice, then provisioning). Next: re-request when provisioning works; enqueue on green.
- 603 (admission probe cleanup), 598 (announcement snapshot + replay defect; three review rounds), 611 (awaited-response outbox grace; nine findings fixed): final-reviewed; remote red on the coordinator deadline / deleted-run upload. Next: no-bump rebases after the first landing, re-request, enqueue.
- 616 (receive lanes): fixed and final-reviewed but CI red on the pre-existing main test defect = decision 1.
- 594 (connection-listener isolation diagnostics): held on decision 2; review rv594 found five MUST-FIX items (see
finals/rv594-final.md); staged fixes onwip/594-listener-isolation-fix-handoff-2026-10-02(unverified). - 628 (draft reproducer, this campaign), 622 (sibling session; rv622 review posted with a MUST-FIX on description references), 604, 574, 618, 617, 631, 632, 635, 637: other sessions'. 585, 580, 578 closed by triage (585 with nine re-proposals: M12 = 636, M2 = listener-failure tests below, M3 and six more pending; assessment in
findings/s585-findings.md).
Flake inventory and what was fixed
- Fixed in open PRs: root-close reconciliation ticks (630), pending relay registration (634), FakeFileSystem fixture race (633), mDNS startup stall (629), admission probe (603), relay cache/replay (598), keepalive on injected Clock (636).
- New, unfixed, with evidence (see
inventory-addendum.mdin the artifacts branch; counts are over today's five full remote runs):testFiniteBackpressureWaitDoesNotRestartForOtherTrafficProgress3/5 ("A finite 200 ms admission budget must not restart when unrelated channel writes drain. Expected <0>, actual <1>");testRelayEmitsWithdrawalAtExpiryAfterProviderTransportClose2/5 ("The observer must register with the real relay before the provider announces its service");testNettyFrameDeadlineBeforeBodyIsTerminal2/5 ("Expected ExecutionException ... but was completed successfully", testNettyFrameDeadlineBeforeBodyIsTerminal.kt:127). Briefsp-fbackpressure.md,p-fwithdrawal.md,p-fnettyframe.mdare written; the first two ran 15 minutes and were stopped (scaffolding only on their wip branches). - Root close racing event-loop shutdown (
testProtocolRootCloseCallerRacingEventLoopShutdown, from PR 605, failed on 621's Actions run at iteration 42 "the root is closed"): a deterministic diagnostic on branchwip/root-close-racing-event-loop-shutdown-handoff-2026-10-02reproduces 1/1 on main: the Netty event loop terminates without closing the root, the physical close is rejected (RejectedExecutionException wrapped in IOException), every caller shares that failed result, and the root stays open. Invariant to implement: after any caller's close completes, success or failure, the root is closed, and the shared failed result still carries the rejection. Briefp-frootclose2.md. Ten amplified copies all timed out at 30 s (not a usable reproducer). - Listener failure closes its own transport (585 re-proposal M2): seven public-API tests fail 7/7 on unchanged main (three invariants: close transport before the listener exits; preserve the listener cause and suppress cleanup failures; fatal Error still propagates); on
wip/relay-listener-failure-closes-transport-tests-handoff-2026-10-02; no fix yet. Sub-handoff: https://www.handoff.wasmserver.com/handoffs/hf-2026-10-02-finish-the-relay-listener-startup-and-fatal-failure-fix. - Also pending from reviews: 585 M3 (listener settles before close observer), a stale-selector race noted by lane f2, fdial3 instrumentation (low priority).
Relevant PRs / refs (branches pushed at stop; all verified with git ls-remote)
| Repo | Branch | Remote head SHA | PR | What is on it | State |
|---|---|---|---|---|---|
| CodexCoder21Organization/PlanRepository | handoff-artifacts/urlprotocol-main-green-2026-09-20 | 4a8219f (short; see branch) | no PR | handoffs/artifacts/urlprotocol-main-green-2026-09-20/resume-2026-10-02/: PLAN.md (campaign log), COMMON.md (lane rules), every brief (p-*.md, REVIEW/FLAKE/REBASE templates), every lane's findings and final message, inventory-addendum.md, the 21:30 report, Actions logs, patches (594 fix, 629 proof, 634 draft, 585 listener tests) |
documents only |
| CodexCoder21Organization/UrlProtocol | wip/594-listener-isolation-fix-handoff-2026-10-02 | c95a9e46c (short) | https://github.com/CodexCoder21Organization/UrlProtocol/pull/594 (not pushed to it) | staged review fixes for 594 | compiles unknown, never run; held on decision 2 |
| CodexCoder21Organization/UrlProtocol | wip/root-close-racing-event-loop-shutdown-handoff-2026-10-02 | 6b3f0595 (short) | no PR | deterministic diagnostic test (reproduces 1/1 on main) + ten amplified copies (time out; delete) | reproducer only, no fix |
| CodexCoder21Organization/UrlProtocol | wip/relay-listener-failure-closes-transport-tests-handoff-2026-10-02 | 817a7a7e (short) | no PR | seven listener-failure tests | fail-first 7/7 on main, no fix |
| CodexCoder21Organization/UrlProtocol | wip/finite-backpressure-wait-restart-flake-handoff-2026-10-02 | 3a9efea4 (short) | no PR | 15 minutes of investigation scaffolding | no findings yet; safe to discard |
| CodexCoder21Organization/UrlProtocol | wip/relay-withdrawal-observer-order-flake-handoff-2026-10-02 | 5515c793 (short) | no PR | 15 minutes of investigation scaffolding | no findings yet; safe to discard |
| CodexCoder21Organization/UrlProtocol | PR branches above (fix/..., see table) | per PR table | per PR table | the campaign's fix PRs | see per-PR notes |
Deployed or published but not merged: nothing. This campaign published no artifact and deployed nothing (checked: no publish lane ran; kotlin.directory was not written to; the coordinate stays 0.0.569 on main).
Next steps
- Confirm the build service provisions again (https://buildtest.kotlin.build/api/health, recent runs reaching TESTING and COMPLETED); if not, that is the blocker and belongs to the service owner (invprov findings give the starting evidence; see memory notes on droplet-service RPC stalls and the droplet limit).
- Re-run Actions on 621 if still red; fix 636's Actions failures (brief p-f636.md) and re-review 636 (p-rv636.md) and 634 (p-rv634b.md) and 621 (p-rv621c.md); delegates never merge - the final review and the enqueue stay with the orchestrator.
- No-bump rebase (REBASE-TEMPLATE.md) and re-enqueue 626 first; then 629, 634, 598, 603, 611, 621 in any order (independent); re-request 630 and 633 and enqueue on green; after the batch, one publish PR bumping the coordinate.
- Flake queue, ordered by recurrence: root-close fix (p-frootclose2.md), fbackpressure, fwithdrawal, fnettyframe, the listener-failure fix (M2), M3, stale-selector race.
- Get the operator's answers to decisions 1 and 2; act on them (616 and 594 depend on them).
Reusable / operational knowledge
- Lane rules that held: at most three test-running lanes share the host's single kompile lock (
flock SCRATCH/local-build.lock scripts/test.bash --local --test ...; batch selectors into one invocation; never hold the lock idle); 120-minute budgets; checkpoint unverified work as patches or wip branches; one push per lane; reviews run against the final head and are void if it moves. gh pr editfails on this repo (deprecated Projects field): edit descriptions withgh api -X PATCH repos/CodexCoder21Organization/UrlProtocol/pulls/N -F body=@file.- Build-service ground truth is on the host:
ssh -p 23 root@198.199.106.165,/root/buildtest-data/runs/<id>*/{run.json,test-events.jsonl,build.log}; the API hides runs older than the newest window. Classify a red remote check from its summary first: "Provisioning deadline exceeded", "Provisioning stopped", "Upload ... failed", "interrupted before its run ID was recorded", "exceeded its overall deadline of 180 minutes", "SSH session is not connected" are all infrastructure; only named FAILED tests are the PR's. - Re-request a check suite with
gh api -X POST repos/.../check-suites/<id>/rerequestafter warming https://github-webhooks.wasmserver.com/; when the coordinator stays bound to a deleted run, push an empty commit instead. - build-watchman 0.0.21 as a tracked background task per PR (
--prplain,--until-mergedfor the queue); its raw-log probe against the host fails today, which blocks its stall verdict but not polling. - Do not
pkill -fa string that appears in your own shell command (it kills your shell); kill delegate lanes by the pid that owns their-o <lane>-final.mdflag, then sweep their detached test loops.