← Priority list
Blocked

Establish the healthy-route first-bytecode failure mechanism and fix its demonstrated cause

Establish why the initial bytecode controller failed on the healthy CI route, reproduce that condition deterministically, then fix its owner and bring a justified PR to green. Main is unchanged; original passed once, open forwarding controls passed twice, held handshake phase probes failed three times before application dispatch. This is phase evidence only; both evidence branches are pushed, no product fix or PR exists.

Required artifact: a deterministic public-API regression identifying the incorrect first-connection operation. Setup timeout symptoms were reproduced; no implementation defect or verified fix was established.

Handoff document

Markdown

Handoff: Establish the healthy-route first-bytecode failure mechanism and fix its demonstrated cause

Written 2026-10-04 16:45 UTC; retains the earlier census and investigation chain.

RE-VERIFY: All state is a write-time snapshot. Re-check handoff-cli claims before updates, gh pr view <url> --json state,headRefOid,mergeStateStatus,statusCheckRollup,mergedAt, and git ls-remote for main and both evidence branches before acting. The supervisor owns every merge, queue, deployment, publication and completion decision. This lane stopped at CHECKPOINT under the explicit instruction to stop if the healthy-route mechanism cannot be reproduced deterministically. No product fix or new PR exists.

Mission

The user asked for a census of approximately the latest 15 UrlResolver runs, then a fix of the most frequent unowned intermittent test, because unrelated failures were making existing PRs red. The corrected census selected testSandboxedProxyRedialsAfterSilentPersistentRpcTimeout; its failures occur during first bytecode connection setup, before the deliberately silent application handler. L27 was asked to verify the old analysis, compare existing fixes, reproduce the actual mechanism and bring a justified fix to green PRs. The root-cause precondition remains unsatisfied; a deterministic phase experiment is preserved without claiming it explains the natural failure.

L27 verified findings and actions

  1. Claimed only this handoff as fable-hq-20261004. Claims immediately before this update showed only that live agent. Main remains 1053ace97e38d3a3ee85294b07a3f8465adf6fd0; original census branch remains 48662f60dedd40c163befd5b0e79be7b254cf166. No changes were deployed or published, and no PR was merged, queued, closed or opened.
  2. Inspected all supervisor-listed PRs and additional relevant redial/timeout/gossip changes. None changes the named sandbox test; none has been demonstrated to fix its first-bytecode failure. Their own scopes should be preserved. The queued relay branch was not touched.
  3. Inspected initial bytecode/controller observation, lazy bootstrap setup, and pending-dial ownership. The 1000ms connection budget waits for a complete controller, including root handshake and child negotiation. Saved TimeoutException traces do not identify which stage was unfinished. Source links are in source analysis.
  4. Added discovery sources outside maintained tests: the original public lazy sandbox and real provider run through a real TCP forwarding route. One control forwards immediately; one holds handshake bytes until the first public proxy invocation returns. Both retain the original 90s limit, 1000ms budget, complete application-timeout assertion, fresh connection identity and exact method counts. The held case fails that application-timeout assertion with initial-bytecode errors, while upstream TCP was accepted and bytecode/application request counts were zero.
  5. Completed targeted totals: unchanged original 1 PASS/0 FAIL; open forwarding control 2 PASS/0 FAIL; held-handshake phase experiment 0 PASS/3 expected FAIL. This demonstrates setup/application phase separation, not an unexplained defect on a healthy route. Admission-retirement logs also occur in the held case, refuting their sufficiency as evidence of a retirement bug. No warmup, retry, increased deadline or narrowed assertion was adopted.
  6. Separate test-comprehensiveness and adversarial review passes are in findings. Review found unpropagated forwarding-thread failures in the discovery fixture; both sources now collect them and report them during cleanup. Re-verification preserved the expected split without attached cleanup errors. The natural CI mechanism and adjacent lifecycle behaviors are not certified by this experiment.
  7. One remote request had a successful health check, then no exposed run ID/verdict for approximately eight minutes. Its client stack was polling after submission. The exact client JVM was stopped; an unidentified submitted remote run might have continued. No remote result is counted. One review invocation was aborted because rebase failed and its memory reading exceeded the lane threshold; it is not a result. Three completed targeted commands returned the expected mixed results. No full local suite or fat-jar build was run.
  8. Salvage both branches. Original census, full nonempty traces and genuine-concurrency baseline remain useful; L27 adds phase controls, full stacks, source links, live PR comparisons and a remote run-ID challenge draft. No branch or PR is recommended for closure. Do not add the deliberately failing phase probes directly to the passing test suite.

Census and correction

Read census-including-attempts.md, not census.md, for the final ranking. The original table counted only final FAILED rows. The API also stores FAILED attempts beneath a final PASSED row. Including those attempts changes the top item. The corrected table counts a test once per run with any failed row or attempt, and separately reports the number of failed observations. all-failed-attempts.json carries full returned traces; selected-runs.json carries run metadata. combined-failed-census.json and sandbox-census-rows.json preserve the derived ranking and all rows for the top item.

Top items by failing runs: testSandboxedProxyRedialsAfterSilentPersistentRpcTimeout 7 (14 failed observations); testRemovedDeadPeerGossipLaneObjectCensus 6 (6); testPruneVerificationRequiresUrlRpcProtocol 5 (12, flUR3 owns); stressTestConcurrentRelayDisabledBootstrapRecoveryRemainsPrompt 4 (4); stressTestUnverifiedPeerAnnouncementsDoNotBecomeGossipDestinations 4 (4). The table contains 122 test names.

The 15 selected CI runs include paired required/informational runs at a head; counts are per run, not per independent commit. Four selected runs executed zero tests and contain provisioning errors, which are infrastructure observations, not code failures. Every selected result projection still reported complete=false and reconciliationPending=true at collection. All pages through totalCount were fetched, using one-based pages of at most 100 rows. The census is therefore a census of returned observations, not a promise that server reconciliation is complete. The seven sampled PR heads have the same main merge base, 1053ace97e38d3a3ee85294b07a3f8465adf6fd0. No sampled main or merge-queue run establishes a before/after regression boundary, so this sample cannot pin a first bad main commit.

Shared signatures do not yet prove one cause: host creation pending during close spans 21 test names and 7 runs (40 failed observations), bytecode connection setup spans 2 tests and 8 runs (15), child-process READY absence appears in one test across 5 runs (5), and deep-map guest invocation exceeding its limit appears in 3 runs. See shared-signatures.md. The missing READY output needs child-process evidence, rather than assuming that a slow or overloaded host is the cause.

Correct top item: sandbox first-call connection

The 13 nonempty returned traces for testSandboxedProxyRedialsAfterSilentPersistentRpcTimeout fail during initial bytecode connection setup before the deliberate silent application RPC. The fourteenth failed observation has no error text. The test expects the full persistent-RPC timeout; instead it gets UrlResolutionException saying its sole provider failed bytecode fetch after connection attempts timed out. No source-level mechanism has been established for those connection failures.

The original single-path test passed once locally. sandbox-eight-path-discovery.kts exercises eight independent instances of the original public lazy-sandbox path concurrently, each using a distinct service URL and real provider/client. It preserves the 1,000ms connection budget, first-handler silence, second-call success, connection-identity assertions, request counts, and original 90s test limit. It joins every worker before restoring shared public-IP state and records cleanup failures. Its initial baseline and five measured repetitions all passed with JAVA_TOOL_OPTIONS='-XX:+UseSerialGC -XX:ActiveProcessorCount=2': 0/6 failing runs across 48 scenario instances. sandbox-eight-path-trials.json holds the five measured results, and individual-local-test-observations.json includes the first baseline. These are unchanged-baseline results, not post-fix evidence. This path did not meet >=50% failure reliability, so no fix, 20-pass gate, full-suite gate, or PR is justified.

Useful source points in UrlResolver.kt: addBootstrapPeer around 7201 only starts a caller-owned exchange when the host already exists and the server is running; openSandboxedConnection around 30344/30373 is lazy and starts bytecode/RPC setup on the first method call. The original test creates a client without eager join. These are observations, not a demonstrated bug. Do not add an arbitrary warmup ping or enlarge the deadline to make its setup pass without first reproducing and establishing the correct contract.

Earlier gossip investigation and ruled-out hypotheses

Before the census correction, the final-row-only top tie led to stressTestUnverifiedPeerAnnouncementsDoNotBecomeGossipDestinations. Its unchanged test passed 22 times locally: one initial baseline, ten SerialGC runs, one initial fixed-processor run, ten SerialGC/fixed-processor runs. Check unchanged-serial-gc-trials.json, unchanged-serial-gc-fixed-processor-trials.json and individual-local-test-observations.json.

concurrent-pairs-discovery.kts uses eight real sender/receiver pairs, 1,024 announcements per pair. concurrent-pairs-discovery-16.kts uses sixteen pairs, 4,096 announcements per pair. They passed with the normal processor count. The sixteen-pair version repeatedly failed with ActiveProcessorCount=2; the first three failures delivered no services to the receiver and reported zero successful batches. Thread samples and full stacks are retained. The remote eight-pair run https://buildtest.kotlin.build/run?id=4ba2db3e also failed: all 19 pages contain 1 FAILED and 1,888 FILTEREDOUT rows. amplified-remote-selected.json contains the selected test result. A separate original bootstrap backoff test passed remotely in https://buildtest.kotlin.build/run?id=9d610b3f (1 PASSED, 1,887 FILTEREDOUT across 19 pages).

concurrent-pairs-first-message-observation.kts additionally traces the first message and queries public receiver observations. It demonstrated a separate accounting discrepancy: five receivers learned 20-50 services while the sender's successful-batch count remained zero. The drain increments successes only when a whole batch has no unsent messages; a successful prefix is released without incrementing that counter. This explains the diagnostic count in the partial-delivery run, but not the earlier zero-delivery failures. Do not treat a counter-only change as a fix of all these failures. unverified-first-message-stack.txt preserves the observation.

A real TCP proxy experiment in pending-first-contact-discovery.kts pauses the first handshake across a public gossip observation window. It observed deferred sends, reuse of the same accepted TCP connection, then successful flush and receipt. Thus the simple hypothesis that the first observation always cancels the only pending first-contact dial is refuted on that public path. The first version of this discovery fixture read receiver state immediately after a local-flush barrier and failed; adding the existing public receipt barrier corrected the fixture and it passed. That initial fixture failure is not a product reproducer. Thread samples likewise did not establish a resolver worker wait cycle; workers progressed or became idle, and some transport workers temporarily contended during address-resolver class initialization.

Relevant PRs / refs

Branch Remote head Contents and recommendation
Original census evidence 48662f60dedd40c163befd5b0e79be7b254cf166 corrected census, full traces, real-concurrency and gossip discoveries; keep; no PR
L27 phase evidence 1f757f186249f034f692ca07c6161f9dfcf72420 reviewed held/open handshake sources, results, source analysis and findings; keep; no product change or PR
PR State / head Checks
https://github.com/CodexCoder21Organization/UrlResolver/pull/1149 OPEN; d802668b97193095fbeae091cf2eaebfc0140032 bld-build: SUCCESS; kotlin.build (kompile-remote-build): FAILURE; kotlin.build (remote): SUCCESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1183 OPEN; 68f268041dd9121d22babb7735e3ebb89c0d6f64 bld-build: SUCCESS; kotlin.build (kompile-remote-build): FAILURE; kotlin.build (remote): FAILURE
https://github.com/CodexCoder21Organization/UrlResolver/pull/1189 OPEN; 4d93ec6cdbd075c869770ffeb80895d8c46b5a62 bld-build: SUCCESS; kotlin.build (remote): FAILURE; kotlin.build (kompile-remote-build): IN_PROGRESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1190 OPEN; df96365a88deeb2559e494dafd22ebb4dc1b762d bld-build: SUCCESS; kotlin.build (remote): FAILURE; kotlin.build (kompile-remote-build): IN_PROGRESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1163 OPEN; 55f89d6a0288b2987c2d95f160c56a1880d81486 bld-build: SUCCESS; kotlin.build (remote): FAILURE; kotlin.build (kompile-remote-build): IN_PROGRESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1185 OPEN; 8c00203b41d8e51f83647f354902680d0e6e0ec8 bld-build: IN_PROGRESS; kotlin.build (kompile-remote-build): IN_PROGRESS; kotlin.build (remote): IN_PROGRESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1186 OPEN; 70c7908e2fdffa97dd9c686061c1a0ae2b7c7aee bld-build: SUCCESS; kotlin.build (kompile-remote-build): IN_PROGRESS; kotlin.build (remote): IN_PROGRESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1187 OPEN; ed065eaad96e51ae165ea3ae00661502e7fd7dff bld-build: SUCCESS; kotlin.build (remote): FAILURE; kotlin.build (kompile-remote-build): IN_PROGRESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1188 OPEN; fe23943fa820904053dde9195e5c249f60dcc4e5 bld-build: SUCCESS; kotlin.build (kompile-remote-build): FAILURE; kotlin.build (remote): SUCCESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1119 OPEN; 50ecd68f18f7368212da4157e966e6da05e8d349 bld-build: SUCCESS; kotlin.build (kompile-remote-build): FAILURE; kotlin.build (remote): FAILURE
https://github.com/CodexCoder21Organization/UrlResolver/pull/1138 OPEN; 40348c2b545407a0d94077b7dd1d45fb12347d6f bld-build: SUCCESS; kotlin.build (remote): FAILURE; kotlin.build (kompile-remote-build): IN_PROGRESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1171 OPEN; 4c6c33f5a527a1b0e95468afb2c26d481bb55704 bld-build: SUCCESS; kotlin.build (kompile-remote-build): FAILURE; kotlin.build (remote): SUCCESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1178 OPEN; bc9157c3260c534fec7d260a2e5d341833b38771 bld-build: SUCCESS; kotlin.build (kompile-remote-build): FAILURE; kotlin.build (remote): SUCCESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1182 OPEN; 6c3a773143007e3f47bbe40bc5cef2f63667f835 bld-build: FAILURE; kotlin.build (kompile-remote-build): FAILURE; kotlin.build (remote): SUCCESS
https://github.com/CodexCoder21Organization/UrlResolver/pull/1157 OPEN; b9c6360f9f3a3fd85e8c41541d5a060963eca1c0 bld-build: SUCCESS; kotlin.build (kompile-remote-build): NEUTRAL; kotlin.build (remote): SUCCESS

PR states/checks were re-fetched at approximately 16:39 UTC. They are comparison candidates belonging to other work, not L27 deliverables; no exact-failure coverage was proved. Keep them for their own review. Full statuses and their source branches are in investigations/L27/prs-final.json.

Next steps

  1. Re-verify current state first: claims, main, evidence SHAs and candidate PR scopes/checks. Read the full branch findings and result traces. Follow the supervisor's ownership of merges, queue actions and completion.
  2. Determine which healthy-route stage was unfinished before the first-bytecode controller expired: TCP acceptance, authenticated root establishment, muxer or child protocol negotiation. The held-route experiment alone cannot answer this. Capture stage evidence from the exact original public path and its first expired controller; do not infer a cause from aggregate memory/load or admission-retirement text.
  3. Once the specific named mechanism is established, force that same condition deterministically through the public path on unfixed main. Retain the original budgets, assertions and genuine concurrent callers. Write the mechanism and evidence before any fix. A passing baseline or imposed non-responsive network does not satisfy this precondition.
  4. Only then fix the owning implementation (upstream if applicable), prove the unchanged reproducer flips with the required statistical gate, run relevant neighbor tests and the remote full suite if shared contracts change, and open a main-targeting PR after immediate fetch/rebase. Perform both independent review passes and get required checks green. Never merge, enqueue, deploy or publish under this lane authorization.
  5. Optionally have the supervisor publish investigations/L27/remote-client-observation.md as an ecosystem challenge. The report-challenge CLI was not invoked because it automatically enqueues/merges, which this brief prohibits.

Operational knowledge

  • Every shell: export PATH=/code/ws/bin:$PATH. The native global coursier has the wrong architecture. The clone-local ignored jars/coursier shim invokes java -jar /code/ws/bin/coursier.jar "$@"; no binary is committed.
  • Shared host ceiling 6 GiB: read /sys/fs/cgroup/memory.current; wait above 4.5 GiB. No large local fat-jar or full suite. At most two remote submissions per lane; L27 used one. Prefer bounded targeted commands and successful fetch/rebase before starting them.
  • Discovery sources live in investigations/L27/. Copy each to tests/ under its declared function name only for execution, then remove the copies. Example after successful fetch/rebase and memory check:
cp investigations/L27/testSandboxedSilentRpcWithHeldInitialHandshake.kts tests/
cp investigations/L27/testSandboxedSilentRpcWithOpenInitialHandshake.kts tests/
timeout 900 env JAVA_OPTS='-Xmx512m -XX:+UseSerialGC -XX:ActiveProcessorCount=2' scripts/test.bash --local --test testSandboxedSilentRpcWithHeldInitialHandshake --test testSandboxedSilentRpcWithOpenInitialHandshake --log /tmp/L27-phases.xml

The expected result is open control PASS and held phase FAIL at the unchanged application-timeout oracle. Those flags constrain the driver; this is not a claim about a specific forked test processor configuration. Reviewed per-test durations were 13.027s and 16.659s. This command is a phase probe, not a healthy-route bug reproducer.

  • Use build-watchman 0.0.21 with a known run ID or a PR's plain mode, tracked and bounded; never --to-merged. The remote CLI observed here stored runId internally without printing it before polling, so no watcher could be attached to that submission. Do not cancel guessed unrelated runs.
  • Read documents: repository README (no AGENTS.md present), handoffs README (triage/format), PHILOSOPHY (root cause/upstream/refutation), TESTING (public API/real dependencies/deterministic conditions), CODE_REVIEW (host-overload claims and weakened tests). Legacy merge wording is overridden by the brief.
  • Clone is clean, with no stash or additional worktree. No maintained test was deleted or disabled. Per-run XML, compiler output and caches are omitted from Git; the branch retains complete relevant failure stacks and structured results as investigation evidence.

No status reports yet.

Add dependency

Complete this handoff

Moves it out of every priority list and into ArchiveArea.