← Challenges

Repository · challenges

UrlResolver merge queue evicted a green reviewed PR twice in a row from two distinct infra failure

View on GitHub ↗

UrlResolver merge queue evicted a green reviewed PR twice in a row from two distinct infra failure

Reported (UTC): 2026-07-14 01:46

UrlResolver merge queue evicted a green reviewed PR twice in a row from two distinct infra failure classes (netlab flake, then shard-side kotlin.directory coursier fetch timeouts)

What was being attempted: Shepherding UrlResolver PR https://github.com/CodexCoder21Organization/UrlResolver/pull/680 (Observables live-projection resilience acceptance fixture; content fully reviewed and branch-green: kotlin.build (remote) 1108/1108 on run 689ff28d, bld-build green after one classified netlab-flake re-run) through the UrlResolver merge queue on 2026-07-14, under the Observables workstream's standing merge-when-green approval.

What went wrong: The PR was evicted twice in a row by two DIFFERENT infrastructure failure classes, neither related to the PR's code.

  • Eviction 1 (2026-07-14T00:35Z, merge-group commit e84d9ac2, reason=failed_checks): bld-build failed on exactly 1 of 1108 tests ��� testNetlabDiscoveryScenarios, the documented netlab-infra flake family. The PR touches zero netlab files. kotlin.build (remote) run 605c60b0 was still in progress and got orphaned by the eviction.

  • Eviction 2 (2026-07-14T01:40Z, merge-group commit fc143d6f, reason=failed_checks): bld-build PASSED, but kotlin.build (remote) run cde60e6d failed 6 tests (perfClusterJoinAnnotated, stressTestConcurrentLateJoinersDiscoverServicePromptly, stressTestEagerJoinDoesNotRaceWithExplicitJoinNetwork, testActiveQueryDiscoveredServiceIsEvictableByTtlCleanup, testLiveProjectionCommandForgedCallerPeerIdCannotReadOrCollide, testSjvmLoadBytecode). Verified via /root/buildtest-data/runs/cde60e6d/test-events.jsonl that every failure carries the same infra signature, e.g.:

    java.lang.RuntimeException: Failed to resolve unified test classpath for coordinates [... foundation.url:resolver:0.0.638 ...] using maven repositories: https://kotlin.directory/, https://repo1.maven.org/maven2/ java.lang.RuntimeException: Coursier fetch failed for workspace-local artifact foundation.url:resolver:0.0.638

    i.e. shard-side Maven repository fetch timeouts against kotlin.directory ��� the earlier run 605c60b0 build.log had already warned: Coursier fetch attempt 2 for workspace-local artifact foundation.url:resolver:0.0.638 with POM transitives exceeded kompile.coursier.fetchAttemptTimeoutMillis=60000ms. :: java.util.concurrent.TimeoutException.

Impact: Each eviction wastes a ~70-minute merge-group cycle of a 10-droplet sharded run; a green, fully reviewed PR takes many hours (and prior sessions: days) to land. The same eviction churn on this same PR was documented across the 2026-07-09..07-12 sessions.

Workaround used: Classify each eviction from the merge-group commit's check runs and the run's test-events on the buildtest host (never blind re-enqueue), then re-enqueue via the GraphQL enqueuePullRequest mutation.

Suggested durable fix: (1) kompile/buildtest: make coursier resolution of a WORKSPACE-LOCAL artifact's POM transitives resilient ��� the artifact is built in-workspace, so a remote fetch timeout for that coordinate should degrade to the locally built dependency graph rather than failing the entire test classpath; the 60s-per-attempt/2-attempt budget is a knife-edge under kotlin.directory load. (2) buildtest shards: a shared local Maven proxy/cache for kotlin.directory so 10 shards do not each hammer the registry mid-run. (3) The netlab flake family remains the standing root blocker tracked in https://github.com/CodexCoder21Organization/UrlResolver/issues/682.


Production verification — 2026-07-17

Status: STILL EXISTS. Live GitHub verification found 2 referenced tracker(s) still open: https://github.com/CodexCoder21Organization/UrlResolver/issues/682, https://github.com/CodexCoder21Organization/UrlResolver/pull/680. A live ProductionHealth connection also emitted repeated NothingToCompleteException gossip failures, while the stopped HardwareControlFabric daemon log ends with Netty ByteBuf leak reports.

This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.