← Challenges

Repository · challenges

UrlResolver consistently-red main CI: aggregate-flakiness root cause + remaining-flake backlog

View on GitHub ↗

UrlResolver consistently-red main CI: aggregate-flakiness root cause + remaining-flake backlog

Reported (UTC): 2026-07-16 02:08

UrlResolver consistently-red main CI: aggregate-flakiness root cause + remaining-flake backlog (2026-07-16)

Root cause of the consistently-red main

It is AGGREGATE FLAKINESS, not a single bug. ~1105 timing/network-sensitive P2P tests each carry a tiny independent failure probability on CPU-starved s-2vcpu-4gb buildtest shards (kompile runs ~4 forked 128 MB test JVMs per shard; stress tests additionally spawn availableProcessors()*2 CPU spinners). With p ~= 0.002/test, P(all 1105 pass) ~= e^-2.2 ~= 11%, so the build is red ~90% of runs even though each test almost always passes. Confirmed empirically: every run fails a DIFFERENT ~2-5 tests (zero overlap between check-run 86809641043 and its rerun 87412496161, and across 4 fresh CI runs + 60 historical check-runs).

Fixed this session (deterministic, locally reproduced, verified)

  1. PeerRegistry.addPeer O(N^2) whole-serviceIndex rescan on every gossip message -> hung testDiscoveredServicesSizeCap past its 30 s limit; fixed to O(peer's own services). A PARALLEL effort landed the same fix independently ??? UrlProtocol #362 (published protocol 0.0.360) + UrlResolver #776 bumped the pin ??? so my UrlProtocol PR #361 was CLOSED as redundant.
  2. testGossipBroadcastStormWithStalledPeersBoundsQueuedMemory ??? a TEST measurement bug, not a production leak. It sampled Runtime.totalMemory()-freeMemory() after System.gc()+fixed sleep, but G1's concurrent explicit-GC had not finished, mislabeling ~18-20 MB of dead garbage as retained. A live GC.class_histogram at the in-flight point showed only 2.43 MB genuinely reachable (256 bounded QueuedGossipMessage, 232 JSONObject, 36 DispatchedContinuation) ??? the per-peer bounded gossip queues work. Fixed to wait for GC collectionCount to advance before sampling. UrlResolver PR https://github.com/CodexCoder21Organization/UrlResolver/pull/777 (10/10 fail -> 10/10 pass, cache-busted). Also inherited #769's two fixes (withdrawal-loss, registration-resync give-up).

Backlog ??? environment-specific flakes NOT reproducible in the dev sandbox

These passed 0/10 locally even under taskset 1-2 CPU on the fast 16-core dev box, so they cannot be fixed test-first here without a CI-matching 2-vCPU droplet reproduction (shipping a guessed fix without a reliable repro violates the CLAUDE.md foundational rule, which is why they were backlogged rather than patched):

  • stressTestBootstrapServiceResyncDoesNotGiveUpUnderSaturation ??? CI shows the confirmed-retry peer-exchange stuck in still-handshaking/dial-failed for the full 75 s after an announcement is lost under relay backpressure; a residual gap after #769. Codex forensics: /home/hostusr/workspace/codex_regsaturation_report.md.
  • stressTestConcurrentRegistrationsReliablyPropagateToRelay.
  • NetLab netem cluster (need Docker/real droplets): testNetlabResolutionUnderLatencyAndPacketLoss (~8x), testNetlabP2PIntegration (~5x), testNetlabDynamicFailureRecovery, etc.
  • Relay-selection ordering: testHealthySameTierRelayWinsBeforeLowerTier.
  • Relay-recovery / NAT re-attach: testRelayRestartSameIdentityReattach.
  • Residual timing-proxy stress tests: stressTestRegistrationAfterWithdrawalRaceUnderContention (~6x) and various latency-bound assertions.

Impact

main CI is red ~90% of runs, blocking merges (PRs cannot get a green window without luck) and masking real regressions.

Suggested durable fix

Provision an s-2vcpu-4gb droplet matching the buildtest shard, reproduce each backlog flake at >=50% there, and fix test-first at the owning layer (mostly UrlProtocol2 in UrlResolver.kt and relay backpressure in foundation.url:protocol). Longer term, the aggregate-flakiness math means the per-test flake budget must be driven very low OR the suite sharded so a single flake does not fail the whole build. Do NOT ship guessed fixes without a reliable reproduction.


Production verification — 2026-07-17

Status: STILL EXISTS. Live GitHub verification found 1 referenced tracker(s) still open: https://github.com/CodexCoder21Organization/UrlResolver/pull/777. A live ProductionHealth connection also emitted repeated NothingToCompleteException gossip failures, while the stopped HardwareControlFabric daemon log ends with Netty ByteBuf leak reports.

This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.