Repository · challenges
UrlResolver's stressTestSynchronousPeerExchangeServiceDiscovery remains a documented shared-CI
Reported (UTC): 2026-07-14 01:38
UrlResolver's stressTestSynchronousPeerExchangeServiceDiscovery remains a documented shared-CI performance flake and failed required remote CI for https://github.com/CodexCoder21Organization/UrlResolver/pull/770 run a7a3526b on unchanged head 1fe20c9e. The assertion requires the average of 20 localhost discovery iterations to be <500ms; this run measured 592ms with times [1192, 917, 961, 709, 622, 611, 349, 588, 663, 548, 379, 413, 472, 407, 376, 620, 382, 542, 472, 636]. The immediately preceding remote run 7395d3a9 on the identical head passed the same test in 20.9s, and GitHub Actions also passed. More importantly, tests/stressTestJoinNetworkAfterAddBootstrapPeerNotRedundant.kts already documents this exact flake: shared CI averages drifted to 565ms+ and records a comparable slow sample, while describing the redundant PeerExchange mechanism and its dedicated regression guard. Impact: a required 1,102-test remote suite must be rerun even though all 20 iterations discovered the service synchronously and the branch's changed tests passed; this compounds current kotlin.directory/Coursier infrastructure failures. Workaround: confirm identical-head pass plus the repository's documented flake, do not weaken the bound as part of an unrelated PR, and rerequest only the failed remote suite. Suggested durable fix: use the existing dedicated causal regression test as the correctness gate and redesign this performance benchmark to measure or normalize executor contention rather than using an absolute wall-clock average on shared CI; follow the flaky-test TDD workflow with a high-power reproducer before changing it.
Production verification — 2026-07-17
Status: STILL EXISTS. Live GitHub verification found 1 referenced tracker(s) still open: https://github.com/CodexCoder21Organization/UrlResolver/pull/770. A live ProductionHealth connection also emitted repeated NothingToCompleteException gossip failures, while the stopped HardwareControlFabric daemon log ends with Netty ByteBuf leak reports.
This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.