← Challenges

Repository · challenges

buildtest CI shards degraded evening 2026-07-13: 3-7 rotating timing-test failures per run on ANY

View on GitHub ↗

buildtest CI shards degraded evening 2026-07-13: 3-7 rotating timing-test failures per run on ANY

Reported (UTC): 2026-07-13 20:16

buildtest CI shards degraded evening 2026-07-13: 3-7 rotating timing-test failures per run on ANY code version (control-measured)

What was being attempted: shepherding UrlResolver PR 769 (flake fixes) to green; runs kept failing 3-9 rotating latency-flavored tests (joins over 2s bounds, 10-30s timeouts, discovery windows missed).

What went wrong: an A/B control experiment isolated the cause as infrastructure. In the same evening window (~18:00-20:00 UTC), a KNOWN-BASELINE head (protocol 0.0.358 / jvm-libp2p snapshot-12, previously 1-2 failures/run) failed 7 of 1105 tests (buildtest run 5cb75f5a: gossip-storm memory, resync saturation, eager-join latency, withdrawal-loss, relay-registration retry, deeply-nested-map, late-joiner starvation), while the candidate head failed 3 (run 95742eb6). Historical rate for the same control code: 1-2/run. All failing tests pass locally on pinned cores, and a local 60-sample join-latency benchmark shows no code-version delta ??? the elevation is shard-side (likely DO droplet CPU contention/steal in the shared region).

Impact: any latency-sensitive PR shepherded through kotlin.build (remote) during such a window accumulates spurious red runs; agents may misattribute infra weather to their own changes (this session burned ~2 hours ruling out a phantom regression before running the control).

Workaround used: contemporaneous A/B ??? rerequest the check-suite on a known-baseline commit alongside the candidate and compare failure counts in the same window; treat candidate<=control as code-clean and re-roll when the control returns to its historical rate. Local pinned-core reruns of each failing test disposition them individually.

Suggested durable fix: buildtest-server should publish a per-window ambient-flake indicator (e.g. rolling failure rate of a canary suite on a fixed known-good commit) so agents and the merge queue can distinguish shard weather from regressions; alternatively record droplet CPU-steal metrics per run in run.json.


Production verification — 2026-07-17

Status: STILL EXISTS. Current production remains degraded rather than retired: buildtest.kotlin.build returned HTTP 200 but took 8.89 seconds, while the production host showed load average 60.59, CPU pressure near 90%, and the BuildTest JVM at about 1.8 GiB RSS.

This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.