Repository · challenges
buildtest whole-build watchdog timeout counts admission-queue wait, failing healthy still-TESTING
Reported (UTC): 2026-07-14 04:51
buildtest whole-build watchdog timeout counts admission-queue wait, failing healthy still-TESTING merge-group runs at exactly 3h from check creation
What was being attempted: Shepherding UrlResolver PR https://github.com/CodexCoder21Organization/UrlResolver/pull/680 through the merge queue (attempt 3) on 2026-07-14.
What went wrong: The merge-group's kotlin.build (remote) check-run (buildtest run 1819fdd1, merge-group commit 222ab8c9) was marked FAILURE at 04:46:36Z ��� exactly 3h00m after the check-run's started_at 01:46:31Z ��� while /root/buildtest-data/runs/1819fdd1/run.json still reported TESTING with live shard event consumption (spool offsets consumed seconds before the failure, verified). Of that 3-hour budget, roughly 80 minutes were consumed by admission-queue wait (run PENDING behind fleet contention until ~03:00Z), then ~40m building and ~70m in the 10-shard test phase. The watchdog therefore failed a healthy, actively-executing run whose only sin was queueing behind other runs, and the merge queue evicted the PR (third infra eviction in a row, each a different class ��� see challenge 2026-07-14-0146).
Impact: A green, reviewed PR lost another full merge-group cycle; the still-running 10-droplet run continues burning capacity for a check that can no longer succeed. Under fleet contention this defect makes the 3h budget structurally insufficient: admission wait + build + long test suites (resilience fixtures with kill/restart cycles) cannot fit, so heavily-loaded periods evict everything slow.
Workaround used: Classified from check-run timestamps vs run.json state, then re-enqueued via GraphQL enqueuePullRequest (attempt 4).
Suggested durable fix: In buildtest-server (BuildRunner/watchdog), start the whole-build timeout clock at admission/provisioning grant rather than check-run creation ��� or track separate budgets for queue-wait and execution ��� so admission backlog does not consume the execution budget. Additionally, when the watchdog fails a check for a run that is still executing and making verifiable progress (fresh spool consumption), prefer extending with a bounded grace or surfacing progress rather than terminal failure.
Production verification — 2026-07-17
Status: STILL EXISTS. Live GitHub verification found 1 referenced tracker(s) still open: https://github.com/CodexCoder21Organization/UrlResolver/pull/680. The production host simultaneously showed load average 60.59, CPU pressure near 90%, ContainerNursery at 238% CPU with 543 threads, and the BuildTest server at about 1.8 GiB RSS.
This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.