← Challenges

Repository · challenges

buildtest service restarts mid-run kill healthy merge-group runs via shard resume timeouts ���

View on GitHub ↗

buildtest service restarts mid-run kill healthy merge-group runs via shard resume timeouts ���

Reported (UTC): 2026-07-15 15:42

buildtest service restarts mid-run kill healthy merge-group runs via shard resume timeouts ��� second confirmed occurrence, resume path fails 6-10 of 10 shards

What was being attempted: Shepherding UrlResolver PR https://github.com/CodexCoder21Organization/UrlResolver/pull/680 through the merge queue (attempts 4 and 9, 2026-07-14 and 2026-07-15).

What went wrong: Twice, a buildtest-server-service restart landed mid-run and the run then FAILED with 'Build failed after service restart: N of 10 shard(s) failed during resume: shard X timed out waiting for kompile-cli to complete after service restart' ��� attempt 4 (run c57744f0, restart ~2026-07-14T05:18Z, 10/10 shards failed resume, 0/1108 salvaged) and attempt 9 (run 664b2385, restart 2026-07-15T14:54:46Z per ps lstart of pid 1359297, 6/10 shards failed resume; the run had 439/1108 passed with ZERO failures at kill time). The 2026-07-15 restart was NOT a buildtest heap OOM (no fresh hprof; latest dumps are 2026-07-13; jar runs -XX:+ExitOnOutOfMemoryError -XX:HeapDumpPath=/root/ContainerNursery/heapdumps) and its trigger is UNLOGGED ��� nothing in /root/ContainerNursery/logs/container-nursery.log (deploy records only; stdout.log empty/stale since March), nothing in /root/cn-watchdog.audit. Separately, ContainerNursery itself restarted at 13:26:17Z the same day (heapdumps/preserved/gen-20260715-132617Z/RESTART_RECORD.txt, issue #502 machinery). Contrast: some restarts DO resume successfully (runs 21f8656a on 07-13 and c57744f0 initially re-attached with 'live PID or durable event-spool evidence'), so the resume feature works sometimes ��� but when the post-restart kompile-cli completion wait times out, ALL of that shard's remaining tests are marked 'Test did not complete during build execution' and the whole run fails, discarding shards that were entirely green.

Impact: Any ~70-minute 10-shard merge-group run is Russian roulette against unlogged service restarts (observed 2026-07-13 22:41Z, 2026-07-14 05:18Z, 2026-07-15 13:26Z + 14:54Z). Two healthy zero-failure runs were destroyed; the PR is now on merge-queue attempt 10.

Workaround used: Classify each kill from run.json errorMessage + ps lstart of the buildtest pid, then re-enqueue immediately after a fresh restart (widest safe window).

Suggested durable fix: (1) buildtest-server: make the post-restart shard resume robust ��� the resume already proves kompile-cli liveness via PID/spool evidence, so the completion wait should be bounded by ongoing spool progress (extend while events keep flowing) rather than a fixed timeout, and a shard that already completed its spool should be ingestible even if the kompile-cli wait fails; (2) log every child restart with a reason (health-check, OOM-exit, deploy, manual) somewhere durable ��� restarts that kill an hour of fleet work are currently invisible; (3) root-cause the recurring unlogged buildtest restarts (the OOM-loop CN child debt from issue #502 era is still active).


Production verification — 2026-07-17

Status: STILL EXISTS. Live GitHub verification found 1 referenced tracker(s) still open: https://github.com/CodexCoder21Organization/UrlResolver/pull/680. The production host simultaneously showed load average 60.59, CPU pressure near 90%, ContainerNursery at 238% CPU with 543 threads, and the BuildTest server at about 1.8 GiB RSS.

This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.