← Challenges

Repository · challenges

Buildtest run 1dff71c6 reattached after shard recovery then wedged with completed tests and frozen

View on GitHub ↗

Buildtest run 1dff71c6 reattached after shard recovery then wedged with completed tests and frozen

Reported (UTC): 2026-07-14 05:46

Buildtest run 1dff71c6 reattached after shard recovery then wedged with completed tests and frozen run files.

What was being attempted: Second remote CI attempt for https://github.com/CodexCoder21Organization/UrlResolver/pull/773 after pre-test infrastructure failure 8df2fb29.

What went wrong: Run https://buildtest.kotlin.build/run?id=1dff71c6 completed all 100 partitions (98 passed, 2 failed), then remained pending. Host ground truth showed test-events.jsonl frozen at 2026-07-14 05:12:39 (1,630,298 bytes) and build.log frozen at 2026-07-14 05:19:02 (2,834,218 bytes). The final log lines said shard 9 was re-provisioned, SSH connected, a live PID or durable event spool was found, and it was 'Attaching to the runner event spool at line 1705'; nothing was appended afterward. At 05:42 both files had been frozen over 20 minutes while GitHub still showed pending, satisfying the documented wedge rule. The run also contained infrastructure failure perfClusterJoinAnnotated because three 60-second Coursier fetch attempts for foundation.url:resolver:0.0.632 timed out, plus unrelated testActiveQueryDiscoveredServiceIsEvictableByTtlCleanup returning [] after 10.18 seconds. The PR changes only testHealthySameTierRelayWinsBeforeLowerTier.kts, which did not fail.

Impact: A second full remote attempt consumed roughly 100 minutes and never published a conclusion, despite all partitions completing. CI could not become green and required another full rerequest.

Workaround used: HardwareControlFabric could not be used because the documented client certificate path is absent (separately reported). As the single rerun owner, I used read-only SSH to confirm both run files frozen over 20 minutes, force-failed 1dff71c6 with buildtest-cli delete, warmed the webhook receiver, rerequested suite 79299779877, verified the details URL changed to bc79d16e and that bc79d16e materialized PENDING, then re-armed build-watchman.

Suggested durable fix: Make shard/event-spool reattachment resume tailing or fail explicitly instead of parking after 'Attaching'; add a watchdog that terminalizes runs once all partitions are complete and both spools are frozen; preserve resolved workspace artifacts across shard reprovisioning to avoid redundant Coursier fetch timeouts; and expose host spool mtimes in the run API so operators do not need SSH.


Production verification — 2026-07-17

Status: STILL EXISTS. Live GitHub verification found 1 referenced tracker(s) still open: https://github.com/CodexCoder21Organization/UrlResolver/pull/773. The production host simultaneously showed load average 60.59, CPU pressure near 90%, ContainerNursery at 238% CPU with 543 threads, and the BuildTest server at about 1.8 GiB RSS.

This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.