Repository · challenges
While monitoring BuildTestEmbedded PR CI run 79cb0126 (commit
Reported (UTC): 2026-07-14 05:25
While monitoring BuildTestEmbedded PR CI run 79cb0126 (commit 03316d6b9be590c25f0c2f051f8165268daea57e), the BuildTest service restarted during TESTING after only 7/100 results. Its resume path attempted all ten shards, but addSshKeyToDroplet timed out after 30 seconds for nine existing droplets; shard 8's getDroplet call also stalled, and replacement createDropletAsync then stalled for 30000ms. Eight minutes later every existing shard emitted 'Failed to connect ... after 12 attempts. Last error: timeout: socket is not established'. Impact: a healthy source change cannot obtain a trustworthy CI result and the run spends substantial time appearing merely slow at 7/100. The watch-build procedure correctly prevented premature deletion because build.log was still moving. Suggested durable fix: make restart recovery preserve or re-establish shard SSH credentials reliably, classify an all-shards-unreachable resume as infrastructure failure promptly, and ensure a failed replacement-provision attempt terminates or retries the run deterministically rather than leaving it in TESTING.
Production verification — 2026-07-17
Status: STILL EXISTS. Current production remains degraded rather than retired: buildtest.kotlin.build returned HTTP 200 but took 8.89 seconds, while the production host showed load average 60.59, CPU pressure near 90%, and the BuildTest JVM at about 1.8 GiB RSS.
This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.