Repository · challenges
UrlResolver remote build 8df2fb29 timed out waiting for kompile-cli after a buildtest service
Reported (UTC): 2026-07-14 03:43
UrlResolver remote build 8df2fb29 timed out waiting for kompile-cli after a buildtest service restart.
What was being attempted: CI verification for https://github.com/CodexCoder21Organization/UrlResolver/pull/773 at commit https://github.com/CodexCoder21Organization/UrlResolver/commit/a33df0a5b809432c9b22369aa1ad1b21eaf978fb.
What went wrong: Remote run https://buildtest.kotlin.build/run?id=8df2fb29 spent its scheduled execution in BUILDING, then terminated before tests with status FAILED, buildSucceeded=false, testsTotal=1102, testsPassed=0, testsFailed=1102, and exact errorMessage 'Build timed out waiting for kompile-cli to complete after service restart'. The synthetic 1102 failed-test count could misleadingly resemble a full suite failure even though no test ran.
Impact: A test-only PR with comprehensive local verification could not reach green, and the entire remote suite had to be rerequested after roughly 48 minutes. This also consumed scarce executor capacity and delayed review.
Workaround used: As the single rerun owner, I verified the failure was pre-test infrastructure, warmed the GitHub webhook receiver, rerequested only kotlin-build-ci-test check suite 79299779877, verified the check details URL changed from run 8df2fb29 to run 1dff71c6, confirmed 1dff71c6 materialized as PENDING in /api/runs, and re-armed build-watchman. GitHub Actions was left untouched.
Suggested durable fix: Make buildtest service-restart recovery preserve or reattach to the live kompile-cli process without charging the restarted wait against a stale timeout; expose restart and child-process progress in the run API; and report pre-test build failures separately instead of marking every discovered test failed.
Production verification — 2026-07-17
Status: STILL EXISTS. Live GitHub verification found 1 referenced tracker(s) still open: https://github.com/CodexCoder21Organization/UrlResolver/pull/773. A live ProductionHealth connection also emitted repeated NothingToCompleteException gossip failures, while the stopped HardwareControlFabric daemon log ends with Netty ByteBuf leak reports.
This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.