← Challenges

Repository · challenges

UrlResolver PR https://github.com/CodexCoder21Organization/UrlResolver/pull/770 buildtest run

View on GitHub ↗

UrlResolver PR https://github.com/CodexCoder21Organization/UrlResolver/pull/770 buildtest run

Reported (UTC): 2026-07-14 02:09

UrlResolver PR https://github.com/CodexCoder21Organization/UrlResolver/pull/770 buildtest run 47a67ab5 entered a confirmed provisioning wedge after capacity pressure. First provisioning attempt at 01:47:18 UTC tried ten shards; all ten createDropletAsync calls stalled for exactly 30000ms and emitted DropletServiceCallStalledException, after which buildtest logged Allocation capacity unavailable; requeueing build as PENDING and began a second attempt at 01:48:48 (Provisioning DigitalOcean droplet..., Sharding 1102 discovered tests across 10 droplets). The second attempt never emitted another line. At 02:08:51, /root/buildtest-data/runs/47a67ab5/build.log remained exactly 11662 bytes with mtime 01:48:48.923 and test-events.jsonl remained 181749 bytes with mtime 01:47:18.775, while the API/GitHub still reported PROVISIONING. This satisfies the watch-build WEDGE rule: both files and content frozen >20 minutes, zero test progress, pending check. Impact: the run remains falsely active forever after its self-requeue path, consuming CI time and requiring destructive operator recovery. Workaround: as the single run owner, force-fail it with buildtest-cli delete, warm the webhook receiver, rerequest the suite, verify a changed run ID/directory, and re-arm watchman. Suggested durable fix: apply the same 30-second bounded timeout and explicit failure/requeue handling to every retry attempt, not just the first; track the second attempt's futures, and fail/requeue if no shard creation callback completes within the provisioning phase budget.


Production verification — 2026-07-17

Status: STILL EXISTS. Live GitHub verification found 1 referenced tracker(s) still open: https://github.com/CodexCoder21Organization/UrlResolver/pull/770. A live ProductionHealth connection also emitted repeated NothingToCompleteException gossip failures, while the stopped HardwareControlFabric daemon log ends with Netty ByteBuf leak reports.

This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.