← Challenges

Repository · challenges

buildtest dynamic lease dispatch (PR 329): first production run re-executed every test ~11x (6,468

View on GitHub ↗

buildtest dynamic lease dispatch (PR 329): first production run re-executed every test ~11x (6,468

Reported (UTC): 2026-07-16 07:07

buildtest dynamic lease dispatch (PR 329): first production run re-executed every test ~11x (6,468 starts for 583 tests) despite green CI and correct final results ��� rolled back

What was being attempted: final verification measurement after deploying buildtest-embedded 0.0.318 (dynamic lease-based test dispatch, merged PRs 329+331) to production.

What went wrong: production run 56ce8190 completed 583/583 with 0 failures and hit the target peak concurrency of 40, but its test-events.jsonl records 6,468 test_started events (~11 attempts per test), a 1,808-second test window (vs ~5-6 minutes for comparable static-dispatch runs the same night), TESTING reached only at 13.66 minutes (missing the 10-minute SLO), and 44.6 minutes total. The attempt-aware first-completion-wins projection made results correct while masking ~11x wasted execution ��� with re-lease and tail-rescue shipping dark and zero failures recorded, the churn most plausibly comes from lease adoption/requeue-on-death false positives or duplicate lease issuance in the coordinator/SSH-queue path.

Impact: ~11x test-execution compute burned per run and a net wall-clock regression versus the static path the feature replaced; would have degraded every CI run had it stayed deployed under daytime load.

Workaround used: production rolled back to buildtest-embedded 0.0.306 (static dispatch) at ~2026-07-16 07:05 UTC ��� the rebuilt rollback jar's sha256 (711c79db���) exactly matched the prior known-good deploy, confirming reproducible builds. The dispatch code stays merged in main but MUST NOT be redeployed until fixed.

Suggested durable fix: a codex gpt-5.6-sol TDD session is dispatched to reconstruct the run's dispatch_decision history (the PR's own explainability log), name the duplicating mechanism, and fix it with an initially-failing hermetic repro. Ecosystem lessons: (1) same as the shared-build dead-path finding hours earlier ��� e2e suites must include production-shaped scenarios (here: a full-scale multi-shard run with realistic spool latencies; the hermetic tests never surfaced attempt inflation); (2) 'results correct' metrics can hide efficiency regressions ��� run-level attempt-count (started_events / tests) is a cheap invariant worth asserting in CI and alerting on in production.


Production verification — 2026-07-17

Status: STILL EXISTS. Current production remains degraded rather than retired: buildtest.kotlin.build returned HTTP 200 but took 8.89 seconds, while the production host showed load average 60.59, CPU pressure near 90%, and the BuildTest JVM at about 1.8 GiB RSS.

This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.