RE-VERIFY (2026-10-01 ~01:40Z, lane K58 verify-and-disposition pass): Everything in the table below was checked firsthand at that time against GitHub (GraphQL), a fresh clone of BuildTestEmbedded main dd7b4641a7a9b6a336b03546efbaee0826f38d53 (2026-10-01 00:31Z, PR 1179), https://buildtest.kotlin.build/api/runs?limit=500, and raw run logs. Older sections further down are history only. Re-check grep -n "shardCount > 1 && cacheBuildRules.isNotEmpty()" src/buildtest/embedded/BuildTestEmbeddedService.kt and DEFAULT_IMAGE in src/buildtest/embedded/DropletManager.kt before acting.
Remaining work
- Decision (owner of BuildTestEmbedded shared preparation): does one-shard cross-run workspace cache reuse still pay for itself after the PR 1082 / PR 1168 redesign? Precondition is met (PR 1082 merged). Evidence below leans toward won't-do; if the owner decides to do it, it is a port onto current main, not a rebase of the old branch.
- Image rebake (production action, needs operator authorization; owned by the BuildTestEmbedded / droplet-image owner): the fleet image 245388985 was baked with buildtest-runner 0.0.104 warm, but production now runs runner 0.0.119, so every run's shard 0 downloads the runner closure cold during workspace preparation. This is a NEW reason to rebake; the rebake this handoff originally named (runner 0.0.96, image 242809349) is long superseded.
Verified state table
| Item |
State (verified 2026-10-01) |
Disposition |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/747 |
MERGED 2026-09-23 18:54Z as test-only ("Cover quoted runner cache paths", merge 9d4f771f). Branch feat/single-shard-workspace-cache-reuse head 960ef7e48ed8f19a76827a69c27f3c22ed4c5862, diverged from main (ahead 1, behind 114). |
DONE (as repurposed); the single-shard change it once carried did NOT land. |
Original single-shard commits b3f113d0 (fail-first test crossRunModuleBuildCacheReusesSingleShard) .. 9edb6fcba717d08eb5fec607f79564a33dbd43eb |
Still fetchable by SHA from GitHub but on NO branch (dangling; may be garbage-collected). |
STILL-OPEN only if the decision is "do it": the owner should push a rescue ref first (e.g. git push origin 9edb6fcb:refs/heads/rescue/single-shard-cache-reuse-9edb6fcb). This verify lane pushed nothing by instruction. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1082 (precondition) |
MERGED 2026-09-27 02:59Z (merge a4201b79). |
DONE; the block on it is lifted. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1084 |
MERGED 2026-09-24 02:39Z. |
DONE (context only). |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1168 (one build-only runner invocation for all shared rules) |
MERGED 2026-09-30 19:08Z (merge 4da8f85a). Its 17-20 min preparation times vs the 1200 s deadline are tracked by hf-2026-09-30-buildtest-restart-recovery-... |
OWNED-BY-ANOTHER-ACTOR. |
| Single-shard gate on main |
Still present: BuildTestEmbeddedService.kt:20361 val sharedBuildCache = if (shardCount > 1 && cacheBuildRules.isNotEmpty()). No crossRunModuleBuildCacheReusesSingleShard test on main. Main has tests/e2eSingleShardBuildDoesNotStartSharedCacheServer.kts (asserts one-shard runs start no cache server and fetch no shared cache; it does not forbid a restore from the persisted store, so it does not by itself settle the decision). |
STILL-OPEN as a decision (item 1). Not enabled on main. |
| Single-shard reuse in production |
Not happening. One-shard run 3af77e00 (ScreenshotTestServerService @0e4b8208, 26 tests) logs no "Restored ... workspace module cache" or "Published workspace module cache" line; build+tests ran inline in the work-queue runner, droplet-ready 20:30:38 -> verdict 20:34:10 (about 3.5 min). Ten-plus identical-commit reruns of that same run each built cold. |
Confirms the gap exists; see value evidence below. |
| Multi-shard cross-run cache |
Working in production: runs 3ba805a4, 9628a572, eb01045e each log "Preparing reusable workspace module cache on shard 0" and "Published workspace module cache <fingerprint> for reuse by later runs"; preparation took 6.4 / 7.3 / 8.0 min. No restore hits seen in these (fingerprint covers the whole source archive, so only byte-identical reruns hit). |
DONE (not this handoff's scope). |
| Image rebake (original: runner 0.0.96 warm, image 242809349 from PR 722) |
Superseded: main DropletManager.kt:979 DEFAULT_IMAGE = "245388985" ("buildtest-ci-ubuntu2404-warmcache-v7-20260914-runner-0.0.104", PR 997 2026-09-14); prior v6 245352603. All 380 non-pending rows of the last 500 runs report dropletImage 245388985. |
OBSOLETE (the 0.0.96 rebake). |
| Image vs current runner |
Main pins runner 0.0.119 (BuildDriver.kt:658, kompile-manager 0.0.188); image warms 0.0.104. Shard 0 of run 3ba805a4 logs 684 "Downloading" lines starting with buildtest-runner/0.0.119/buildtest-runner-0.0.119.pom inside workspace preparation (followers 1-3: 84 total). |
NEEDS-OWNER-ACTION: rebake with runner 0.0.119 (+ kompile-manager 0.0.188) warm, following the procedure in DropletManager.kt comments; operator authorization required. Relevant to the 17-20 min preparation investigation. |
Value evidence for the item-1 decision
- Frequency: 31 of the last 500 runs (~12 h window, 9021 total) had fewer than 50 tests (single-shard territory at
minTestsPerShard=25); most were one ScreenshotTestServerService commit rerun repeatedly.
- Cost of a one-shard cold build is small (about 2.5 min of droplet time in run 3af77e00) and overlaps test dispatch in the work-queue runner; a restore would add an artifact transfer of hundreds of MB (multi-shard artifacts were 371 MB - 1.17 GB) and only hits on byte-identical archives.
- Recommendation for the owner: complete this handoff as won't-do unless measurements show repeated identical one-shard reruns are a material share of fleet time. If doing it: rescue
9edb6fcb, port onto current main (drop the shardCount > 1 term, pass the cache into the inline one-shard pipeline, keep e2eSingleShardBuildDoesNotStartSharedCacheServer green by restoring from the store without a cache server), fail-first test first.
Next steps
- Owner of BuildTestEmbedded shared preparation: make the item-1 decision; complete as won't-do or port as above.
- Operator: authorize (or decline) a rebake to warm runner 0.0.119; coordinate with hf-2026-09-30-buildtest-restart-recovery-... since preparation time is the live problem there.
- Once both are dispositioned this handoff can be completed.
History (pre-2026-10-01 sections, stale)
RE-VERIFY (2026-09-24, sweep3 re-verification): The body below the 2026-09-24 section was written 2026-08-28 and is STALE: PR 747 no longer carries the single-shard change (it merged test-only), and the branch/SHA/image references describe state that no longer exists. Read "## 2026-09-24 re-verification and rescope" at the bottom first. Before acting, re-check gh pr view for https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1082 and grep current BuildTestEmbedded main for shardCount > 1 && cacheBuildRules.isNotEmpty().
Handoff: Review single-shard workspace cache reuse PR
Written 2026-08-28 05:43 UTC.
RE-VERIFY: This is a write-time snapshot. The authoritative checks at write time were git ls-remote origin refs/heads/feat/single-shard-workspace-cache-reuse, gh pr view for pull requests 714, 741, 747, and 722, and the locked local targeted test commands below. Re-run those commands before taking further action; current GitHub state and fresh test output take precedence.
Mission summary
Review and land the single-shard workspace cache reuse change in https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/747 once its required checks and review are handled. The source change is complete and pushed at 9edb6fcba717d08eb5fec607f79564a33dbd43eb; no merge, queue action, deployment, image change, or rebake was performed.
What was found and done
-
The former gate constructed SharedShardBuildCache only when shardCount > 1. The inline one-shard call to runShardPipeline also omitted the already-created cache object, so one-shard runs skipped restore, preparation, publication, and reuse.
-
The cache gate is now based only on whether the normalized build-rule set is non-empty. The single-shard pipeline receives the cache object, enters the existing shard-zero restore-or-prepare path, and publishes the same cache artifact. Single-shard log wording says for this shard; existing multi-shard wording is preserved.
-
Cache identity and validation were not changed. The existing fingerprint still covers the complete submitted archive, normalized build-rule key, runner version, kompile-manager version, and droplet image. Restore still checks the manifest, artifact size/hash, and valid .buildcache tar structure. Publication remains atomic and bounded by entry and byte limits.
-
No evidence showed that the old gate provided shard isolation or prevented cache poisoning. The single-shard test runs two separate service generations against fresh fake droplets and checks public logs, module-build count at the fake service boundary, and the persisted cache-entry directory.
-
The fail-first test was run on current main before the change with:
flock /tmp/buildtest-remote.lock scripts/test.bash --local --test crossRunModuleBuildCacheReusesSingleShard
Both builds completed with 2 passed, 0 failed, 2 total, but the assertion observed moduleBuilds=2, persisted entries [], and no public preparation or restore messages. This is the valid baseline; earlier empty-XML fixture failures and a same-coordinate stale artifact were corrected before recording this result.
-
After the fix, the focused test completed ALL TESTS PASSED (1/1 tests completed successfully). The existing multi-shard reuse, source-change invalidation, corrupt-artifact fallback, and toolchain-skew invalidation tests also completed successfully, each under the shared lock:
crossRunModuleBuildCacheReusesIdenticalInputs
crossRunModuleBuildCacheMissesWhenSourceChanges
crossRunModuleBuildCacheFallsBackFromCorruptArtifact
crossRunModuleBuildCacheMissesOnToolchainVersionSkew
-
The branch was fetched and rebased onto origin/main immediately before the final focused run and PR creation. It is clean, has no stash or local-only commit, and its remote head matches the local head.
Relevant PRs / refs
| Repo |
Branch (linked) |
Remote head SHA |
PR |
What is on it |
State |
| CodexCoder21Organization/BuildTestEmbedded |
feat/single-shard-workspace-cache-reuse |
9edb6fcba717d08eb5fec607f79564a33dbd43eb |
PR 747 |
Single-shard cache gate removal, pipeline wiring, artifact-coordinate bump, and fail-first regression test |
OPEN; local targeted tests green; required checks dispatched but the write-time snapshot reports both completed with failures; no merge |
| CodexCoder21Organization/PlanRepository |
wip/build-phase-root-cause-handoff-2026-08-26 |
d4e788d545a28c3ebffaeae4e382a8e009590164 |
no PR by handoff policy |
Earlier measured build-phase evidence retained for context |
Pushed documentation artifact; no new work in this lane |
Related pull requests, all write-time verified:
Deployed or published but not merged
This lane deployed nothing, published no production artifact, changed no production route, and rebaked no image. Image 242809349 already exists in sfo3 and nyc3, as recorded in PR 722; no action was taken on the image. The 0.0.68271805 coordinate is the source build coordinate used by the test harness, not a production deployment.
Next steps
- Re-check PR 747 before any branch update. The user explicitly prohibited CI watching in this lane; the write-time check snapshot is recorded above and must not be represented as green.
- Obtain the user's review/merge decision. Do not merge or enqueue automatically.
- If source changes are requested, check the PR is still OPEN, rebase on current
origin/main immediately before any new build/test, rerun the targeted cache tests under flock /tmp/buildtest-remote.lock, and update this handoff with the new remote SHA.
Reusable operational knowledge
rg is unavailable in this checkout; grep was used for repository searches.
- The shared test lock is
/tmp/buildtest-remote.lock; never bypass it. Remote execution initially hit a 30-second getBuildRun RPC timeout, so the fail-first and final verification used the locked local harness. This was infrastructure/harness behavior, not the cache verdict.
- The test's
@Timeout(120) is the same budget used by neighboring multi-node cache tests and was not changed during the fix. No test iteration budget was reduced.
- The test fixture uses an OS-assigned SSH port and cleans the service, fake droplet, temporary data directory, and generated key material in
finally blocks. It uses a fake implementation only at the infeasible external droplet boundary and executes real tar archive commands.
- The final source commits, in order, are
b3f113d0 (fail-first test), ef793794 (gate/log change), ee94d8e5 (artifact coordinate), and 9edb6fcb (single-shard cache wiring).
2026-09-24 re-verification and rescope
Remaining work (new mission): After https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1082 lands, decide whether single-shard runs still need cross-run workspace cache reuse under the new SharedShardBuildCache design. If so, port it onto then-current main: drop the shardCount > 1 condition on the cache gate, pass the cache object into the inline one-shard pipeline, and add a fail-first one-shard reuse test (the old crossRunModuleBuildCacheReusesSingleShard from branch feat/single-shard-workspace-cache-reuse at 9edb6fcb is a starting point). If the new design makes it moot, complete this handoff as won't-do with that reasoning.
Findings (verified 2026-09-24 ~02:25 UTC):
- https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/747 is MERGED (2026-09-23 18:54Z, merge commit 9d4f771f) but was repurposed during the 2026-09-23 old-PR cleanup: it is now titled "Cover quoted runner cache paths" and merged ONLY
tests/buildRunnerCommandExecutesSpecialCharacterCachePath.kts (+87). Its body says the one-shard production change was dropped because main had substantially changed that pipeline, and lists the port as remaining work.
- BuildTestEmbedded main 073b3523 still has the old gate at
src/buildtest/embedded/BuildTestEmbeddedService.kt:17511: if (shardCount > 1 && cacheBuildRules.isNotEmpty()). There is no crossRunModuleBuildCacheReusesSingleShard test on main. One-shard runs still skip the cross-run workspace cache; the original mission did NOT land.
- The same code is being rewritten right now: https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1082 (branch
fix/r218-shared-preparation-coordination, SharedShardBuildCache + shared preparation: followers await shard 0, single common compile; last push 2026-09-24 01:46Z by another live agent; checks red at the time) and https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1084 (touches the shared-cache merge; last push 01:19Z). Porting now would collide with a moving target.
- Rebaking the buildtest droplet image is a production action and requires explicit operator authorization; it is NOT part of this handoff's automated scope.
Supervisor ruling (2026-09-24 ~02:35 UTC): rescope in place (no duplicate handoff) rather than complete-as-superseded or drop. Whether single-shard cross-run reuse is still the right design can only be judged against the PR 1082 rewrite, so this handoff now depends on url://handoff/handoffs/hf-2026-09-23-fix-coordinator-shared-preparation-and-follower-waiting (the open handoff that owns branch fix/r218-shared-preparation-coordination / PR 1082) and stays BLOCKED until that lands.