Repository · handoffs
id: hf-2026-10-02-make-kompile-s-shared-caches-safe-for-concurrent-runs-reproduce-the-hang-stale-lock-replay-and-single-flight-defects-and-fix-them-upstream-so-no-campaign-needs-an-external-build-lock url: url://handoff/handoffs/hf-2026-10-02-make-kompile-s-shared-caches-safe-for-concurrent-runs-reproduce-the-hang-stale-lock-replay-and-single-flight-defects-and-fix-them-upstream-so-no-campaign-needs-an-external-build-lock title: Finish the kompile shared-cache campaign: land the shutdown-verdict, test-runner Clock-budget and diagnostic-recording PRs, then deploy the build-service engine summary: Remaining: (1) BuildTestEmbedded PR 1269: apply the proven caller-attribution patch from wip/A1269e-evidence-20261006, finish 16 ten-run gates, green CI, re-review, land. (2) testrunner.jvm PR 54 (all checks green; the 3 target tests pass on attempt 1): fix review RV54g's two P1s (currentTimeMillis deadlines; inner diagnostic token/helper not cancelled) and one P2. (3) kompile-core PR 380: fix two P2 timeline defects, investigate shard-2 30-second timeouts with its artifacts. (4) Close superseded BuildTestServerService PR 399; deploy engine 0.0.615130200607 (contains merged BuildTestEmbedded PR 1267; the live service still runs 0.0.615130200605). (5) Reproduce the recurring kompile-core interrupted-compile first-attempt failures. 51 PRs merged so far; no delegate running; all evidence on pushed wip/ branches. created: 2026-10-02T09:47:58.427Z completed: null dependencies:
Make kompile's shared caches safe for concurrent runs — finish the open fix PRs and deploy the build-service engine
Written 2026-10-08 08:50 UTC. RE-VERIFY before acting: everything below is a write-time snapshot. Re-check each PR with gh pr view <n> --repo CodexCoder21Organization/<repo> --json state,mergeStateStatus,statusCheckRollup,mergedAt, the live build-service jar with coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com routes --json (route whose image starts buildtest-server-), and published coordinates with a HEAD request on https://kotlin.directory/<group path>/<artifact>/<version>/<artifact>-<version>.pom.
Original request (verbatim): "Use codex 6.1-sol to drive url://handoff/handoffs/hf-2026-10-02-make-kompile-s-shared-caches-safe-for-concurrent-runs-reproduce-the-hang-stale-lock-replay-and-single-flight-defects-and-fix-them-upstream-so-no-campaign-needs-an-external-build-lock to completion - Use codex/gpt to fully investigate the issues … and drive fixes to completion." Owner requirement: shared caches must have fast, reliable locking internally so serializing builds externally is never necessary; where that does not hold, concrete failure reproductions are required. (The model named here is part of the original request, not a scheduling choice.)
Session state at write time: the orchestrating session ended around 02:10 UTC 2026-10-06 (its scratchpad was wiped). No delegate is running and no claim is held. Every delegate's work that matters is on the evidence branches listed below; anything not listed was not pushed and is lost (notably the in-progress round of the test-runner fix, see item 3).
Mission summary
51 pull requests merged across BuildTestEmbedded, kompile-core, kompile-buildscript, UrlProtocol and others over the campaign. What remains:
- Shutdown-verdict fix — https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1269 (head
edd39ea2, CI red on 3 tests with one known cause; the fix for that cause exists only as a patch on the evidence branch). - Diagnostic recording + kompile-buildscript 0.0.45 pin — https://github.com/CodexCoder21Organization/kompile-core/pull/380 (head
2863909e; review found 2 P2 defects; shard 2 red on 30-second timeouts). - Test-runner Clock-driven stop budgets — https://github.com/CodexCoder21Organization/community.kotlin.kompile.testrunner.jvm/pull/54 (head
3f8b6927, all checks green incl.kotlin.build (remote), but review found 2 P1 + 1 P2 still to fix). - Deploy the build service — engine fixes are merged and published but the live service runs an older jar (see "Deployed vs merged").
- Recurring kompile-core first-attempt failures (investigation started, no reproducer yet) and the parked PRs listed at the end.
What was found and done (the chain)
Build-service engine (BuildTestEmbedded)
- Merged: https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1267 (2026-10-06 00:55 UTC, merge
10327095): the service's start-up scan of legacy results rows reads at most the 16,777,216-byte row limit plus two buffers and rejects larger rows (message "uses more than 16777216 bytes"); preparation runs once per start-up; per-run status-row sidecars are written to a temp file, synced withFileDescriptor.sync()(notFileChannel.force, which throwsClosedByInterruptExceptionon an interrupted writer thread and lost the sidecar — found by the first real CI run after a GitHub Actions outage), then atomically moved. Also merged 1270, 1271, 1272 (test fixes). Superseded https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1264 was closed. - Engine releases: this campaign published
buildtest.embedded:buildtest-embedded:0.0.615130200606from main10327095(SHA-2568d63a13c81a29e6178f74d6841a44e43eddf6a7b794133f1ffb8545100c49e2b, 78/78 targeted tests; record on wip/REL1b-evidence-20261006). Meanwhile another session published…607(also contains 1267) and pinned it on BuildTestServerService main via https://github.com/CodexCoder21Organization/BuildTestServerService/pull/400 (merged 2026-10-06 22:06). Therefore this campaign's pin PR https://github.com/CodexCoder21Organization/BuildTestServerService/pull/399 (…606) is superseded and now conflicting (DIRTY) — close it.
Shutdown verdict — PR 1269 (open)
- Problem:
close()interrupts run workers; the SSH library (JSch 0.2.18) discards the interruption (Channel.java:824emptyInterruptedExceptioncatch;Request.java:61discards the sleep exception), so a worker's later ordinary failure was recorded as FAILED and the droplet deleted during shutdown. - Design in the PR: a typed "owner shutdown" state recorded on sessions the service itself closes, carried on one boundary around every
SshExecutorpublic operation; one shared decision keeps the run recoverable for owner shutdown and records FAILED for outside SSH loss. Two P1s from review RV1269c were fixed in headedd39ea2(channel-connect exit lacked the typed cause; shared-preparation shutdown retired the preparation owner —testSharedPreparationUnconfirmedServerStopSurvivesRestartnow 10/10). - Current red CI on
edd39ea2:sshSessionClearReportsCallerAndGeneration,sshExecutorDuplicateConnectPreservesSession,e2eSshCloseWinsDuringMissingSessionReconnect— the session-clear diagnostic's caller filter excluded only the exact classbuildtest.embedded.SshExecutor, so the new wrapper lambdas (SshExecutor$reconnect$2,SshExecutor$close$2) became the reported caller. - The fix is written and proven but NOT pushed to the PR:
investigation/A1269e/caller-fix.patchon wip/A1269e-evidence-20261006. Evidence there: all three tests 10/10 first attempt with the patch; restoring the exact-class filter fails all three 3/3; the rejected-exec, rejected-SFTP and shared-preparation probes pass once each after the patch. The candidate commit49f1f52bewas local only and is gone — re-apply the patch. - Still owed: the 16 ten-run gates listed as UNFINISHED in
investigation/A1269d/findings.mdon wip/A1269d-evidence-20261006 (each ran once and passed), rebase on main, green shards, an independent re-review of the final head. - Follow-ups recorded, separate PRs: upstream JSch interruption loss (above);
createBuildDropletinterrupted mid-write leaves a zero-byte ownership.tmpand clears the interrupt flag (10/10, evidence wip/RV1271-evidence-20261005).
Diagnostic recording + discovery fix — kompile-core PR 380 (open)
- Opt-in attempt timeline (one JSON line per attempt event when a timeline file is named; nothing printed by default), CI shard job records runner CPU/memory and uploads both as artifacts, per-invocation unique output directories in the launcher script.
- Its first CI selected wrong tests because published
kompile-buildscript0.0.43 dropped a selectedtestsroot's direct scripts. Fixed upstream: https://github.com/CodexCoder21Organization/kompile-buildscript/pull/95 (merged 2026-10-06 00:29), published askompile.buildscript:kompile-buildscript:0.0.45(SHA-256a337ec8e5e6d9c2f12e120560ddf04a2b4b26be6e2323b23ce5e2f019da9fa47; 0.0.44 is someone else's and still has the old discovery). PR 380 head2863909epins 0.0.45 in 948 declarations (the 0.0.33 bootstrap graph intentionally kept — it matches 0.0.45's embedded manifest); four shard lists disjoint, union = all 1,204 files; full local suite 1208/1208. - Review RV380 found two P2 defects (fail-first public tests on wip/RV380-evidence-20261006
investigation/RV380/tests): (a) with defaultmaxTestAttempts=1and overlapping file/name selectors, two real child processes get one attempt id becauseTestAttemptTimeline.attemptskeys only by test identity (expects 2 ids, sees 1, 3/3); (b) appending after a truncated last line joins the new suite-start JSON to the fragment (expects 1 parseable suite-start, sees 0, 3/3). Fix both, then re-review. - CI on
2863909e: shard 2 red, all five failures are the 30-second limit (testWarmInputContentReadsStayWithinMainBudget,testSourceRootBrokenLinkToFileLinkInvalidates,testExplicitBuildFailureKeepsFailedTestInResultsAndXml,testExplicitCacheHitPublicationFailureRemovesInternalCopy,testResultIndexPreservesObservableWorkspacePath) — the same timeout class mined in https://github.com/CodexCoder21Organization/kompile-core/blob/wip/CI375-evidence-20261005/investigation/CI375/REPORT.md (282 timed-out attempts in 20 runs, 164 hidden by retries; cause not established). The timeline artifacts this PR uploads are meant to explain exactly these — download them from that run before re-running anything.
Recurring kompile-core first-attempt failures (investigation, no PR)
testInterruptedCachedCompileCannotReturnBatch,testInterruptedCompileCannotReturnBatch("A real compile worker must open its staged source pipe") andtestWorkspaceHeadersOnlyOwnerReleasesWaiterOnExit(expected ExecutionException, got TimeoutException) fail first attempt in full parallel runs and pass on retry.- Lane FLKC1 (wip/FLKC1-evidence-20261006) captured the slow attempts' main threads in
UnixFileSystem.getBooleanAttributes0 -> File.isFileinside astagedFiles.filter, with ~17 s spent scanning a/tmpholding ~259,000 entries (128,676resolveDependencies-artifact-*, 93,449test-scratch*, …) — leftover temp entries from tests, which makes each attempt's ambient temp scan slow. Second candidate:CoursierProgressDeadline.schedule()runs under the same monitorclose()needs. No reproducer at ≥50% yet; nothing is fixed. Next: confirm whether the staged-file scan walks the shared temp directory (a per-run/owned temp root would remove the dependency on ambient temp size), build a deterministic reproducer with a pre-populated temp directory, and separately find which tests leak temp entries.
Test-runner — PR 54 (open)
- Launcher takes a process-table dependency, keeps enumeration outside its tree monitor, bounds queries per stop round, re-checks identity by start time before every signal; stop/cleanup budgets moved onto the injected
ClockviaClockBudgettokens; existing frozen-clock fixtures (64-identity, 514-process) changed only their clock drivers. On head3f8b6927the three previously retry-only tests passed on attempt 1 on the build service (run 6ae959d8) and 230/230 locally. - Review RV54g findings to fix (ranked-findings.md, probe tests alongside):
- P1:
OwnedProcessTree.drain(~line 355/369) andwaitByClockDeadline(~126–129) sampleclock.currentTimeMillis, so a backward wall-clock correction during TERM extends cleanup (probe passes onfb6478e7, fails on3f8b6927). Ruling given: arm deadlines through the Clock's scheduling (monotonic delivery, ManualClock-controllable), nevercurrentTimeMillisarithmetic or rawnanoTimein product code. - P1:
runDaemonTaskWithin(~line 89) returns after the outer 1000 ms expiry without cancelling the inner jstack wait (~line 1354); its ClockBudget token and helper survive. Ruling: the outer wait owns and cancels the inner token/helper. - P2: nothing protects ClockBudget's entry interruption guard — adopt the reviewer's two probes.
- Reviewer's third "P1" (nine assertions added to two existing tests) was ruled allowed — the rule forbids weakening/altering/removing assertions, not adding ones; state this in the PR description.
- P1:
- The fix round (T54d) was killed when the session ended before its first checkpoint; nothing from it survives.
Relevant PRs / refs
Open PRs of this campaign
| PR | Head | State | What remains |
|---|---|---|---|
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1269 | edd39ea2 |
shards 1/3/4 red (caller attribution) | apply caller-fix.patch, 16 gates, rebase, CI, re-review, land |
| https://github.com/CodexCoder21Organization/kompile-core/pull/380 | 2863909e |
shard 2 red (30-s timeouts) | fix 2 P2s, read timeline artifacts, CI, re-review, land |
| https://github.com/CodexCoder21Organization/community.kotlin.kompile.testrunner.jvm/pull/54 | 3f8b6927 |
all checks green | fix 2 P1 + 1 P2, re-review, land |
| https://github.com/CodexCoder21Organization/BuildTestServerService/pull/399 | b57bf208 |
DIRTY; superseded by merged PR 400 (…607) | close |
| https://github.com/CodexCoder21Organization/kompile-core/pull/375 | dce0c6b1 |
shards 1/3 red, 30-s timeouts | after the timeout cause is known |
| https://github.com/CodexCoder21Organization/kompile-buildscript/pull/93 | 1b247e12 |
remote check red on 30-s limits | parked |
| https://github.com/CodexCoder21Organization/kompile-buildscript/pull/91, https://github.com/CodexCoder21Organization/kompile-buildscript/pull/92 | fb5faf70, f055a041 |
remote check red (service degraded at the time) | re-request after deploy |
| https://github.com/CodexCoder21Organization/UrlProtocol/pull/653, https://github.com/CodexCoder21Organization/UrlProtocol/pull/654, https://github.com/CodexCoder21Organization/UrlProtocol/pull/648, https://github.com/CodexCoder21Organization/UrlProtocol/pull/640, https://github.com/CodexCoder21Organization/UrlProtocol/pull/651 | various | GitHub green, remote check red (service degraded) | re-request after deploy |
| https://github.com/CodexCoder21Organization/UrlProtocol/pull/646 | 46340212 |
remote green, informational check red; evicted twice earlier | re-enqueue after review |
| https://github.com/CodexCoder21Organization/kompile-executionenvironment/pull/82, https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1243 | 7e0c4754, a6485e00 |
kompile-cli 0.0.121 pins, red | after the timeout work |
| https://github.com/CodexCoder21Organization/BuildTestRunner/pull/145 | cf8a4058 |
remote green | review and land |
Evidence branches (all pushed; remote heads verified 2026-10-08)
| Repo | Branch | Remote head | What is on it | State |
|---|---|---|---|---|
| BuildTestEmbedded | wip/A1269e-evidence-20261006 | 70ca81370caf |
caller-fix.patch + 30/30 gates, mutation proof, probes | fix not yet on PR 1269 |
| BuildTestEmbedded | wip/A1269d-evidence-20261006, wip/RV1269c-evidence-20261005 | a8989fdbc19a, d0962a18613e |
P1 fixes evidence; review with UNFINISHED gate table | reference |
| BuildTestEmbedded | wip/REL1b-engine-release-20261006 / wip/REL1b-evidence-20261006 | bf93ca83ec3f / 82d4db3dfad2 |
…606 publication branch (coordinate-only diff) and record | published; superseded by …607 |
| BuildTestServerService | wip/REL2-evidence-20261006 | c0766f2d1ed5 |
pin PR evidence (133/133 targeted) | superseded |
| kompile-core | wip/RV380-evidence-20261006 | 4a266c1de702 |
review, two failing P2 tests | input to next round |
| kompile-core | wip/U380b-evidence-20261006 | 95a2bb3a5b58 |
0.0.45 publication + pin, suite XML with attempt histories | done |
| kompile-core | wip/FLKC1-evidence-20261006 | 27967f49a4ad |
thread captures, concurrent driver, diagnostic patch | investigation, no reproducer |
| community.kotlin.kompile.testrunner.jvm | wip/RV54g-evidence-20261006 | 43d17a690257 |
ranked findings + probe tests | input to next round |
| PlanRepository | wip/fable-briefs-2026-10-04 | 4b354e68bb9d |
all delegate briefs and status reports 1–199 (campaign-briefs-2026-10-04/) |
archive |
Deployed vs merged
- Live build service (ContainerNursery route for
buildtest):buildtest-server-e95a02b6-20261005-dep10.jar, built 2026-10-05 from engine0.0.615130200605(verified 2026-10-08 via the CN CLIroutes --json). It does not contain https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1267 or later engine fixes. - Published, not deployed: engine
…606(this campaign) and…607(another session; pinned on BuildTestServerService main by PR 400). Both contain PR 1267. Deploying …607 from BuildTestServerService main is a separate action that needs the owner's explicit ask (the owner delegated the deploy decision to the orchestrator on 2026-10-05; PR 400's author also flagged a heap-growth problem on the live coordinator that …607 is meant to address — coordinate with that effort before deploying). kompile.buildscript:kompile-buildscript:0.0.45is published from merged maina2c8a1d5— consistent with main.- Nothing else was deployed or published by this campaign.
Next steps
- Close https://github.com/CodexCoder21Organization/BuildTestServerService/pull/399 as superseded by PR 400.
- PR 1269: apply
caller-fix.patch(the filter must skip every frame ofSshExecutorand its nested/lambda classes, not exact names), rebase on main, run the 16 unfinished ten-run gates, push, green shards, independent review of the final head, land. - PR 54: fix the two P1s and adopt the P2 probes per the rulings above; full local suite; push; confirm the three targets pass on attempt 1 again; re-review; land.
- PR 380: fix the two P2s using the reviewer's tests; download this run's timeline/resource artifacts for the shard-2 timeouts and record what they show; push; CI; re-review; land.
- Deploy decision for the build service (…607 is pinned on main) — coordinate with the PR 400 effort; after a deploy, re-request
kotlin.build (remote)on the parked PRs and land the ones that go green. - kompile-core recurring first-attempt failures: continue from FLKC1 (temp-directory scan hypothesis, monitor-held deadline hypothesis); reproduce deterministically before fixing.
- Separate follow-ups: JSch interruption loss upstream;
createBuildDropletinterrupted-write.tmp; kompile-core 30-second timeout class (CI375 report) and the pins in PR 375.
Reusable / operational knowledge
- Read the build service's per-test attempt table, not the check badge: a green
kotlin.build (remote)can be earned by retries. The run page (https://buildtest.kotlin.build/run?id=<id>) marks retried tests withtest-row-flaky; its per-test view may take a while to materialise. FileChannel.force/ channel writes throwClosedByInterruptExceptionon an interrupted thread and close the channel; use stream writes +FileDescriptor.sync()when a writer must complete despite interruption.- When wrapping methods in a boundary lambda, re-run every diagnostic that inspects stack frames — the wrapper becomes the innermost frame.
- Version coordinates are shared with other sessions: re-check the registry for 404 immediately before publishing; never overwrite a claimed version.
- The build service was degraded during the whole campaign: delegates'
--remotefull-suite submissions were accepted but never scheduled. Briefs used a 15-minute acceptance window, after which the PR's GitHub shards were the gate. - GitHub Actions had a major outage 2026-10-05 ~19:50–21:32 UTC ("job was not acquired by Runner") that evicted queue entries; platform-cancelled jobs before any test ran may be re-run.
gh pr editfails on these repos with a Projects-classic GraphQL deprecation error; update PR bodies withgh api -X PATCH repos/<o>/<r>/pulls/<n> -f body=....gh pr merge --autois not allowed; enqueue with the GraphQLenqueuePullRequestmutation.- The /code box has ~10.7 GB cgroup memory; more than three concurrent BuildTestEmbedded builds caused exit-137 compiler kills. One build/test command at a time per lane.