← Priority list
In progress

BuildTest: "Out of date" site notice, coordinator memory, droplet utilization and slow build rules ��� paused; single projection half landed, utilization/build-rule/memory fixes open

Skeleton (1276) and fragments (1278) merged; 1277/1281/374 corrections and root wiring in progress; utilization PRs 1280/1283 open with mechanisms proven; four kompile compiler-worker drafts; memory PRs 1282/1273/1268 open; coordinator restarted 3x under the memory rule; session winding down on owner instruction with four lanes finishing.

Claimed by fable-code-2026-10-06c

Handoff document

Markdown

BuildTest: "Out of date" site notice, coordinator memory, droplet utilization and slow build rules — state at 2026-10-06 08:56 UTC (session paused on the owner's instruction; no lanes running)

Owner instruction at 08:27 UTC (verbatim): "Update the handoff with the latest details, including pushing any work-in-progress to github. Do not launch more delegates, but let the existing delegates finish on their own schedule and update the handoff with their details/results, let things wrap up gracefully and do not start next steps." Nothing new is dispatched after this point; the four lanes still running (listed under "In flight") finish on their own and their results are appended to this handoff as they land.

Mission

Active goal (verbatim): "Fix url://handoff/handoffs/hf-2026-08-24-fix-startup-compaction-delta-heap-exhaustion-then-revalidate-the-extreme-suite and bring droplet utilization to 100% specified in https://github.com/CodexCoder21Organization/BuildTestEmbedded/blob/main/docs/UTILIZATION_ARCHITECTURE.md and fix the projections to stay fast and up-to-date by fixing the underlying root cause of slow projections as specified in https://github.com/CodexCoder21Organization/DocumentationRepository/blob/main/playbooks/SLOW_SERVICES.md#fix-1-persist-expensive-projections-update-them-incrementally and https://github.com/CodexCoder21Organization/DocumentationRepository/blob/main/architecture/STATE_MANAGEMENT.md - maxinize parallelization of delegates (delegate to gpt 6.1-sol) and race PRs." Owner priorities stated during the work: the "Out of date" notice on https://buildtest.kotlin.build/ is the single most important thing; the coordinator must not run out of memory; fix causes, never mitigate; "We need to fix the real underlying issue (test lane utilization, tests not triggering timeouts), changing droplet limit is not resolving the underlying issue which is the real problem." Every production fix is reproduced first by a test that takes production's path (https://github.com/CodexCoder21Organization/DocumentationRepository/blob/main/playbooks/FIX_DID_NOT_HOLD.md ).

Terms

The coordinator (BuildTestEmbedded engine inside BuildTestServerService, route url:buildtest:) schedules test runs onto cloud machines ("droplets"). The site (BuildTestWui, https://buildtest.kotlin.build/ , route https:buildtest.kotlin.build:443) keeps its own stored copy (replica) of the coordinator's run list by reading the coordinator's change feed. The single projection is the owner-chosen design (decision 2026-10-05 ~23:00 UTC: "the single projection design is clearly superior, go with single projection design"): one persisted table of current rows in the coordinator, each stamped with the position of its last change; the feed answers "rows changed after P"; a reset is the same query from zero. Utilization is the share of a droplet's test slots busy while tests remain to run.

Artifacts of this session (all pushed)

  • Orchestrator artifacts (every brief, every lane's findings and final report, status reports r24–r48, decisions log, queue, tools, coordinator evidence incl. per-run test dumps): https://github.com/CodexCoder21Organization/BuildTestEmbedded/tree/wip/handoff-2026-10-06/orchestrator-artifacts/investigations/handoff-2026-10-06 — read decisions.log, reports/r47.md and the *-orchestrator-findings*.md first.
  • Frozen single-projection contract and the staged plan: dz2-contract.md (with amendment A1: identity elements/runId ≤ 16,384 UTF-16 units, escaped rowKey ≤ 65,536 bytes, so every admissible row fits the 983,040-byte page cap) and dz2-fastpath.md in the same directory.
  • Finished lanes' local-only branches snapshotted under wip/handoff-2026-10-06/<lane>/… in BuildTestEmbedded and BuildTestWui (review probes, mutation variants, evidence checkpoints). Each lane also pushed its own wip/<lane>-2026-10-0{5,6} branch.
  • Previous session's artifacts: https://github.com/CodexCoder21Organization/BuildTestEmbedded/tree/wip/handoff-2026-10-05b/orchestrator-artifacts (background only).

Merged this session

  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1276 (04:24 UTC) — single-projection store skeleton (SQLite current rows, packed payloads, group-committed admission/publication, byte-capped pages).
  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1278 (07:51 UTC) — fragment delivery for oversized rows, convergence while maintenance runs, amendment A1 admission bounds, interrupted-first-open recovery.
  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1263 , https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1255 (2026-10-05) — on-disk classification index; late-verdict merge.
  • https://github.com/CodexCoder21Organization/BuildTestServerService/pull/398 (2026-10-05, deployed) ; https://github.com/CodexCoder21Organization/BuildTestWui/pull/371 and https://github.com/CodexCoder21Organization/BuildTestWui/pull/367 (deployed to the site 01:17 UTC from main 117d528a). No engine has been published and nothing from 1276/1278 is deployed: the single projection is not wired into the running coordinator until https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1281 lands.

Open pull requests — state, what is proven, what remains (heads verified 08:30 UTC)

Single projection (engine) — the "Out of date" fix

  1. https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1277 — COMPLETE coverage, RUN_TOMBSTONE, build-rule aliases, atomic replacement. Head 647dda20 was green (152/152, 4/4 mutations) and passed the orchestrator's final review (notes: 1277-orchestrator-findings-b.md). The final-head reviews (ra1277b, rt1277b) found 2 BLOCKING (grouped reduction misses an alias prepared in the same group; a late legacy alias arriving after its terminal execution is shown live), 1 SHOULD-FIX (COMPLETE re-tombstones already-removed rows), 4 test gaps. Correction lane spd6 is IN FLIGHT (174-scenario gate running, mutations 3/3 red); it rebases onto main (1278 landed) and pushes. After it pushes: re-review the correction delta, orchestrator final review, enqueue. Needs a rebase after 1278 in any case (DIRTY).
  2. https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1281 (draft) — startup enrollment of every writer (rules 5.1–5.3 met, 20 tests) + the real root wiring. Head c25d157a; DIRTY; CI red (expected, mid-rework). Decisions already made: the root module COMPILES the projection module sources directly (community.kotlin.buildtest.projection.api/src and …embedded/src as extra source roots; only SQLite added to the POM) because every other route (published artifact, buildLocalArtifact bridge) made the root artifact's POM reference an unpublished artifact that all existing root tests fetch; the SQLite BUSY between catalog and store (two DEFERRED read-then-write handles on current-v1.sqlite) is fixed by taking write ownership before reads (BEGIN IMMEDIATE) in ProjectionStoreTransactions.kt and the catalog. Lane sp6 ended at its time box with: 6.4 UNMET (the proof that a new reader survives the legacy retention ceiling — this is the proof that the original notice mechanism is gone), SQLite fix on wip only (wip/sp6-2026-10-06), gate 1/25, carries 1277's head 647dda20, DECISION NEEDED in its findings (read findings/sp6-findings.md). Remaining: rules 6.1–6.4 as written in dz2-fastpath.md ("PR 6"), the SQLite fix moved into the PR, rebase onto main + 1277, gate, reviews, final review.
  3. Site: https://github.com/CodexCoder21Organization/BuildTestWui/pull/374 — durable schema-3 replica with partial-row storage, epoch switch only on a complete copy, bounded scalars, one connection per page. Head b20db59e (lane sp7f's push), CI GREEN at 08:10; earlier head 1f5aecbb green. sp7f closed the final-head review's BLOCKING (a malformed final fragment — duplicated optional key — was accepted and replaced the good display) and should-fixes (numeric grammar before conversion; EOF after the root JSON value; exactly-once close; observer re-entry), and the test gaps (incl. "the run tombstone alone drops every child row" — the engine hides a run's test/rule/utilization rows on RUN_TOMBSTONE without per-row tombstones). Remaining: re-review of the sp7f delta against b20db59e (findings/sp7f-findings.md), orchestrator final review, enqueue. Then the site reader (plan "PR 8", brief briefs/spe.md, not started) which actually switches the site to the new feed.
  4. Release order once 1277, 1281 and 374 + PR 8 are merged: publish the engine → pin in BuildTestServerService → deploy coordinator (measure time-to-ready against the supervisor's 900 s allowance and disk high-water on a data copy first) → deploy the site → confirm outOfDate=false stays and retained memory stays flat. Rollback per dz2-fastpath.md Q4.

Utilization (owner's runs 1fcb4d1a, 5baf542a, c7c595aa, 5bf1b96a)

Measured (dumps in coordinator/*-all.json): 2.7 and 6.7 tests in flight against ~40 slots; rows "RUNNING" for hours with end times recorded. Mechanisms proven (findings/ut1-findings.md, ut2-findings.md): (M1) follower droplets expired at the provider's 1-hour lease because the watchdog sweep renews serially and was parked on whole-snapshot persistence — only the primary was renewed; (M2) completed non-authoritative failures stayed publicly RUNNING with no retry verdict; (M3) each test completion ran an fsync-bearing attempt-history reconciliation before the scheduler got its slot back (median dispatch-to-start 18 s, max 310 s); (M4) coordinator-side per-test deadline enforcement was deliberately unarmed by a staged rollout and pinned by tests/e2eCoordinatorDeadlineRecyclesOnlyExpiredShard.kts (its messages say the arming step "is not included"). 5. https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1280 — M1–M3 (new ActiveDropletLeaseRenewals per-droplet timers; bounded retry decision with prioritizeFailedRetries; reconcileAttemptHistoryBeforeReturn=false on the dispatch path). Head 05c77df7; 369-file matching selection ran; CI RED on one existing test reviewOccupancyIngestionKeepsOwnedStoragePolicy ("Refused ingestion must emit its warning"): the storage-refusal warning was emitted by the reconciliation M3 moved off the dispatch path — a real behaviour change; the refusal must still be reported at the history-materialization boundary (review rt1280 recommends exactly that). Orchestrator review (1280-orchestrator-findings.md): O1 close() joins provider calls without a bound; O2 first renewal after restart is at now+interval (state the margin); O3 the retry authority must not read the now-lazily-reconciled index; O4 retries jump the queue (bounded by the retry target). Reviews ra1280 and rt1280 finished (read them). Also to fold in: the sibling-renewal defect (siblings of a droplet whose deletion is queued are not renewed; failing reproducer on wip/ut3-2026-10-06) and the 5bf1b96a evidence (37 rows with end times still RUNNING; see ut1-orchestrator-notes.md). Correction round not started. 6. https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1283 — M4 deadline arming: deadlines registered after durable dispatch recording and before queue publication; expired test recorded as timed out, slot recycled at timeout + grace. Six assertions of the pinning test replaced under the owner's approval ("You are approved to change e2eCoordinatorDeadlineRecyclesOnlyExpiredShard.kts using your best judgement"; before/after in findings/ut3-findings.md). 403/403 first attempts; two mutations red. CI running at 08:30 (watch: out/watch-1283.log in the artifacts). Reviews not run (briefs briefs/ra1283.md, rt1283.md ready). 7. Stale site rows for run 5bf1b96a (owner question 08:0x): the coordinator's run.json says 1,880 passed / 4 failed / 1,891 total while the site shows 1,853 passed / 0 failed / 37 RUNNING — the site's projection is stale per row (the retried attempts' terminal verdicts never reached the replica) even though its freshness flag is false. Lane sv1 is IN FLIGHT on the exact mechanism (brief briefs/sv1.md; hypotheses: the projection-change for a retried attempt is not emitted/keyed for the row; the replica ignores the authoritative verdict; a reset replaced the rows). Its result is appended below when it lands. 8. Not started: droplet-release correction https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1274 (brief f1274b.md: findings M1–M4 of 1274-orchestrator-findings-b.md), compact run record https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1275 (brief env1b.md), dispatch-estimate stage 2 (est1b.md; needs owner approval to change one existing test's work-count expectations), runner stage mode switch to PIPELINE in production after the runner change lands.

Build rules (owner's run c7c595aa: 11 min "build rules")

Measured (findings/br1-findings.md, design compiler-lifetime-design.md on wip/br1-2026-10-06): each cache-missing rule starts a fresh build-script JVM with cold compiler state — eight support-module compilations cost 90.3 s in fresh JVMs vs 12.9 s in one warm JVM; the two-at-a-time limit is memory admission; the run page's "CPU time" for rules is summed wall time (rule CPU is not recorded). Decisions: warm compiler workers owned by the workspace, isolated build-script children kept (authorized); standalone/workspace-less rule calls get a temporary per-invocation compilation owner (option A); helper processes owned as Linux process groups (taken 08:00 UTC, not yet briefed). 9. Four cross-linked DRAFT pull requests from lane br3: https://github.com/CodexCoder21Organization/kompile-core/pull/381 , https://github.com/CodexCoder21Organization/kompile.executionenvironment.bridge.interface/pull/30 , https://github.com/CodexCoder21Organization/kompile-executionenvironment/pull/83 , https://github.com/CodexCoder21Organization/build-kotlin-jvm/pull/71 . Counting invariant PASSES (16 modules → one worker, W=1). Five of seven ownership rows proven; remaining: helper-process cleanup + memory-lease boundary (process groups), full core/JVM suites, consumer pin adoption (BuildTestRunner admission accounts the shared worker once), publishing. Read findings/br3-findings.md and finals/br3-final.md.

Coordinator memory

Root cause chain (heap dump 01:08 UTC, mem4 findings in the 2026-10-05b artifacts): ~1.1 GB of the journal's parsed tail and baseline-cache drafts retained in heap. The coordinator is restarted under a rule (retained after two consecutive full collections > 2,594 MB); restarts this session at 00:13, 02:13, 04:51 UTC (the last because getFleetStatus timed out at 30 s with old generation 98% and a full collection every 30 s, producing the site's error notice). Retained 1.9–2.1 GB at 08:20; a further restart is likely; watch: jstat -gcold <pid> (procedure in tools/retained-watch.py, restart command below). 10. https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1268 — parsed tail bounded. 367/368 at the f1268g product branch (wip/f1268g-product-2026-10-06, includes the review's R1 adoption fix). The one red is the main test stressProjectionChangeFeedReadDoesNotWaitForCommittedRewrite (asserts resetRequired=true as evidence of commit; S2 keeps the reader's cursor). Owner approved 08:12 UTC changing that test to observe the commit via the publication observer and assert reader continuation. Lane f1268h is IN FLIGHT applying it (brief briefs/f1268h.md). The CI OutOfMemoryError in tailAccelerationSurvivesDeathOfOriginalRunnerE2E did not reproduce in four runs; peak heap is the embedded compiler's disposer bookkeeping; flake brief fl1.md not started. 11. https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1282 — pre-existing defect: closing the journal while a durable rewrite is pending leaves it over its row bound (main first-fails 41 > 25). Reproducer red→green 3/3, soak 3/3, 347/347. CI RED on one existing test projectionStartupFoldGenerationSwitchKeepsReadersAndSidecarConsistent ("the folded generation exists only in memory; journal.jsonl must keep the committed generation") — the restore-on-open fold must persist the folded generation before readers/sidecar see it. Review ra1282 IN FLIGHT; rt1282 finished (6 gaps before landing; read it). Correction round not started. 12. https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1273 — baseline cache keeps no drafts (A1 count 0 at both sizes; gate 99/100, the soak being 1282's defect). Rebase onto 1282 once it lands, re-gate, reviews. 13. https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1279 — coordinator shutdown race found by the same work; green; reviews not run. 14. https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1138 (cache seeding) waits for an upstream kompile-core fix (brief kup1.md: string-path reads untracked for rule reuse; input records stored by absolute path); https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1254 (hold mechanism) is superseded by the single projection — close it when PR 8 exists.

Were in flight at 08:35 UTC — all four have since finished; see "Appended results" at the end

  • spd6 — PR 1277 second correction (started 07:21; 174-scenario final gate running). Branch wip/spd6-2026-10-06; pushes to spd/complete-coverage-tombstones-aliases-and-atomic at its gate.
  • ra1282 — adversarial review of PR 1282 (started 07:33). Findings → out/ra1282-findings.md (artifacts branch refreshed when it lands).
  • sv1 — stale-row mechanism for run 5bf1b96a (started 08:17; brief briefs/sv1.md). Branch wip/sv1-2026-10-06.
  • f1268h — PR 1268 landing round with the approved test change (started 08:20, by the queue refill a moment before the wind-down instruction arrived). Pushes to the PR branch at its gate. Queued but NOT started (briefs in briefs/): ra1283, rt1283, fl1, kup1, f1274b, env1b, spe, est1b, f125d.

Next steps when the work resumes (not started, in order of value for the notice)

  1. Land 1277 (after spd6's push: re-review, final review, enqueue) and 1281 (finish 6.1–6.4, move the SQLite fix in, rebase, gate, reviews). Land 374 (re-review sp7f's delta) and build PR 8 (site reader). Release engine → deploy coordinator → deploy site → confirm the notice stays away.
  2. Land 1280 (correction: refusal warning at the materialization boundary, O1–O4, sibling renewal, 5bf1b96a test), 1283 (reviews), the sv1 fix.
  3. Land 1282 (fold persistence + rt1282 gaps + ra1282 findings) → rebase 1273 → land 1268 (after f1268h) → 1279. Then coordinator restarts stop.
  4. Finish the four kompile drafts (process-group ownership, suites, pin adoption, publish) → BuildTestRunner admission → measure build-rule time on a production run.
  5. 1274, 1275, est1b (owner approval needed for one existing test's expectations), kup1 → 1138; switch runner stage mode to PIPELINE; rerun the extreme suite; complete this handoff.

Operational knowledge

  • Site freshness: curl -s 'https://buildtest.kotlin.build/api/runs?limit=1' (outOfDate, refreshFailing, upToDateAsOfEpochMs); health incl. fleet: /api/health (fleet.asOfEpochMs staleness = the coordinator not answering getFleetStatus); per-run rows: /api/test-results?id=<run>&page=N&limit=100 (limit ≤ 500; 503s under load — retry); build rules: /api/build-rules?id=<run>; phases: /api/run-timing?id=<run>.
  • Coordinator process on the host: pgrep -f buildtest-server-.*dep10.jar (the parent, not the -XX:ActiveProcessorCount=1 test-bundle child); memory: jstat -gcold <pid>; restart: cd /tmp && ~/bin/cs launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com restart --route-key 'url:buildtest:' ("Failed: Unknown response format" = success; new process in ~40 s). Logs: /root/ContainerNursery/data/container-logs/buildtest-server-*.stdout.log and buildtest-wui-*.stdout.log. Run records: /root/buildtest-data/runs/<run>/run.json, test-events.jsonl.
  • Host access: ssh -p 23 -i ~/.ssh/id_ed25519_new root@198.199.106.165, read-only except the authorized actions (site jar upload with a rollback copy, coordinator route restart). Never print route listings or lines containing token/secret/password/key. The fleet dropletLimit reads 200 (no lane changed it; the owner cancelled the limit-change lane).
  • Build/test capacity on the orchestration box: two local slots (tools/run-slot.sh); named tests of a committed head run on GitHub machines in ~5–45 min via tools/gha-run.sh (BuildTestEmbedded) / evidence/gha-run-wui.sh on wip/e2e1-2026-10-05 (BuildTestWui). --remote is the production service under repair — do not use it.
  • Delegation discipline that mattered: stop the armed refill before killing a lane by hand (it refills on the kill); run the orchestrator's final review as soon as a head is final, in parallel with delegate reviews; every PR that touches a test on main needs the owner's approval for the assertion change (two were granted this session, quoted above).

History (condensed; details in the 2026-10-05b artifacts and reports r24–r48)

  • 2026-10-05: mechanism of the notice established (coordinator writes ~50 records/s, retains 32 MiB for a lagging reader, the site's snapshot read takes 15–60 min → "reset required" loop); first-baseline and startup-cost MUST-FIX items; site lock fix deployed; owner chose the single projection; the eight-PR fast path started 23:58 UTC.
  • 2026-10-06 00:06: the build machine's disk filled and killed every lane; checkouts were deleted without the pushed-work check (lesson recorded; tools/reap.sh added). 01:08 heap dump → memory root cause. 01:17 site deploy. Coordinator restarts 00:13 / 02:13 / 04:51. 03:xx–08:30: everything above.

Appended results of the finishing lanes

  • 08:36 UTC — spd6 finished (PR 1277 second correction): pushed head 5944b510 to spd/complete-coverage-tombstones-aliases-and-atomic, rebased onto main 040fb44b (after 1278), description updated. B1 (same-group alias lookup), B2 (late legacy alias after terminal execution), S1 (no re-tombstone of removed rows) and test gaps G1–G4 all met; gate 176/176 first attempts (152 original + 12 new + 10 merged-fragment controls + 2); the 16 new/changed scenarios three passes each; mutations 3/3 red. PR CI on 5944b510 was running at 08:40 (https://github.com/CodexCoder21Organization/BuildTestEmbedded/actions/runs/37436857754 ). Remaining for 1277: re-review of the correction delta 647dda20→5944b510 (brief pattern: briefs/ra1278b.md), orchestrator final review of that delta, then enqueue.
  • 08:24 UTC — ra1282 finished (adversarial review of PR 1282 at 7ccc3c00): 2 BLOCKING / 1 SHOULD-FIX / 1 NOTE — not landable. B1: a recovery write failure during the restore-on-open fold discards valid committed journal state (ProjectionChangeJournal.kt:5332, open handler 4851–4858, empty replacement 5670–5673). B2: one expired status deletes unrelated baseline test identities (:5291, :5294, normalized fold 5313–5317). S1: startup restoration work scales with the retained baseline rather than the overflow (:5295, :5305, :5322). Reproducers on wip/ra1282-2026-10-05/investigations/ra1282. Together with the CI failure (projectionStartupFoldGenerationSwitchKeepsReadersAndSidecarConsistent: the folded generation must be persisted to journal.jsonl) and rt1282's six gaps, the correction brief for 1282 should state the invariant: restore-on-open is a durable, failure-atomic rewrite of only the overflow, applied to the journal file before any reader or sidecar sees the folded generation, never deleting rows outside the folded range; a failed rewrite leaves the pre-fold journal intact.
  • 08:43 UTC — sv1 finished (run 5bf1b96a "37 running on one droplet"): no product defect in the feed or the site was demonstrated; this CORRECTS item 7 above. The 37 RUNNING rows are tests whose physical attempt had ended (endedAt set) and which were waiting for a retry; the logical row stays RUNNING until an authoritative verdict. The four counted failures were decided by the coordinator between 08:09:55 and 08:11:36 — after the site response captured at 08:06 and before the run.json read at ~08:14 — and for both rows traced, the terminal update was emitted and stored by the site at its exact feed sequence. Hypotheses "update not emitted" and "replica ignores the verdict" are refuted; "a reset replaced the rows" is unproven because the 08:06 capture kept no cursor. Evidence: https://github.com/CodexCoder21Organization/BuildTestEmbedded/blob/184af6353ae2fbcb558b8f2979a58caf32ffe95d/evidence/sv1-findings.md . What remains true and is the real defect: rows waited 35–168 minutes for their retry on a run reduced to one droplet — that is the utilization problem (lost followers, slow retry decision), owned by https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1280 . A display follow-up worth considering: the site labels "attempt finished, awaiting retry" as RUNNING, which reads as 37 tests executing.
  • 08:45 UTC — CI on https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1283 (deadline arming) is RED at cede66cc: tailAccelerationStopFiresFromWriterWhenNoRunnerSendsAnotherRecord and tailAccelerationNeverStopsShardOwingBuildResult fail with "The idle shard never received a speculative copy of the straggler" — the dispatch log shows memory_skip (predicted peak 2,147,483,648 bytes exceeds the droplet) and the straggler recorded as an authoritative FAILURE with maxTestAttempts=1. Not diagnosed. First question for whoever resumes: do these fail on the 1280 branch alone (then it is M2's prioritizeFailedRetries / retry finalization changing tail-acceleration behaviour) or only with the deadline arming on top; reproduce locally before touching CI.
  • 08:44 UTC — f1268h (PR 1268) reports a dependency, lane still finishing its mutation proofs: with the owner-approved test change applied, its gate first-fails the unchanged soak at 49 rows against a bound of 25 — the pre-existing defect that https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1282 fixes. So 1268 cannot go green before 1282 lands (same as 1273). Order: 1282 (after its correction round) → rebase 1273 and 1268 → gates.
  • 08:55 UTC — f1268h finished (PR 1268 landing round); ALL lanes have now exited. The owner-approved test change is applied and the review's adoption fix (R1) is in; approved test and R1 probe 3/3; both mutations red; gate 367/368 — the one red is the unchanged soak (49 rows against its 25-row bound), i.e. the defect https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1282 fixes. NOT pushed to the PR branch (still 2f1a11a2). The ready head is wip/f1268h-product-2026-10-06 @ d78a21cf (findings wip/f1268h-2026-10-06). When 1282 lands: rebase d78a21cf onto main, rerun its gate, push to the PR branch, then reviews.
  • State at hand-over (08:56 UTC): no delegate running; nothing queued was started; CI on https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1277 head 5944b510 still running (four shards pending); coordinator retained memory 2,276–2,479 MB against the 2,594 MB restart rule and NOBODY IS WATCHING IT after this session — check jstat -gcold <pid> and restart per "Operational knowledge" if the site shows an error notice or /api/health fleet data goes stale. Site reports outOfDate=false.

Latest status report

3 reports
Running fable-code-2026-10-06c 50fdf0b6 13 minutes ago

Status report r125 — 2026-10-08 02:01:25 UTC

Original request (active /goal, verbatim): "Finish url://handoff/handoffs/hf-2026-08-24-fix-startup-compaction-delta-heap-exhaustion-then-revalidate-the-extreme-suite and then get droplet utilization to 100% with consideration of https://github.com/CodexCoder21Organization/BuildTestEmbedded/blob/main/docs/UTILIZATION_ARCHITECTURE.md and also solve the underlying root cause of any projection update slowness with consideration of https://github.com/CodexCoder21Organization/DocumentationRepository/blob/main/playbooks/SLOW_SERVICES.md#fix-1-persist-expensive-projections-update-them-incrementally and https://github.com/CodexCoder21Organization/DocumentationRepository/blob/main/architecture/STATE_MANAGEMENT.md - Parallelize work assigned to delegates and race PRs where possible and deploy after merges. Flaky tests encountered during the work should be diagnosed and fixed using the flakey test skill."

🟡 Status: RUNNING — three questions for you; the rest of the work continues without them

Q1. May I restart the kotlin-build-ci container on ContainerNursery, with no new code?

  • Background: kotlin-build-ci is the service that runs the kotlin.build (remote) checks. It is half-started: its startup thread died, so its reconciliation loops never began. Merge queues across repositories are therefore stalled; for example, https://github.com/CodexCoder21Organization/UrlResolver/pull/1213 has not moved from queue position 11 in over 2 hours.
  • The catch: the code fix, https://github.com/CodexCoder21Organization/kotlin-build-ci/pull/301, cannot get its own kotlin.build (remote) check while the service is broken.
  • Options:
    • (A) Restart now with no code change. Small and reversible: if the restart succeeds, checks flow again everywhere.
    • (B) Deploy 301's branch before it merges. This ships unmerged code.
    • (C) Merge 301 on a full remote test run standing in for the check, then deploy on your approval.
  • Recommendation: A.

Q2. Approve one changed assertion in https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1295?

  • projectionAdvertisedResetRecoveryIoFailureRequiresRepair would now expect writing committed row to the advertised reset failed: java.io.IOException: No space left on device. On main it expects org.json.JSONException: Unable to write JSONObject value for key: run.
  • It is still a full-text match and still checks the root disk-full error.
  • Recommendation: approve, because the new message names the real cause.

Q3. Approve one changed assertion in https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1300?

  • stressPaginatedReadsDuringDurationProjectionRebuildKeepsDurationOrdering requires exactly 32 source rows to be read by 16 concurrent page reads. The PR's shared page cache lets those reads share one load, so only 2 rows are read.
  • Proposed: require between 1 and 32 rows instead.
  • Mutation evidence: with the change, the test still fails on each defect it guards:
    • a stale cached page after the ordering changed;
    • ignored duration averages;
    • a whole-snapshot rescan (652 rows read).
  • With the change, it passes 3 runs in a row.
  • Recommendation: approve. The assertion counts a cost the PR legitimately removed.

Progress since r124 (01:41)

  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1295: both regressions are fixed and pushed (head 465006df).
    • Cause: at startup, every journal row was copied into a temporary holding stage, and each copy also maintained an on-disk index that nothing ever reads, one disk read or write per node.
    • Fix: the temporary stage now skips that index.
    • Result: the two timed-out tests now take 23–29 s. Gate: 28/28 tests on 3 runs, with no retries.
    • Next: CI and a fresh Opus review of the new head are running.
  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1300: two of the three new failures are fixed in product code, and the gate passed 22/22 three times.
    • First fix: the snapshot writer built its page index by re-reading the ~8 MB file twice; it now builds it from the rows it writes.
    • Second fix: a small file's malformed-row error message was restored to main's.
    • Third failure: needs the Q3 decision. The work is saved on a branch and not yet pushed.
  • https://github.com/CodexCoder21Organization/BuildTestServerService/pull/402: its last red test is explained.
    • Mechanism: when the engine library's service fails during construction, its cleanup waits on a background thread with the interrupt flag set. That wait throws immediately, so cleanup never releases the control-directory lock, and the replacement service then fails with "already has a writer in this process".
    • The fix is only in https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1297.
    • 402's path: 1297 lands, then an engine version is published from main, then 402's dependency pin is bumped to it.
  • https://github.com/CodexCoder21Organization/kotlin-build-ci/pull/301: both new tests (one where recovery throws, one where it throws an Error) fail on the old startup order and pass on the new one. A mutation that trips a tripwire is now caught (6/6 red), where before 4 of 6 tests swallowed it. The full remote suite is running.
  • 1292: shards 2–4 are green; shard 1 is still running at ~42 minutes, which is typical.
  • Coordinator memory: retained heap is 1.3–2.1 GB, under the 2.6 GB rule; no restart needed.

PRs

✅ Merged

  • https://github.com/CodexCoder21Organization/PlanRepository/pull/10741 (challenge record, earlier this session)

⏳ Pending (verified 02:01)

  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1292 — OPEN; review clean; shard 1 of 4 still running
  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1295 — OPEN; new head 465006df; CI and review running; needs Q2
  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1297 — OPEN, CLEAN at 07e8b08d; being superseded by the Opus fix
  • https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1300 — OPEN; fix on a branch; needs Q3
  • https://github.com/CodexCoder21Organization/BuildTestServerService/pull/402 — OPEN; waits on 1297, an engine publish and a pin bump
  • https://github.com/CodexCoder21Organization/UrlResolver/pull/1213 — in the merge queue at position 11; stalled behind the CI service (Q1)
  • https://github.com/CodexCoder21Organization/kotlin-build-ci/pull/301 — OPEN, CLEAN at bfd9eaa7; fixes in progress
  • https://github.com/CodexCoder21Organization/kotlin-build-ci/pull/302 — OPEN, red; mostly superseded by 301

Delegated work at a glance

⏳ Running - Opus - ~12 minutes total time - 5 minutes since last update - Adversarial review of https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1295 at 465006df. It is checking that the temporary stage with the index switched off can never be published, persisted or read for classification, and that reordering the hooks breaks no caller. Two targeted runs are in progress. Next: a verdict.

⏳ Running - Opus - ~40 minutes total time - 4 minutes since last update - Takeover of https://github.com/CodexCoder21Organization/kotlin-build-ci/pull/301. The new tests fail before the fix and pass after it, and the tripwire mutation is now caught. The full remote suite is running from a0910b6d. Next: push to the PR.

⏳ Running - Opus - ~20 minutes total time - 3 minutes since last update - Finishing https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1297 from the round-12 work branch.

  • It found that a "compile failure" in one baseline run was a defect in its own test runner, not in the code: the GitHub checkout has no git history. It relaunched with bundled sources.
  • The re-entrant-close test passed first time on 3 of 3 runs.
  • It was told that 402 depends on this PR's constructor-cleanup fix.
  • Next: diagnose the expanded-set failures.

⚪ Not running (finished) - Opus - ~27 minutes total time - Fixed 1295's two regressions and pushed 465006df. Nothing further expected.

⚪ Not running (finished) - Opus - ~26 minutes total time - Fixed two of 1300's three failures. The third is a proposed test change awaiting Q3; it will resume on your answer.

⚪ Not running (finished) - Opus - ~3 minutes total time - Diagnosed 402's last red test as the engine's failed-construction cleanup, fixed only in 1297. No changes made. Nothing further expected.

How the fixes work

  • 1295: journal rows are edited as streamed text, and the journal index is paged to disk, so heap use stays flat as the journal grows. The temporary holding stage used during startup now skips the on-disk index it never reads, which makes startup fast again. Only a clean stop counts as a stop; any real failure is still reported.
  • 1300: concurrent page reads share one cached load, and warm reads reuse decoded rows, so a one-result read no longer reads the whole 5 MB result file. The writer now builds each snapshot's page index from the rows as it writes them, instead of reading the file back. A malformed small file reports the same error as on main.
  • 301: startup recovery of saved builds runs on a background thread, started only after every reconciliation loop is running. A recovery step that throws or hangs can no longer leave the service half-started. Each failure now logs its stack trace, and test tripwires are no longer swallowed.

🟡 Blockers

  • Q1–Q3 above, which need your answers. They gate 1213 landing, 1295 landing and the 1300 push respectively.
  • Infrastructure: the half-started CI service, which is Q1.

Plan changes

  • 402 now has a real dependency on 1297 (code that only 1297 contains), so it waits for that rather than for a sequencing preference.
  • Self-check: genuine progress this cycle — two fixes pushed or proven, one root cause found, and three gates turned green.

Landing parallelization sweep: all landing gates are already running in parallel. 1292 is waiting only on its last shard. 1295's review and CI are both running. 301, 1297 and 1300 each have a fixer working or are waiting on your answer.

Next steps: land 1292 when its CI is green; land 1295 after its review, CI and your Q2 answer; push 1300 after Q3; land 301 and 1297, then publish the engine and bump 402's pin; unstick the CI service per Q1; then revalidate the extreme suite, work on droplet utilization and fix the root cause of projection slowness.

Top lessons

  1. When a check stops reporting across a whole organization, the service that dispatches it may be half-started: it serves requests, but its background loops never began. Check its thread list for those loops before blaming individual PRs. Run fragile recovery work only after the loops are installed, and off the startup thread.
  2. A scratch or temporary data structure should not carry the bookkeeping of the published one. Here an on-disk index nothing ever read was maintained for 400,000 temporary rows, and the cost appeared as timeouts in unrelated-looking heap tests. When a slowdown shows up, profile which writes are actually read back.
  3. Before calling a test failure "pre-existing", confirm the baseline you compared against is really main and not an earlier head of the same PR. A lane here mislabelled two real regressions for exactly that reason.

Earlier reports

Running fable-code-2026-10-06c 50fdf0b6 32 minutes ago
Running fable-code-2026-10-06c 50fdf0b6 53 minutes ago

Add dependency

Complete this handoff

Moves it out of every priority list and into ArchiveArea.