← Priority list
Unclaimed

Land the remaining BuildTestEmbedded extreme-condition PRs and finish the open rounds

Land the 10 open BuildTestEmbedded campaign PRs: 1364, 1355 and 1358 are in the merge queue; 1363 and 1352 need re-enqueue after checks; 1361 and 1372 need review then enqueue; 1360 must be reworked after it failed in the merge queue against main (oversized results must stay readable through some bounded route); 1371 and 1374 have red branch CI. Then open a PR for branch fix/F7fu-r5-pretest-phase-times-ordered and get the user's answer on two setup-only edits to flaky tests. 43 campaign PRs are already merged.

Handoff document

Markdown

Land the remaining BuildTestEmbedded extreme-condition PRs and finish the open rounds

Written 2026-10-09 14:40 UTC. RE-VERIFY before acting. Every PR state, queue position and CI result below is a snapshot from that moment. Re-check each PR with gh pr view <n> --repo CodexCoder21Organization/BuildTestEmbedded --json state,mergeStateStatus,statusCheckRollup,mergedAt,headRefOid. Re-check the merge queue with gh api graphql -f query='{repository(owner:"CodexCoder21Organization",name:"BuildTestEmbedded"){mergeQueue(branch:"main"){entries(first:20){nodes{position state pullRequest{number}}}}}}'.

Mission summary

The user asked: "Consider https://github.com/CodexCoder21Organization/DocumentationRepository/blob/main/architecture/TESTING.md#extreme-condition-testing-load-fuzzing-failure-injection-and-chaos and Using at most 6 opus subagents, we need to identify and introduce tests and fix as many of these problems as possible. We have a huge problem where behavior at scale is under-tested and it has bitten us in too many ways to count in recent weeks. I suspect most of the problems have been in https://github.com/CodexCoder21Organization/BuildTestEmbedded so let's fix all potential situations in that repository and make sure the tests are comprehensive so such conditions never arise again." They followed up with: "Do 5 rounds of identify+test+fix".

The campaign ran from 2026-10-08 ~17:00 UTC to 2026-10-09 ~13:00 UTC. It used up to six lanes, each covering one area: admission, torn writes, crash and restart, malformed input, performance and locks, clock steps, and run-ID validation. Every round followed the same pattern: identify one defect; write an end-to-end test through the public API that fails on main because of it; fix the root cause; open one PR per round. Every PR description leads with the user's request quoted above.

Each PR then went through an independent adversarial review of its final head and the orchestrator's final review, and only then into the merge queue. Delegates never merged or enqueued.

What remains: land the 10 open PRs listed below, finish F7's last round (branch pushed, no PR), and act on two open questions for the user. The "Further candidates" list holds suspected defects nobody has started.

Results so far

43 campaign PRs are merged (https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1308 through https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1365; list: gh pr list --repo CodexCoder21Organization/BuildTestEmbedded --state merged --search 'in:body "under-tested"', excluding 1082 and 1189, which predate the campaign). The themes:

  • Bounded admission and output. Queues, error texts, read responses and feed rows have hard limits. Anything cut is replaced by an explicit "omitted" marker. A rule set during review: no limit may silently drop tests or text.
  • Crash and restart safety. State is written atomically. Half-written temporaries are removed at startup. Receipt deletion and projection eviction are serialized.
  • Clock steps. The watchdog, deletion retries, the reconciler and run completion keep a sane order after the clock steps backward.
  • Malformed runner input. Unparseable lines, non-numeric telemetry and nameless events end as explained failures instead of crashing the run.
  • Performance. Scheduling, deletion drain and duration lookups went from quadratic rescans to indexes. Their tests count work done instead of timing it.

Open PRs (state at 14:35 UTC)

PR Branch @ head What it fixes Gates done What is left
https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1364 fix/F7fu-r2-run-id-length-control @ e35cca3cd033fdfe65184d1bffa7bee1977d644f Refuses run IDs that cannot name a directory: over 255 bytes, unpaired surrogate, or new control characters CI green, review CLEAN, orchestrator final CLEAN In merge queue, position 1. Watch it land.
https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1355 f5fu/r1-unreadable-preparation-build-rule-line @ 907a3792fca8c0308486d4496c77e25b47485e2b One unreadable build-rule line no longer fails shared-cache preparation; a rule whose only completion line is unreadable ends FAILED with the line named All gates CLEAN In merge queue, position 2.
https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1358 f4fu/r5-unparseable-test-script @ 95cc3bcb14ef64161b442bce546992959d5c86d6 A test script that cannot be parsed becomes a failing placeholder entry instead of being silently dropped; other scripts still shard All gates CLEAN In merge queue, position 3.
https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1363 fix/F7fu-r1-phase-times-backward-clock @ 194c83654bf034e80195879cab26caed94c1b7e2 Run phase times and droplet destroy times never go before an earlier recorded time All gates CLEAN, but removed from the queue at 13:38 UTC with failed_checks Read the merge-group failure: the queue entry was added 12:28, so look for the merge-group run around then. If it is the known fixture flake (see Open questions), re-enqueue. If it is an integration conflict with main, rebase, re-review and re-enqueue.
https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1352 fix/F3fu-r5-forensics-storage-margin @ e5139ccc7e8cf0b176b3d0c76d0ffef0e37308fe Forensic copies of deleted runs keep the free-space margin Branch CI green after one rerun of the shard that hit the known flake. The patch is byte-identical to a head reviewed CLEAN. Confirm it is not in conflict with main, then do the final review and enqueue. It was dropped from the queue earlier as DIRTY and rebased.
https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1361 bound-abandoned-droplet-service-calls @ 489d3161a3c84c730e28bd33ef4bb73e854026cf Limits droplet-service calls that time out but keep running (64 per kind of call) CI green; reviewer pre-read matched the invariant below Formal review verdict and final review on this head, then enqueue. Invariant: calls to a healthy provider are never refused; timed-out calls that are still running stay bounded (at most 64 + the callers already in flight). Tests that must exist: responsiveProviderServesConcurrentCallsWithoutRefusal (200 concurrent calls, 0 refused) and concurrentCallersToAStalledProviderLeaveABoundedNumberParked.
https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1360 f5fu/r3-bound-names-in-projection-feed @ 035827452a8ba9614cd49de8935771080311c2e6 Caps oversized test names, errors and fields in the projection feed so one row can't stall readers Failed in the merge queue at 12:41 UTC against current main. The rework was never pushed; this is still the old head. See "1360 rework" below.
https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1371 f5fu/r4-bound-rule-names-in-projection-feed @ 87442d903186776427e2a314a2a76f71fadd3480 The same limits for build-rule rows in the feed Branch CI red: bld-build and shard 1 (https://github.com/CodexCoder21Organization/BuildTestEmbedded/actions/runs/37926018049). Cause not yet read. Read the failure. It likely shares code with 1360, so do the 1360 rework first and align this PR with it.
https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1372 f4fu/r6-pending-deletion-temporaries @ 28a91bdeb18ca7efa6569eb67fdeadac5146bec4 Removes deletion-intent temporaries a stopped process left behind, at startup CI green; reviewer pre-read fine Review, final review, enqueue.
https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1374 f4fu/r7-run-temporaries-outside-fs @ 77eac52bf2b14757b3cf8d795b46303b79ab1ca6 Removes result and dispatch-mode temporaries from run directories at startup Branch CI red: bld-build plus shards 3 and 4 (https://github.com/CodexCoder21Organization/BuildTestEmbedded/actions/runs/37926443246). Cause not yet read. Read the failure and fix it. The bld-build failure suggests a compile or merge problem.

Branch with no PR

Repo Branch Remote head PR What is on it State
CodexCoder21Organization/BuildTestEmbedded fix/F7fu-r5-pretest-phase-times-ordered 060f20df65e56dd25a38e0fbdda80e38ac8b5858 no PR synchronizePreTestPhaseTimestamps now clamps projected pre-test phase times so they are never before the droplet's createdAt Committed and pushed at 12:38 UTC. Whether its fail-first evidence and neighbour tests were completed is unknown. Re-run its new test red (on main) and green (on the branch), then open the PR.

Nothing else is unpushed. The orchestrator's local scratch directory, with the lane clones and findings files, was wiped when the session ended. Every lane pushed at each milestone, so the remote branches above are the complete record. No production deploys or artifact publishes were made by this effort; it was code and tests only.

1360 rework (the most important open item)

The merge-queue failure (run https://github.com/CodexCoder21Organization/BuildTestEmbedded/actions/runs/37924741872, shard 1) came from the PR's own test:

java.lang.IllegalStateException: getTestResults cannot return 2 rows within its 1048576-byte complete response budget even after recoverable detail fields were omitted; the remaining response uses 1049097 bytes. Read the results in pages with getTestResultsPaginated.
  at buildtest.embedded.ReadRpcResponseBudgetKt.fitRecoverableRowDetailsToReadRpcBudget(ReadRpcResponseBudget.kt:405)
  at buildtest.embedded.BuildTestEmbeddedService.getTestResults(BuildTestEmbeddedService.kt:9900)
  at ...projectionFeedBoundsOversizedTestNameAndKeepsReadersMoving(...kts:125)

Branch CI was green only because the branch predates https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1351, which added the read budget.

The invariant to satisfy:

  1. Every stored test result must be readable through some size-limited public read route.
  2. The feed's truncation marker must name a route that actually returns the complete value. The current marker says getTestResults, which now refuses this case, and getTestResultsPaginated refuses a single row over the budget.
  3. Either limit names at ingestion, or let the results routes page past or omit a single oversized row with a marker. The test must exercise the chosen path end to end.

Rebase onto main first and reproduce locally.

Open questions for the user (no answer yet)

  1. Two intermittently failing tests: fix their setup? Both are races in the test setup, not in production code, and neither edit changes an assertion. The project rule says existing tests are not edited without the user's approval.
    • tests/testProjectionCompactionMixedHistoryMatchesMain.kts line 97: add service!!.latestProjectionCursor; between cancelBuildRun and archiveBuildRun. Cancel and archive can land in one journal batch, and then only the last row keeps its fingerprint, so one field differs from the expected file.
    • truncatedEventJournalNeverFinalizesAsPassed: truncate the event journal only after both runners have received drain, using an AtomicInteger and CountDownLatch, and hold their exit-code writes until then. Today it truncates while the second shard is still writing, so the failure message varies.
    • This test has already failed https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1352 once, and is a candidate cause of the 1363 queue removal.
    • Recommendation: approve.
  2. SSH test: count distinct channels destroyed instead of destroy calls? The test is e2eSshCloseDuringReconnectClearRejectsCompetingConnect. The double count comes from Apache MINA SSHD's channel close (2.12–2.15), which only the test uses. Our SshExecutor behaves correctly, and 150 plain runs did not reproduce the double count.
    • Recommendation: approve.
  3. Lower priority: oversized test scripts. Should the oversized-script path from https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1319 use the same placeholder approach as 1358? That would change four merged tests that assert the current exceptions.

Process rules this campaign used (keep them)

  • Rounds: one round per PR. Every round starts with a hermetic end-to-end test through the public API that fails on main. Use FakeDropletService, ManualClock and temp directories, the default 128 MB heap, no mocks, no reflection, and no widened visibility.
  • Verdicts count work rather than time it. Use counted bounds or orders-of-magnitude separation, never machine-tuned ones.
  • Never raise a timeout, reduce an iteration count, or edit an existing test without the user's approval.
  • Caps: no limit may silently drop text or tests. It either fails loudly or leaves an explicit omitted marker. A fallback may not collapse a whole run onto one shard. Each kind of droplet-service call gets its own budget.
  • Gates per PR: green CI, then an independent adversarial review of the exact final head, then the orchestrator's final review, then enqueue. Reviews include diffing existing test files against origin/main; a deleted assertion is a blocking finding (one PR overwrote an existing test file with the same name). A head that changes voids earlier reviews.
  • Merge queue:
    • CI is bld-build plus bld-test-shard 1–4. A queue run takes about 40–60 minutes.
    • Enqueue with gh api graphql -f query='mutation($id:ID!,$oid:GitObjectID!){enqueuePullRequest(input:{pullRequestId:$id,expectedHeadOid:$oid}){mergeQueueEntry{position}}}' -f id=<node id> -f oid=<head sha>.
    • Treat a queue failure as a real integration result. Read the shard log for TESTS FAILED and the stack trace above it.
  • Flakes: rerun the shard once (gh run rerun <id> --failed), record it, and never debug through CI.

Further candidates (suspected, not started)

Feed and results routes

  • getTestResultsPaginated refuses a page holding a single over-budget row. A test with a name over 1 MiB has no size-limited route to its results. This is tied to the 1360 rework.
  • Rows already in the journal from before the feed caps can still stall readers at their cursor. This needs a limit applied when the journal is read.

Event journal and deletion queue

  • When a live test-events.jsonl shrinks under the coordinator, the run fails with "java.io.EOFException: null" and nothing says the journal shrank.
  • DeletionWorker.rescanPendingIntentsLocked never rechecks a name it has already indexed. If an intent file is replaced by a directory, it stays in the index.

Clock order

  • provisioningAttemptStartedAt, admissionGrantedAt and admissionQueuedAt are still raw clock readings, so a provisioning attempt can appear to start before its admission grant. The watchdog budget relies on these values, so this needs its own design.

Run IDs and file names

  • Run IDs containing Unicode format characters (U+202E, U+200B, U+FEFF) are still accepted as directory names.
  • On a JVM whose file-name encoding is ASCII, a non-ASCII run ID makes getBuildRun/deleteBuildRun throw InvalidPathException instead of refusing the ID with a clear message.

Performance and limits

  • analyzeArchiveProjectIdentityIsLinearInProjectCount (from https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1327) timed out at 60 s twice on CI shard 2, stuck in isPackageNamePart / extractProjectNamesFromArchive, but ran in about 13 s locally. Count the work across input sizes to tell a path slower than linear from a machine-tuned limit. Do not raise the timeout.
  • LeaseScheduler: poll() walks every lease ever issued (~LeaseScheduler.kt:1738), isFinished() checks every test, and recordCompletedPeakMemory sorts all samples on each completion. These measured about 3.6× per doubling, which fell short of the orders-of-magnitude bar.
  • Resource estimates are rebuilt from the full history on every runner event (BuildTestEmbeddedService.kt ~19225/19302, DynamicDispatchOrchestrator.kt:1401). Not observable through the public API yet.

Upstream

  • kompile-buildscript BuildscriptCache.kt:2789 throws a NullPointerException with no location for a top-level function without a name. Fix it upstream in that library.

No status reports yet.

Add dependency

Complete this handoff

Moves it out of every priority list and into ArchiveArea.