Land the remaining BuildTestEmbedded extreme-condition PRs and finish the open rounds
Written 2026-10-09 14:40 UTC. RE-VERIFY before acting. Every PR state, queue position and CI result below is a snapshot from that moment. Re-check each PR with gh pr view <n> --repo CodexCoder21Organization/BuildTestEmbedded --json state,mergeStateStatus,statusCheckRollup,mergedAt,headRefOid. Re-check the merge queue with gh api graphql -f query='{repository(owner:"CodexCoder21Organization",name:"BuildTestEmbedded"){mergeQueue(branch:"main"){entries(first:20){nodes{position state pullRequest{number}}}}}}'.
Mission summary
The user asked: "Consider https://github.com/CodexCoder21Organization/DocumentationRepository/blob/main/architecture/TESTING.md#extreme-condition-testing-load-fuzzing-failure-injection-and-chaos and Using at most 6 opus subagents, we need to identify and introduce tests and fix as many of these problems as possible. We have a huge problem where behavior at scale is under-tested and it has bitten us in too many ways to count in recent weeks. I suspect most of the problems have been in https://github.com/CodexCoder21Organization/BuildTestEmbedded so let's fix all potential situations in that repository and make sure the tests are comprehensive so such conditions never arise again." They followed up with: "Do 5 rounds of identify+test+fix".
The campaign ran from 2026-10-08 ~17:00 UTC to 2026-10-09 ~13:00 UTC. It used up to six lanes, each covering one area: admission, torn writes, crash and restart, malformed input, performance and locks, clock steps, and run-ID validation. Every round followed the same pattern: identify one defect; write an end-to-end test through the public API that fails on main because of it; fix the root cause; open one PR per round. Every PR description leads with the user's request quoted above.
Each PR then went through an independent adversarial review of its final head and the orchestrator's final review, and only then into the merge queue. Delegates never merged or enqueued.
What remains: land the 10 open PRs listed below, finish F7's last round (branch pushed, no PR), and act on two open questions for the user. The "Further candidates" list holds suspected defects nobody has started.
Results so far
43 campaign PRs are merged (https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1308 through https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1365; list: gh pr list --repo CodexCoder21Organization/BuildTestEmbedded --state merged --search 'in:body "under-tested"', excluding 1082 and 1189, which predate the campaign). The themes:
- Bounded admission and output. Queues, error texts, read responses and feed rows have hard limits. Anything cut is replaced by an explicit "omitted" marker. A rule set during review: no limit may silently drop tests or text.
- Crash and restart safety. State is written atomically. Half-written temporaries are removed at startup. Receipt deletion and projection eviction are serialized.
- Clock steps. The watchdog, deletion retries, the reconciler and run completion keep a sane order after the clock steps backward.
- Malformed runner input. Unparseable lines, non-numeric telemetry and nameless events end as explained failures instead of crashing the run.
- Performance. Scheduling, deletion drain and duration lookups went from quadratic rescans to indexes. Their tests count work done instead of timing it.
Open PRs (state at 14:35 UTC)
| PR |
Branch @ head |
What it fixes |
Gates done |
What is left |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1364 |
fix/F7fu-r2-run-id-length-control @ e35cca3cd033fdfe65184d1bffa7bee1977d644f |
Refuses run IDs that cannot name a directory: over 255 bytes, unpaired surrogate, or new control characters |
CI green, review CLEAN, orchestrator final CLEAN |
In merge queue, position 1. Watch it land. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1355 |
f5fu/r1-unreadable-preparation-build-rule-line @ 907a3792fca8c0308486d4496c77e25b47485e2b |
One unreadable build-rule line no longer fails shared-cache preparation; a rule whose only completion line is unreadable ends FAILED with the line named |
All gates CLEAN |
In merge queue, position 2. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1358 |
f4fu/r5-unparseable-test-script @ 95cc3bcb14ef64161b442bce546992959d5c86d6 |
A test script that cannot be parsed becomes a failing placeholder entry instead of being silently dropped; other scripts still shard |
All gates CLEAN |
In merge queue, position 3. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1363 |
fix/F7fu-r1-phase-times-backward-clock @ 194c83654bf034e80195879cab26caed94c1b7e2 |
Run phase times and droplet destroy times never go before an earlier recorded time |
All gates CLEAN, but removed from the queue at 13:38 UTC with failed_checks |
Read the merge-group failure: the queue entry was added 12:28, so look for the merge-group run around then. If it is the known fixture flake (see Open questions), re-enqueue. If it is an integration conflict with main, rebase, re-review and re-enqueue. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1352 |
fix/F3fu-r5-forensics-storage-margin @ e5139ccc7e8cf0b176b3d0c76d0ffef0e37308fe |
Forensic copies of deleted runs keep the free-space margin |
Branch CI green after one rerun of the shard that hit the known flake. The patch is byte-identical to a head reviewed CLEAN. |
Confirm it is not in conflict with main, then do the final review and enqueue. It was dropped from the queue earlier as DIRTY and rebased. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1361 |
bound-abandoned-droplet-service-calls @ 489d3161a3c84c730e28bd33ef4bb73e854026cf |
Limits droplet-service calls that time out but keep running (64 per kind of call) |
CI green; reviewer pre-read matched the invariant below |
Formal review verdict and final review on this head, then enqueue. Invariant: calls to a healthy provider are never refused; timed-out calls that are still running stay bounded (at most 64 + the callers already in flight). Tests that must exist: responsiveProviderServesConcurrentCallsWithoutRefusal (200 concurrent calls, 0 refused) and concurrentCallersToAStalledProviderLeaveABoundedNumberParked. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1360 |
f5fu/r3-bound-names-in-projection-feed @ 035827452a8ba9614cd49de8935771080311c2e6 |
Caps oversized test names, errors and fields in the projection feed so one row can't stall readers |
Failed in the merge queue at 12:41 UTC against current main. The rework was never pushed; this is still the old head. |
See "1360 rework" below. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1371 |
f5fu/r4-bound-rule-names-in-projection-feed @ 87442d903186776427e2a314a2a76f71fadd3480 |
The same limits for build-rule rows in the feed |
Branch CI red: bld-build and shard 1 (https://github.com/CodexCoder21Organization/BuildTestEmbedded/actions/runs/37926018049). Cause not yet read. |
Read the failure. It likely shares code with 1360, so do the 1360 rework first and align this PR with it. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1372 |
f4fu/r6-pending-deletion-temporaries @ 28a91bdeb18ca7efa6569eb67fdeadac5146bec4 |
Removes deletion-intent temporaries a stopped process left behind, at startup |
CI green; reviewer pre-read fine |
Review, final review, enqueue. |
| https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1374 |
f4fu/r7-run-temporaries-outside-fs @ 77eac52bf2b14757b3cf8d795b46303b79ab1ca6 |
Removes result and dispatch-mode temporaries from run directories at startup |
Branch CI red: bld-build plus shards 3 and 4 (https://github.com/CodexCoder21Organization/BuildTestEmbedded/actions/runs/37926443246). Cause not yet read. |
Read the failure and fix it. The bld-build failure suggests a compile or merge problem. |
Branch with no PR
| Repo |
Branch |
Remote head |
PR |
What is on it |
State |
| CodexCoder21Organization/BuildTestEmbedded |
fix/F7fu-r5-pretest-phase-times-ordered |
060f20df65e56dd25a38e0fbdda80e38ac8b5858 |
no PR |
synchronizePreTestPhaseTimestamps now clamps projected pre-test phase times so they are never before the droplet's createdAt |
Committed and pushed at 12:38 UTC. Whether its fail-first evidence and neighbour tests were completed is unknown. Re-run its new test red (on main) and green (on the branch), then open the PR. |
Nothing else is unpushed. The orchestrator's local scratch directory, with the lane clones and findings files, was wiped when the session ended. Every lane pushed at each milestone, so the remote branches above are the complete record. No production deploys or artifact publishes were made by this effort; it was code and tests only.
1360 rework (the most important open item)
The merge-queue failure (run https://github.com/CodexCoder21Organization/BuildTestEmbedded/actions/runs/37924741872, shard 1) came from the PR's own test:
java.lang.IllegalStateException: getTestResults cannot return 2 rows within its 1048576-byte complete response budget even after recoverable detail fields were omitted; the remaining response uses 1049097 bytes. Read the results in pages with getTestResultsPaginated.
at buildtest.embedded.ReadRpcResponseBudgetKt.fitRecoverableRowDetailsToReadRpcBudget(ReadRpcResponseBudget.kt:405)
at buildtest.embedded.BuildTestEmbeddedService.getTestResults(BuildTestEmbeddedService.kt:9900)
at ...projectionFeedBoundsOversizedTestNameAndKeepsReadersMoving(...kts:125)
Branch CI was green only because the branch predates https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1351, which added the read budget.
The invariant to satisfy:
- Every stored test result must be readable through some size-limited public read route.
- The feed's truncation marker must name a route that actually returns the complete value. The current marker says
getTestResults, which now refuses this case, and getTestResultsPaginated refuses a single row over the budget.
- Either limit names at ingestion, or let the results routes page past or omit a single oversized row with a marker. The test must exercise the chosen path end to end.
Rebase onto main first and reproduce locally.
Open questions for the user (no answer yet)
- Two intermittently failing tests: fix their setup? Both are races in the test setup, not in production code, and neither edit changes an assertion. The project rule says existing tests are not edited without the user's approval.
tests/testProjectionCompactionMixedHistoryMatchesMain.kts line 97: add service!!.latestProjectionCursor; between cancelBuildRun and archiveBuildRun. Cancel and archive can land in one journal batch, and then only the last row keeps its fingerprint, so one field differs from the expected file.
truncatedEventJournalNeverFinalizesAsPassed: truncate the event journal only after both runners have received drain, using an AtomicInteger and CountDownLatch, and hold their exit-code writes until then. Today it truncates while the second shard is still writing, so the failure message varies.
- This test has already failed https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1352 once, and is a candidate cause of the 1363 queue removal.
- Recommendation: approve.
- SSH test: count distinct channels destroyed instead of destroy calls? The test is
e2eSshCloseDuringReconnectClearRejectsCompetingConnect. The double count comes from Apache MINA SSHD's channel close (2.12–2.15), which only the test uses. Our SshExecutor behaves correctly, and 150 plain runs did not reproduce the double count.
- Lower priority: oversized test scripts. Should the oversized-script path from https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1319 use the same placeholder approach as 1358? That would change four merged tests that assert the current exceptions.
Process rules this campaign used (keep them)
- Rounds: one round per PR. Every round starts with a hermetic end-to-end test through the public API that fails on main. Use FakeDropletService, ManualClock and temp directories, the default 128 MB heap, no mocks, no reflection, and no widened visibility.
- Verdicts count work rather than time it. Use counted bounds or orders-of-magnitude separation, never machine-tuned ones.
- Never raise a timeout, reduce an iteration count, or edit an existing test without the user's approval.
- Caps: no limit may silently drop text or tests. It either fails loudly or leaves an explicit omitted marker. A fallback may not collapse a whole run onto one shard. Each kind of droplet-service call gets its own budget.
- Gates per PR: green CI, then an independent adversarial review of the exact final head, then the orchestrator's final review, then enqueue. Reviews include diffing existing test files against
origin/main; a deleted assertion is a blocking finding (one PR overwrote an existing test file with the same name). A head that changes voids earlier reviews.
- Merge queue:
- CI is bld-build plus bld-test-shard 1–4. A queue run takes about 40–60 minutes.
- Enqueue with
gh api graphql -f query='mutation($id:ID!,$oid:GitObjectID!){enqueuePullRequest(input:{pullRequestId:$id,expectedHeadOid:$oid}){mergeQueueEntry{position}}}' -f id=<node id> -f oid=<head sha>.
- Treat a queue failure as a real integration result. Read the shard log for
TESTS FAILED and the stack trace above it.
- Flakes: rerun the shard once (
gh run rerun <id> --failed), record it, and never debug through CI.
Further candidates (suspected, not started)
Feed and results routes
getTestResultsPaginated refuses a page holding a single over-budget row. A test with a name over 1 MiB has no size-limited route to its results. This is tied to the 1360 rework.
- Rows already in the journal from before the feed caps can still stall readers at their cursor. This needs a limit applied when the journal is read.
Event journal and deletion queue
- When a live
test-events.jsonl shrinks under the coordinator, the run fails with "java.io.EOFException: null" and nothing says the journal shrank.
DeletionWorker.rescanPendingIntentsLocked never rechecks a name it has already indexed. If an intent file is replaced by a directory, it stays in the index.
Clock order
provisioningAttemptStartedAt, admissionGrantedAt and admissionQueuedAt are still raw clock readings, so a provisioning attempt can appear to start before its admission grant. The watchdog budget relies on these values, so this needs its own design.
Run IDs and file names
- Run IDs containing Unicode format characters (U+202E, U+200B, U+FEFF) are still accepted as directory names.
- On a JVM whose file-name encoding is ASCII, a non-ASCII run ID makes
getBuildRun/deleteBuildRun throw InvalidPathException instead of refusing the ID with a clear message.
Performance and limits
analyzeArchiveProjectIdentityIsLinearInProjectCount (from https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1327) timed out at 60 s twice on CI shard 2, stuck in isPackageNamePart / extractProjectNamesFromArchive, but ran in about 13 s locally. Count the work across input sizes to tell a path slower than linear from a machine-tuned limit. Do not raise the timeout.
LeaseScheduler: poll() walks every lease ever issued (~LeaseScheduler.kt:1738), isFinished() checks every test, and recordCompletedPeakMemory sorts all samples on each completion. These measured about 3.6× per doubling, which fell short of the orders-of-magnitude bar.
- Resource estimates are rebuilt from the full history on every runner event (BuildTestEmbeddedService.kt ~19225/19302, DynamicDispatchOrchestrator.kt:1401). Not observable through the public API yet.
Upstream
- kompile-buildscript
BuildscriptCache.kt:2789 throws a NullPointerException with no location for a top-level function without a name. Fix it upstream in that library.