← Priority list
Unclaimed

Review CI and deploy the merged coordinator with engine 0.0.615130200608

Finish the current full Actions gate, then obtain the orchestrator merge decision and deploy the locally built coordinator to url:buildtest: with a verified rollback backup and full readiness/feed/admission observation. The pin PR is OPEN at 3374a1315945e98623fb421784d38a1352131251, all six local targeted checks passed first attempts, and the failure fixture is repaired with a positive stale-file recovery counterpart. Engine 0.0.615130200608 is published and verified; no production upload was performed. Orchestrator owns CI watch/review/merge and separate upload dispatch.

Handoff document

Markdown

Handoff: Review CI and deploy the merged coordinator with engine 0.0.615130200608

Written 2026-10-09 03:57 UTC. RE-VERIFY: This is a write-time snapshot; before acting read the orchestrator decision, query the PR and remote branch heads, and inspect nursery health. Fresh state wins.

Original scratchpad: /tmp/claude-1000/-code/50fdf0b6-d89e-42ff-8356-33fffc39f1da/scratchpad; brief briefs/deploy1340.md, findings out/deploy1340-findings.md, authoritative current decision out/deploy1340-decision.md. Re-check current orchestrator instructions if that filesystem is gone; the remote decision snapshot is historical.

Mission

The user’s standing request is “deploy after each merge.” This lane was authorized to publish BuildTestEmbedded main, open the coordinator engine-pin PR, then only after the orchestrator writes “MERGED — proceed to deploy” build merged coordinator main and upload exactly url:buildtest:. Never merge/enqueue in this lane, change config or other routes, manually restart, deploy WUI, or submit runs. SSH is read-only and scoped to copying FROM the host, uptime/free, and durable-record observation if needed. Original lane began 02:39 UTC with an 04:04 deadline; a successor needs its own time budget and current orchestrator direction.

What was found and done

  1. Fresh engine main was 0be27701e9c3e236afddd7ec15bd2d080591e6b0. It includes the independent projection-reader descriptor fix, clock ordering, memory skip-record and duration statistics changes requested by the brief, and other earlier merged changes including the deletion recovery fix. Build unchanged main locally with scripts/build.bash --local buildtest.embedded.buildMaven <artifact-dir>. Its ordinary project POM version was already-published 0.0.69274749, so only the built POM root version, names and checksums were changed to a new publication coordinate. Source was not changed.
  2. Published buildtest.embedded:buildtest-embedded:0.0.615130200608. HTTPS-fetched POM and jar matched local bytes. Jar 5,232,916 bytes, SHA-256 0e9363ebeff3540bdcbffe4f81e55cbf54197b08c1fe5ab0fa4c9ad45301ec80; POM SHA-256 92110aefbebdf12528500bc8edaf85881c14065a402b1bc74d8d957641909808. Both fetchable under https://kotlin.directory/buildtest/embedded/buildtest-embedded/0.0.615130200608/ with the matching versioned filenames. Published code is already on engine main; its new artifact coordinate is intentional.
  3. Coordinator pin PR changes four one-line version references: production engine pin, two direct test artifact pins and the default-runner composition comment. Resolver 0.0.1259, protocol 0.0.527, runner 0.0.123 and manager 0.0.205 remain unchanged. The four targeted local composition/linkage/bind tests passed 4/4 twice, including after final rebase. Independent review found no actionable issues.
  4. Full Actions CI passed 239/240 tests and failed uploadSessionsEvictExpiredEntries. Exact-head local reproduction failed 0/1. It expects a regular runs/<id>/run.json.tmp file to force deletion persistence failure. The earlier merged https://github.com/CodexCoder21Organization/BuildTestEmbedded/pull/1332 (commit https://github.com/CodexCoder21Organization/BuildTestEmbedded/commit/8f54cb7ddb93217ce3c2003eaf0c0b961c6700a7) deliberately deletes that stale regular file under the publication lock and retries. That makes the fixture obsolete. This is not a deferred/swallowed write or a moved tombstone path: deleteBuildRunNow still synchronously calls replaceRunSnapshotIfUnchanged, writes/renames run.json, and wraps/throws persistence errors inline.
  5. The repaired failure fixture now creates a directory at run.json.tmp and asserts the actual full java.io.FileNotFoundException: <path> (Is a directory) cause. A local experiment that retained the old literal proved assertFailsWith<IllegalStateException> passed and only the exact cause comparison failed. The corrected fixture keeps the complete wrapping message, retained-session retry, archived/bypass states and exact cancel count. It matches the upstream real-directory failure test. No timeout/iteration/assertion was weakened.
  6. The positive counterpart uses real public handler/service calls: it copies a canceled run's valid snapshot to a stale regular temporary, advances the handler's ManualClock past TTL, verifies successful fresh-session creation and temporary removal, checks archived old run, rejects old-session chunks with the full message, and accepts a fresh chunk. Six targeted local checks passed 6/6 first attempts after the complete repair; independent final review found no actionable issues.
  7. Orchestrator initially required unchanged exact literal and raw XML. Clarification at 03:49 permits changing ONLY the observed cause, retaining full exact-message/type checks, and accepts raw Actions per-test result lines because the existing workflow has no XML upload. No workflow change. Complete repair was fetched/rebased (main unchanged), PR OPEN checked, decision reread, committed and pushed at 3374a1315945e98623fb421784d38a1352131251. Full Actions run https://github.com/CodexCoder21Organization/BuildTestServerService/actions/runs/37881167887 is IN_PROGRESS at write time; no first-attempt verdict yet. PR body includes “Round 2 — fixture repair for engine 0.0.615130200608,” old/new fixture text, observed cause and 6/6 local evidence. Orchestrator owns the watch/review/merge and separate upload dispatch; this lane does not upload before deadline or skip startup observation.
  8. No production upload, restart or new run was performed. Route was rechecked at 03:52–03:53 UTC and is RUNNING with /root/ContainerNursery-uploads/jars/buildtest-server-e95a02b6-20261005-dep10.jar, port 33915. At 03:03 load was 25.05/27.68/28.55, available memory 2,787 MiB, swap used 1,107/3,071 MiB; refresh before deployment. /api/runs was outOfDate=true, refreshFailing=false, so WUI run state was not trustworthy. No rollback copy exists yet because upload gate was not reached.

Relevant PRs and refs

Repo Branch Remote head SHA PR What is on it State
BuildTestEmbedded wip/deploy1340-evidence 29e51f2c3f9c269e0a44450e91978b68cdd15e15 no PR Findings, publication/test logs and XML, review, initial proposed patch, final PR body, decision snapshots Documentation/evidence only; product code unchanged from merged main
BuildTestServerService wip/deploy1340-engine-pin 3374a1315945e98623fb421784d38a1352131251 https://github.com/CodexCoder21Organization/BuildTestServerService/pull/404 Four pin references, directory failure fixture with full cause, positive stale-file counterpart OPEN, 6/6 local first-attempt pass; bld-build IN_PROGRESS
BuildTestServerService wip/deploy1340-round2-evidence 3374a1315945e98623fb421784d38a1352131251 no separate PR; same code as pin PR Complete repaired coordinator source checkpoint Identical to PR head; recoverable remotely

Remote heads were compared with local git rev-parse HEAD; both repositories had empty status/stash/unpushed-commit sweeps and only one own worktree each. No local-only code remains.

Published but not deployed: engine coordinate above, built from merged engine main. Coordinator PR is unmerged; new engine is NOT live on the coordinator as a result of this lane. Re-check health/current image before any upload.

Next steps

  1. Re-verify current state first: decision file, PR state/checks, remote branches and live route. Do not infer approval from elapsed time. Decision snapshots are on the evidence branch; current orchestrator file is authoritative.
  2. The source repair is COMPLETE and on the pin PR branch. Do not reapply the initial proposed patch. Clone the PR branch fresh and inspect the actual head; branch/table above links all source/evidence. Scope wording and evidence alternatives were resolved. Current full Actions run is the one to watch; do not submit another run as a debugging strategy.
  3. Orchestrator explicitly owns current CI watch/review/merge/upload. Coordinate that ownership before starting a duplicate watcher. Use plain build-watchman 0.0.21 with a tracked hard deadline for the full current Actions gate; raw per-test console lines are accepted first-attempt evidence. Count pass/fail lines and give retries no credit. If red, diagnose and reproduce at cause before changes or retries. For any changes: fetch/rebase, targeted local verification, inspect PR still OPEN, commit with required footer, read decision before push. No full local suite on the shared host.
  4. On green, write READY FOR ORCHESTRATOR MERGE <exact head> in findings and await the orchestrator’s actual merge decision. Never merge/enqueue yourself. No production action before “MERGED — proceed to deploy”.
  5. After that decision: pull coordinator main, confirm the engine pin, take a slot and build fatjar LOCALLY (scripts/build.bash --local buildtest.server.buildFatJar <jar>; >32 MiB cannot be returned by remote builder). Record jar bytes/SHA-256. Copy live jar FROM host using scp -P 23 root@198.199.106.165:/root/ContainerNursery-uploads/jars/buildtest-server-e95a02b6-20261005-dep10.jar <backup>; compare SHA-256 with the live file. Read-only ssh -p 23 root@198.199.106.165 'uptime; free -m'; count coordinator durable records if needed. Refresh exact image first in case another authorized lane deployed.
  6. Read decision AGAIN before upload. Use nursery CLI upload-jar --file <jar> --route 'url:buildtest:' --url https://api.nursery.wasmserver.com. It overwrites the route image and restarts coordinator; no manual restart/config/WUI change. Ensure time budget for full observation BEFORE upload.
  7. Watch health-detailed STARTING to RUNNING and startup log. Prior startup recovery took 629–674s; nursery kills at 900s and two prior generations hit that deadline. If killed/relaunched at 900s, do not act a second time; immediately put ## ROLLBACK RECOMMENDED plus evidence in findings, recommending the verified backup for orchestrator to decide. Observe feed refreshFailing clearing ~2min after ready, no STARTING loop, and an organically admitted new run reaching TESTING within 20min. Never submit a run. Capture dashboard screenshot if Chromium is available, otherwise HTML banner state.
  8. Report exact version, PR/head, deployed jar SHA, readiness time, observations, rollback path and unproven items; remove only this lane’s checkouts/cache namespaces/processes. Keep backup/evidence out files.

Operational knowledge

  • Use $HOME/bin/coursier; /usr/local/bin/coursier is broken.
  • Nursery CLI 0.0.20 health-detailed, containers, server logs and provider logs work. container --route-key url:buildtest: gives endpoint-not-found; child container-logs gives JSONObject parse error or hits deadline with both 0.0.20 and 0.0.21. Provider logs are empty for this route. Fresh rotated startup logs may work after upload, but that was not proven. Never deploy just to repair a log command. Fabric cert directory is absent. Generic server logs can provide lifecycle events; child startup evidence still needs inspection.
  • Nursery docs mention /containers/<route>/logs, but current source handles /logs/container/<route>. CLI health works independently. This mismatch was recorded for orchestrator; report-challenge auto-merges outside this lane, so it was not invoked.
  • Headless Chromium executable was not found at standard locations; HTML banner fallback is available.
  • Local slots live at scratchpad/out/build-slot-{1,2,3}, retry30s with 10min limit and trap release. Never hold across waiting; never remove another lane’s slot. Own cache namespaces end _workspace_workspace_deploy1340_engine and _workspace_workspace_deploy1340_coordinator.
  • Handoff CLI health passed during this lane. Create/complete through CLI, not handoff git PR/merge. No production operations on the handoff route are authorized.
  • Evidence branch includes findings, publication log, raw CI failure, local XML/logs, unapplied patch, independent reviews and draft PR body. Jar outputs/caches and credentials are deliberately not in Git. Published jar is recoverable from Maven; live rollback backup must be taken fresh before any deployment.

Lane cleanup

At 03:57 UTC both own fresh checkouts and their matching ~/.aibuildcaches namespaces were removed, along with redundant local publication jar/POM/artifact output. No process was active under either checkout; every tracked command exited. Source, tests, results and findings needed to resume are on the linked remote branches; other lanes and their caches/slots were untouched. No backup jar exists yet.

No status reports yet.

Add dependency

Complete this handoff

Moves it out of every priority list and into ArchiveArea.