← Priority list
Unclaimed

Decide the UrlResolver sandbox-retirement design (PR 1226), then land the TinySite outage fixes, redeploy SFS/filedrop/TinySite, and publish the sazabi.com capture

Remaining: (1) owner picks the UrlResolver PR 1226 design for timed-out sandbox calls (recommended: retire for new calls only, keep already-returned objects readable) and it gets implemented, reviewed, published; (2) review+merge UrlResolver PR 1227 (resolver 0.0.1306 already published); (3) fix remaining review findings and land SimpleFileSystemServiceServer PR 25, FiledropEmbedded PR 14 + FiledropServiceServer PR 27 (and PR 26), TinySiteWui PRs 1+3, each after clean gpt-6.1-sol + Opus reviews per the owner's rule; (4) redeploy SFS-vnext, filedrop-server, tinysite-wui from pulled main (owner-approved) and resolve the current filedrop.wasmserver.com 500 (filedrop cached a failed SFS open at 01:57Z); (5) deploy the sazabi.com capture (Good-UI-Designs branch wip/dark-landing-sazabi-handoff-2026-10-09) to TinySite for owner approval. State: ContainerNursery main dab37c78 (EOF connection-leak fix) is deployed as systemd unit containernursery-dab37c78-20261008T154648Z and verified; CN PR 652 merged; evidence on PlanRepository branch handoff-artifacts/2026-10-08-tinysite-outage.

Handoff document

Markdown

Decide the UrlResolver sandbox-retirement design, then land the TinySite outage fixes and redeploy

RE-VERIFY: write-time snapshot 2026-10-09 ~02:00 UTC. Every PR state, CI result, review verdict and production fact below was re-checked at write time but will drift. Re-check each PR with gh pr view <full URL> --json state,headRefOid,mergeStateStatus,statusCheckRollup, branches with git ls-remote, production with HardwareControlFabric (see Operational knowledge) and coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com health-detailed.

Mission

On 2026-10-08 the owner asked to capture https://www.sazabi.com/ into Good-UI-Designs (good-ui-designs skill). Publishing the capture to TinySite (https://tinysite.wasmserver.com) failed repeatedly (HTTP 502/503, chunk 0 is missing), and the owner said: "What is going wrong with tinysite? If we are running into a problem, we need to fix it." This handoff carries that fix effort plus the paused sazabi capture.

Owner decisions/approvals already given (carry them forward; do not re-ask):

  • Merging: a PR may be merged after both an adversarial review by gpt-6.1-sol (the owner's own instruction — via codex) and an independent Opus review come back clean on the exact final head, plus green CI and a final review.
  • Deploys approved after merges for: ContainerNursery, simple-filesystem-vnext, filedrop-server, tinysite-wui. Always git pull the latest main before building a deploy.
  • ContainerNursery heap: keep -Xmx768m for now (live set measured ~180 MB after full GC on 2026-10-08, above the stored 128m directive); investigate getting back to 128m separately.
  • The ContainerNursery watchdog cron (/root/cn-watchdog.sh every minute) was removed per the owner's no-auto-restart directive (backup /root/crontab.backup-20261008T150706Z). Do not reinstate.
  • Approved test changes: UrlResolver testLazyTimeoutRetirementRejectionClosesSjvmExactlyOnce may change only its trigger (busy guest instead of interruptible RPC), assertions unchanged; FiledropEmbedded testPublishedChunkedPayloadRecoveryAfterMarkerCleanupFailure assertion updated to "a retried commit of an already-published upload returns the file".
  • OPEN DECISION (owner asked for this handoff to enable it): UrlResolver PR 1226 design — see "Decision needed" below.

What was found (root-cause chain, verified)

  1. ContainerNursery (CN) URL facade leaked one backend TCP connection per streamed response drained to EOF. Deployed build a2fc6415 bundled UrlResolver 0.0.1203, which deregistered a finished response stream at EOF without closing it; the stream owned the CN→backend TCP connection, so it stayed open and the client's later __stream_close could not reach it. Evidence: simple-filesystem-vnext (SFS) had 7,168 threads = 3,576 idle TcpServer-Reader/Writer pairs, all from 127.0.0.1 owned by the CN pid (ss -tnp), growing ~50–80/min; SFS logs showed 407 v2_readChunkStream and 0 __stream_close. Fail-first test: 40 EOF reads leave 40 backend connections open on a2fc6415, 0 on main. Fixed by https://github.com/CodexCoder21Organization/ContainerNursery/pull/632 (merged 2026-09-27, never deployed until 2026-10-08). Deployed and verified in production: after deploy, 27 reads → 27 __stream_close, 0 CN-held SFS connections, SFS 16 threads; TinySite file serving went from 12–30 s to 0.3–0.6 s.
  2. SFS (url://simple-filesystem-vnext/, -Xmx128m, UrlResolver 0.0.1095, embedded 0.0.18) OOM-crashed (OutOfMemoryError: Java heap space 2026-10-08 10:33:04Z, hprof /root/ContainerNursery/heapdumps/java_pid3588199.hprof). Contributors: (a) each leaked connection ~12 KB heap (+ direct buffers); (b) SFS resolver 0.0.1095 retains streamed chunk data (up to 16 MiB each) in StreamAwareServiceHandler.pullableStreams until __stream_close/connection close — never arrived; (c) embedded 0.0.18 re-parses/rewrites a ~3 MB namespace-event file per mutation (~82 MB garbage each, bounded, not a leak). (a)/(b) neutralised in production by the CN deploy; SFS's own fix is PR 25 (+ resolver 0.0.1306 from PR 1227); (c) is PR 23.
  3. FiledropEmbedded 0.1.11 (filedrop-server): every metadata read deleted one session marker per committed chunked file (thousands of v2_delete/min while serving/publishing); commit not idempotent (duplicate commit racing an in-flight one → chunk 0 is missing); construction always rewrote index.json; remote I/O under one global listingLock. Fixed in FiledropEmbedded PR 14 (several review rounds).
  4. TinySiteWui: one failed/timed-out call tore down and rebuilt the shared UrlResolver, failing unrelated concurrent requests; commit retry re-sent commits that might still be running; page aborted sessions on unknown outcomes. Fixed in TinySiteWui PR 3 (stacked on PR 1, the unmerged branch production runs).
  5. UrlResolver (upstream): (a) any sandboxed call deadline retires the whole SJVM, making objects returned by sibling calls unusable (Cannot invoke method on closed instance proxy: filedrop/api/DropImpl.getId()) — a successful commit can answer 502 — PR 1226, design open; (b) StreamAwareServiceHandler shutdown not terminal and close failures swallowed — PR 1227, done, published 0.0.1306.

Ruled out: host capacity (host MemAvailable 13.7 GB; SFS heap analysis shows leak/churn, not under-provisioning); filedrop expired-drop purge as the delete-storm source (deletes were non-recursive v2_delete from the read path, not deleteRecursively).

Incident caused during this effort (resolved) — read before touching CN

Restarting CN from a script started via HardwareControlFabric launch put CN and all its app JVMs into hcf-daemon.service's cgroup (MemoryMax=512M, CPUQuota=120%). CN's memory guard then refused/killed containers ("RAM avail 0MB") → kotlin.directory, buildtest, kotlin.build, handoff, tinysite, filedrop etc. down ~15:07–15:48Z 2026-10-08. Rolling back the jar did not help. Fixed by relaunching CN in its own transient systemd unit. Recorded: https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-08-1743-restarting-containernursery-from-a-hardwarecontrolfabric.md. Never (re)start CN or app services directly from an HCF-launched shell; use systemd-run (recipe in Operational knowledge).

Decision needed: UrlResolver PR 1226 design

Problem: in UrlResolver 0.0.1302, when one sandboxed service-proxy call hits the SJVM per-call deadline (30 s), invokeSandboxMethod retires the whole SJVM (SandboxedProxyGenerator.kt timeout branch ~1396-1404, from https://github.com/CodexCoder21Organization/UrlResolver/pull/713), and every object previously returned by that sandbox (e.g. TinySite's DropImpl/DroppedFile) then throws "Cannot invoke method on closed instance proxy". Reproducer: tests/testReturnedObjectSurvivesSiblingCallTimeout.kts on the PR branch (fails on main 4fe0615 with the production message); TinySite-side reproducer branch https://github.com/CodexCoder21Organization/TinySiteWui/tree/w3b/slow-storage-call-reproducer.

PR 1226's current implementation (head fc51e8a): if the timed-out call is inside ServiceBridge.rpc, interrupt only that call and keep the sandbox; otherwise retire. Both adversarial reviews (findings: reviews/out/ur1226-sol-findings.md, reviews/out/ur1226-opus-findings.md in the artifacts branch) showed holes: deadline firing in the post-RPC tail (effects, stream resolution, result marshaling) or a swallowed interrupt lets the abandoned guest keep computing in the live sandbox (probe: sandboxActiveAfterTimeout=true, worker at ~985 CPU-ms/s); effect-resume can send a new RPC after abandonment; the changed test leaves a CPU-burning guest thread running; the "no new RPC after abandonment" guard is untested. Also observed: retirement on main never actually stopped a runaway guest either (the interpreter has no cancellation points; SJVMImpl.close() only cancels a coroutine scope).

Lane analysis (from SJVM source, sandboxjvm main / 0.0.55 lineage): host exceptions thrown into the guest are catchable (SJVMThreadImpl.kt:466 converts to guest InternalError); the interpreter loop has no cancellation check; a host java.lang.Error would unwind uncatchably but skips guest finally/monitorexit, leaving monitors held (Monitor.kt has no per-thread release). So "abandoned call never does more work" (INV-A) and "guest locks/state stay consistent" (INV-B) cannot both be guaranteed inside UrlResolver alone.

Options:

  1. Retire for new calls only (orchestrator's recommendation, not yet chosen by the owner): keep main's "any deadline retires" rule, but retirement only (a) routes NEW calls to a fresh sandbox (LazyReconnectingMethodDispatcher already reconnects) and (b) refuses further host/remote calls from the retired SJVM; objects already returned stay readable on the retired SJVM until unreachable/idle-reaped (then the SJVM closes). Rationale: retirement never stopped runaway guests anyway, so keeping the retired SJVM's heap alive for already-returned value objects adds no new execution risk, needs no new time bound, and fixes the production symptom. Cost: retired SJVMs live as long as their returned proxies (needs a Cleaner/idle path); methods on returned objects that themselves need remote calls would fail after retirement.
  2. Bounded extra window, then retire (lane option 2): only the pure remote wait counts as "waiting on the service"; after the deadline every exit path throws into the guest, enter()/effect-resume/stream callbacks refuse new host calls, and if the worker hasn't exited within one more method-deadline window the sandbox retires. Adds a new time bound (needs explicit owner approval under the timeout rules); a misbehaving guest can compute up to one window in a live sandbox.
  3. Upstream SJVM cancellation (lane option 1): per-thread cancellation checked at loop back-edges/calls + monitor-releasing unwind + handler detection in https://github.com/CodexCoder21Organization/sandboxjvm, then force-unwind when safe, retire otherwise. Correct and also fixes "retirement never stops a stuck guest", but large; guest-implemented locks (ReentrantLock) would still force retirement.
  4. Keep always-retire, fix only the client (TinySite copies fields out of returned objects immediately / tolerates failures). Contradicts the fix-upstream rule.

Whatever is chosen: the remaining asked-for work is ready (INV-C test hygiene: guest waits released in finally + assert worker exits; fail-first tests for post-RPC-tail deadline, abandoned-call RPC attempt, effect-resume after abandonment; regression test cleanup; README contract), then full suite, CI, publish next unused resolver version (0.0.1305 is declared on the branch but unpublished; 0.0.1306 is taken by PR 1227 — re-check kotlin.directory before publishing), then bump TinySiteWui and add the sibling site-creation case to its timed-out-call test. The idle-reaper variant (sandbox closed after 120 s idle while the caller holds returned objects) is a separate follow-up recorded in the PR.

Relevant PRs / refs (state at write time)

Repo PR Branch @ head CI (kotlin.build (remote)) Review state Next
ContainerNursery https://github.com/CodexCoder21Organization/ContainerNursery/pull/632 merged 2026-09-27 green — deployed 2026-10-08 (see Deployed)
ContainerNursery https://github.com/CodexCoder21Organization/ContainerNursery/pull/652 test/url-facade-many-eof-reads-close-backend-connections @ 1fcb0f664217e7f716094ab6387e8fdb7aee55b7 green sol CLEAN + Opus CLEAN (round 2) MERGED 2026-10-08T18:14Z (merge commit 09e352ffeaa3df11de38d11d34c4eb027b3e574b); test-only, nothing to deploy
SimpleFileSystemServiceServer https://github.com/CodexCoder21Organization/SimpleFileSystemServiceServer/pull/25 fix/stream-release-bump @ ac8ba701fe343c10a55a117c0d986a2387ea3107 green (31/31 local) round 2: Opus CLEAN; sol FINDINGS(2 major: late stream registration after host close; swallowed stream-close failures) — both are UrlResolver defects now fixed by PR 1227/0.0.1306 bump resolver 0.0.1302→0.0.1306 (+POM alignment), add host-close-race + close-failure-retry tests, re-review (sol+Opus), merge, deploy SFS
SimpleFileSystemServiceServer https://github.com/CodexCoder21Organization/SimpleFileSystemServiceServer/pull/23 bump-embedded-0.0.26-upgrade-compatibility @ b5e408b3e2f231e0a43fd63ee930706713574288 red: all 42 tests passed but testDurableBackendLoopback exceeded @Timeout(60) at 62.5 s under CI CPU over-admission (buildtest log: CPU admission "OBSERVE mode" admitted 5573 millicores on a 2-vCPU droplet); evidence comment https://github.com/CodexCoder21Organization/SimpleFileSystemServiceServer/pull/23#issuecomment-6062919496 not reviewed this effort do NOT raise the timeout; reduce CPU cost of the PR's new upgrade tests or fix buildtest CPU admission; version collision with PR 25 (both claim 0.1.1x) → rebase whichever lands second
SimpleFileSystemServiceServer https://github.com/CodexCoder21Organization/SimpleFileSystemServiceServer/pull/24 feat/closeable-simplefilesystem-manager @ 26b0bd554be8c69d77fe4c835e61f5ef0e3bc0cd green not reviewed this effort needed by FiledropServiceServer PR 26 (requires publishing simplefilesystemservice-client 0.1.7)
FiledropEmbedded https://github.com/CodexCoder21Organization/FiledropEmbedded/pull/14 fix/reads-do-not-delete-session-markers @ 9c26a4a788c47182996280f69be1ed60869187fa green (88/88 local) round 3: Opus FINDINGS(2 minor: revision guard upsertNewerListingEntry ~1078-1083 untested/never fires; README overpromises listing freshness after non-retried uploadFile/deleteFile/createDrop listing-write failure); sol FINDINGS(2 major: listing rebuild can overwrite a newer projection across processes; overlapping 0.1.11/0.1.16 writers during upgrade invalidate revision ordering) — reviews/out/fdc-*-findings.md fix round 4, publish next unused version (0.1.14/0.1.15/0.1.16 claimed), repoint PR 27, re-review, merge
FiledropServiceServer https://github.com/CodexCoder21Organization/FiledropServiceServer/pull/27 fix/pin-filedrop-embedded-0.1.14 @ 48771cf2a10a104345e700a62c0cc259ac67b5bb green (24/24) reviewed with PR 14 (test fixture serialization judged a correct fixture fix) repoint to final embedded version; merge after PR 14; deploy filedrop-server
FiledropServiceServer https://github.com/CodexCoder21Organization/FiledropServiceServer/pull/26 fix/filedrop-init-retry-after-failure @ 0e2a3ad60ea92359be46c26110e6d96c2abdc457 red (testE2eSjvmColdStartInitializesAfterProtocolProbe, also flaky on main) not reviewed production-relevant (see Open production issue); blocked on simplefilesystemservice-client 0.1.7 (SFS PR 24)
TinySiteWui https://github.com/CodexCoder21Organization/TinySiteWui/pull/1 tinysite-initial @ 80225f405bc4e28fb60690504d724df5ad407503 green (14/14) round 2 findings below (shared with PR 3) production base (main is a skeleton); old handoff https://www.handoff.wasmserver.com → hf-2026-10-06-merge-the-tinysite-pull-request-once-the-owner-approves tracks it
TinySiteWui https://github.com/CodexCoder21Organization/TinySiteWui/pull/3 fix/storage-call-failure-does-not-fail-concurrent-requests @ 31b10f24ef51c4e5cbbbecc5c17afa1aaf2491e4 (stacked on PR 1) green (18/18) round 2: Opus FINDINGS(1 major: relay RELAY_FORWARD_FAILED after dispatch classified as service-answered → plain 502 + page abort; minors: PR 1's "refused 400" happens after commit; untested NOT_SENT branch; slow page SHA-256 ~27 MB/s; 502/504 without JSON body treated as definitive); sol FINDINGS(1 major: commit response truncated after headers still triggers abort; 1 minor: confirmation can match another site's older file) — reviews/out/tsb-*-findings.md fix round 3; also the RELAY_FORWARD_FAILED isServerError classification should be fixed upstream in UrlResolver; then bump resolver once PR 1226 is published
UrlResolver https://github.com/CodexCoder21Organization/UrlResolver/pull/1226 fix/returned-object-survives-sibling-timeout @ fc51e8aa63ee9bce304dd0581ac3ef50cef4bd46 red: run https://buildtest.kotlin.build/run?id=d6c2e103 1885/1892 (7 failed — not yet triaged); bld-build failed on testDeadConfiguredBootstrapDialsBackOffExponentially (pre-existing intermittent) FINDINGS from both reviewers (round 1) owner design decision, then implement + tests, triage the 7 CI failures, re-review, publish
UrlResolver https://github.com/CodexCoder21Organization/UrlResolver/pull/1227 fix/stream-handler-terminal-shutdown-and-cleanup-failures @ 0d832e1485d482f5da58668e62977ddd15c24b0f green (1897/1897) not yet reviewed sol + Opus review, merge; foundation.url:resolver:0.0.1306 already published from this head

Branches pushed for this handoff (no PR)

Repo Branch Remote head SHA PR What is on it State
CodexCoder21Organization/PlanRepository handoff-artifacts/2026-10-08-tinysite-outage 8bad7f571e83b0f20fef397472d4d48033c13984 no PR handoffs/artifacts/2026-10-08-tinysite-outage/: investigation scripts + remote outputs (jstack summaries, ss/cgroup/memory snapshots, CN live histo), gzipped service logs, every adversarial review prompt + findings + final verdicts, reviewer probe tests and mutation patches, lane notes/findings for all lanes, delegation briefs, prior root-cause findings (agentA SFS heap, agentB filedrop, agentC CN leak + 40-read test) evidence only; raw codex transcripts and duplicate logs omitted (large, no extra content)
CodexCoder21Organization/Good-UI-Designs wip/dark-landing-sazabi-handoff-2026-10-09 1659dbe69853c082c55c389152c02d5c28641d39 no PR dark-landing-sazabi/ (captured site, screenshot, README) + _wip-sazabi-capture-tools/ (record/build/serve/verify scripts; delete before PR) sub-path HTTP copy verified against live site; not owner-approved; not deployed to TinySite; raw 36 MB recording not committed (regenerate with record.js)
CodexCoder21Organization/TinySiteWui w3b/slow-storage-call-reproducer 973ea0c2123c9b7e5dfc2368bdce50c7f6c9c456 no PR reproducer for the UrlResolver returned-object defect reproducer only

All lane clones were swept (git status, git log --branches --not --remotes, stashes): no unpushed commits; the only uncommitted diffs were reviewers' deliberate mutation/probe worktrees, saved as patches under reviewer-mutations/ in the artifacts branch.

Published artifacts (verified on kotlin.directory by SHA-256)

  • filedrop.embedded:filedrop-embedded:0.1.14 — 8ccf122e7c67e1da704ae14aea2eb57a7a5ec80bc4e1b40355cd97c95c194e17 (round-1 head 981fd49; superseded, contains round-1 defects — do not use)
  • filedrop.embedded:filedrop-embedded:0.1.15 — a1059bfa0a33d81f5d481037d5fa79f84ff55a7c91914d907fa8aa9d96a252b8 (round-2 head 7320540; superseded)
  • filedrop.embedded:filedrop-embedded:0.1.16 — cd81f30bf20a1e455fcc84005dd779f35bf20185af0a83c9da435fe38a9b6b37 (round-3 head 9c26a4a; has open round-3 review findings)
  • foundation.url:resolver:0.0.1306 — 0d134fcb961e8ef1030f18d4a409d63fcfb20ccde31be2f6c33efe8e72d73dbb (UrlResolver PR 1227 head 0d832e1; PR not yet reviewed/merged)
  • foundation.url:resolver:0.0.1305 — not published (declared on PR 1226's branch)

Deployed — what is live (verified 2026-10-09 01:59Z via HCF)

  • ContainerNursery pid 1368870, transient systemd unit containernursery-dab37c78-20261008T154648Z.service (KillMode=process, no Restart=), cmd java -Xmx768m -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/root/ContainerNursery/heapdumps/ -jar bin/container-nursery-dab37c78-20261008T1454Z.jar config.json (cwd /root/ContainerNursery), jar SHA-256 e3d7c099b05d4bfce1ef595fb5e83d4de651300b1024ae04f112230c1960c123, built from ContainerNursery main dab37c78 — merged (main has since advanced with PR 652, test-only). Previous jar bin/container-nursery-a2fc6415-20260926T235430Z.jar kept for rollback (but a rollback must also use systemd-run). Root crontab now has no cn-watchdog line. Audit lines appended to /root/ContainerNursery/bin/deploy-audit.log.
  • Unchanged by this effort (all still the pre-fix builds): tinysite-wui.jar SHA-256 e24bd6161d4ebf11436e87d0e59e0e4cfe152d491e0357c8961ed2362b3d160a (built from TinySiteWui tinysite-initial ~dcae604, not on main); filedrop-server.jar 30403d0a16b740e41a8d0d38adc03d4be39698f7d5eee2c4d84468d727fd7f1b (FiledropEmbedded 0.1.11); simplefilesystemservice-server-vnext.jar a78ff4df6437940c68ff7c79c10c69ef77703fece5dd5aaf6d9ff794394edfb7 (resolver 0.0.1095).

Open production issue at write time

https://filedrop.wasmserver.com/ returns HTTP 500 (2026-10-09 01:59Z): filedrop-server's init v2_openFilesystem to simple-filesystem-vnext failed with AmbiguousRpcRequestException ... transport failed after the request write began ... Stream closed while reading message data (read 0 of 4 bytes) at ~01:57Z, and every later listDrops fails with the cached error. Timeline: SFS cold-started 8 times since 23:02Z (idle-stopped/started); at 01:56:07 filedrop started before SFS was ready (Attempt 1/6 failed: All 1 peers ... failed bytecode fetch), connected 01:57:14, then the open failed. FiledropServiceServer PR 26 ("retry filedrop initialization after a failed attempt instead of caching the failure forever") fixes the caching half; why the SFS transport closed mid-request is not yet diagnosed — investigate (investigate-broken-service skill) before restarting filedrop. tinysite.wasmserver.com /health was 200 at the same time.

Next steps

  1. Get the owner's choice for the PR 1226 design (options above). Then implement on PR 1226 with fail-first tests per the review findings, triage the 7 kotlin.build (remote) failures on run d6c2e103, full local suite, CI green, publish the next unused resolver version (verify on kotlin.directory first; 0.0.1306 is taken).
  2. Review and land UrlResolver PR 1227 (sol + Opus on head 0d832e1, final review, merge).
  3. SFS PR 25: bump resolver to 0.0.1306 (align POM pins as before), add tests for the two round-2 sol findings (late registration after host close; close-failure retry), local suite, CI, sol + Opus re-review, merge, then deploy simple-filesystem-vnext (pull main, build fat jar, upload/update route via container-nursery-deploy skill; verify the route's process cgroup is CN's unit, not hcf).
  4. Filedrop: round-4 fixes on FiledropEmbedded PR 14 for the round-3 findings (cross-process listing rebuild overwriting a newer projection; mixed 0.1.11/0.1.16 writers during upgrade — define the upgrade/rollout story, e.g. revision-aware repair or a migration gate; revision guard test or removal; README accuracy), publish next unused version, repoint FiledropServiceServer PR 27, re-review, merge both, deploy filedrop-server. Resolve the open production 500 (step above) — land PR 26 too once SFS PR 24 + client 0.1.7 are published and its red test is root-caused.
  5. TinySite: round-3 fixes on PR 3/PR 1 for round-2 findings (truncated commit response → no abort; relay RELAY_FORWARD_FAILED/RELAY_FORWARD_RETRYABLE classification — fix isServerError upstream in UrlResolver; PR 1's post-commit "400"; NOT_SENT test; non-JSON 502/504; cross-site confirmation), bump resolver once PR 1226 publishes and add the sibling site-creation timeout case, re-review, merge PR 1 then PR 3, deploy tinysite-wui (merge to main first; pull before build).
  6. After deploys: verify end-to-end — publish a multi-file site via ~/.claude/skills/good-ui-designs/scripts/tinysite-deploy.sh, confirm no 502/503, no v2_delete storms in SFS logs, __stream_close ≈ v2_readChunkStream, SFS thread count flat.
  7. Resume the sazabi.com capture (owner's original request): from the WIP branch, deploy dark-landing-sazabi/ to TinySite (tinysite-deploy.sh <dir> --title dark-landing-sazabi), test the deployed site (WebGPU in headless needs the SwiftShader flags in gpu.js), give the owner the link and wait for approval; only then remove _wip-sazabi-capture-tools/, add the repo README entry, open the PR (never put the TinySite link in the PR).
  8. Follow-ups recorded, not started: UrlResolver idle reaper closing a sandbox while returned objects are held; CN heap back to 128m (analyze the ~180 MB live set: 85 MB byte[]); buildtest repeated "chunked upload interrupted" / lost check-suite dispatch faults; buildtest CPU admission in OBSERVE mode over-admitting.

Operational knowledge

  • HCF CLI: coursier launch community.kotlin.hardwarecontrolfabric:cli:0.2.6 -r https://kotlin.directory -- -H 198.199.106.165 -c ~/.config/hardware-control-fabric/client_cert.pem -k ~/.config/hardware-control-fabric/fabric_key.pkcs8 --ca-cert ~/.config/hardware-control-fabric/server_cert.pem <cmd> (the skill's default cert filenames are stale). launch returns no stdout: write a script with write-file, run it with launch --cmd /bin/sh --args /tmp/x.sh, write output to a text file, read-file it. read-file mangles binary and refuses files over ~2 MB; write-file occasionally reports success without the file landing — re-check.
  • CN restart recipe (from an HCF script): stop old CN (SIGTERM, wait), then systemd-run --unit=containernursery-<tag>-<ts> -p KillMode=process -p WorkingDirectory=/root/ContainerNursery -p OOMScoreAdjust=-500 --setenv=CN_RESTART_HANDOFF=1 /bin/sh -c 'exec java -Xmx768m -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/root/ContainerNursery/heapdumps/ -jar bin/<jar> config.json >> /root/ContainerNursery/cn-<tag>-<ts>.log 2>&1'; health curl --resolve api.nursery.wasmserver.com:443:127.0.0.1 https://api.nursery.wasmserver.com/health; verify grep memory /proc/<pid>/cgroup. Build CN with scripts/build.bash containernursery.buildFatJar <out>.jar (no published artifact for main). The ssh-restart-containernursery tool hard-codes -Xmx128m — don't use it while 768m stands.
  • CN container logs: container-logs --route-key '<key>' where keys look like url:filedrop:, url:simple-filesystem-vnext:, https:tinysite.wasmserver.com:443; the API returns only the newest 10,000 lines and ignores --since/--until beyond that.
  • Find a CN app process: grep -l '<ENV>=<value>' /proc/[0-9]*/environ (e.g. FILESYSTEM_UUID=7c680d86 for filedrop) or by listening port from containers; CN launches jars from copied paths so name grep fails.
  • Codex (only if a reviewer wants gpt-6.1-sol per the owner's merge rule): codex exec --model gpt-6.1-sol --dangerously-bypass-approvals-and-sandbox -c model_reasoning_effort='"high"' -C <dir> -o <final.md> "$(cat prompt.md)" </dev/null; review prompts used are in the artifacts branch reviews/*.md (shared rules in reviews/review-common.md).
  • Headless Chromium on the /code box (aarch64): Playwright chromium_headless_shell-1228 + Debian bookworm arm64 libs extracted into ~/.cache/chromium-deps (apt lists are empty; fetch .debs from deb.debian.org Packages.gz); WebGPU via --enable-unsafe-webgpu --enable-features=Vulkan,WebGPUService --use-vulkan=swiftshader --use-webgpu-adapter=swiftshader --enable-unsafe-swiftshader --use-angle=swiftshader.
  • Several CI runs failed on buildtest infrastructure (chunked upload interrupted, tarball download timeout, lost check-suite dispatch); each was re-requested once after confirming zero tests ran.

No status reports yet.

Add dependency

Complete this handoff

Moves it out of every priority list and into ArchiveArea.