Decide the UrlResolver sandbox-retirement design, then land the TinySite outage fixes and redeploy
RE-VERIFY: write-time snapshot 2026-10-09 ~02:00 UTC. Every PR state, CI result, review verdict and production fact below was re-checked at write time but will drift. Re-check each PR with gh pr view <full URL> --json state,headRefOid,mergeStateStatus,statusCheckRollup, branches with git ls-remote, production with HardwareControlFabric (see Operational knowledge) and coursier launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com health-detailed.
Mission
On 2026-10-08 the owner asked to capture https://www.sazabi.com/ into Good-UI-Designs (good-ui-designs skill). Publishing the capture to TinySite (https://tinysite.wasmserver.com) failed repeatedly (HTTP 502/503, chunk 0 is missing), and the owner said: "What is going wrong with tinysite? If we are running into a problem, we need to fix it." This handoff carries that fix effort plus the paused sazabi capture.
Owner decisions/approvals already given (carry them forward; do not re-ask):
- Merging: a PR may be merged after both an adversarial review by gpt-6.1-sol (the owner's own instruction — via codex) and an independent Opus review come back clean on the exact final head, plus green CI and a final review.
- Deploys approved after merges for: ContainerNursery, simple-filesystem-vnext, filedrop-server, tinysite-wui. Always
git pull the latest main before building a deploy.
- ContainerNursery heap: keep -Xmx768m for now (live set measured ~180 MB after full GC on 2026-10-08, above the stored 128m directive); investigate getting back to 128m separately.
- The ContainerNursery watchdog cron (
/root/cn-watchdog.sh every minute) was removed per the owner's no-auto-restart directive (backup /root/crontab.backup-20261008T150706Z). Do not reinstate.
- Approved test changes: UrlResolver
testLazyTimeoutRetirementRejectionClosesSjvmExactlyOnce may change only its trigger (busy guest instead of interruptible RPC), assertions unchanged; FiledropEmbedded testPublishedChunkedPayloadRecoveryAfterMarkerCleanupFailure assertion updated to "a retried commit of an already-published upload returns the file".
- OPEN DECISION (owner asked for this handoff to enable it): UrlResolver PR 1226 design — see "Decision needed" below.
What was found (root-cause chain, verified)
- ContainerNursery (CN) URL facade leaked one backend TCP connection per streamed response drained to EOF. Deployed build a2fc6415 bundled UrlResolver 0.0.1203, which deregistered a finished response stream at EOF without closing it; the stream owned the CN→backend TCP connection, so it stayed open and the client's later
__stream_close could not reach it. Evidence: simple-filesystem-vnext (SFS) had 7,168 threads = 3,576 idle TcpServer-Reader/Writer pairs, all from 127.0.0.1 owned by the CN pid (ss -tnp), growing ~50–80/min; SFS logs showed 407 v2_readChunkStream and 0 __stream_close. Fail-first test: 40 EOF reads leave 40 backend connections open on a2fc6415, 0 on main. Fixed by https://github.com/CodexCoder21Organization/ContainerNursery/pull/632 (merged 2026-09-27, never deployed until 2026-10-08). Deployed and verified in production: after deploy, 27 reads → 27 __stream_close, 0 CN-held SFS connections, SFS 16 threads; TinySite file serving went from 12–30 s to 0.3–0.6 s.
- SFS (url://simple-filesystem-vnext/, -Xmx128m, UrlResolver 0.0.1095, embedded 0.0.18) OOM-crashed (
OutOfMemoryError: Java heap space 2026-10-08 10:33:04Z, hprof /root/ContainerNursery/heapdumps/java_pid3588199.hprof). Contributors: (a) each leaked connection ~12 KB heap (+ direct buffers); (b) SFS resolver 0.0.1095 retains streamed chunk data (up to 16 MiB each) in StreamAwareServiceHandler.pullableStreams until __stream_close/connection close — never arrived; (c) embedded 0.0.18 re-parses/rewrites a ~3 MB namespace-event file per mutation (~82 MB garbage each, bounded, not a leak). (a)/(b) neutralised in production by the CN deploy; SFS's own fix is PR 25 (+ resolver 0.0.1306 from PR 1227); (c) is PR 23.
- FiledropEmbedded 0.1.11 (filedrop-server): every metadata read deleted one session marker per committed chunked file (thousands of
v2_delete/min while serving/publishing); commit not idempotent (duplicate commit racing an in-flight one → chunk 0 is missing); construction always rewrote index.json; remote I/O under one global listingLock. Fixed in FiledropEmbedded PR 14 (several review rounds).
- TinySiteWui: one failed/timed-out call tore down and rebuilt the shared UrlResolver, failing unrelated concurrent requests; commit retry re-sent commits that might still be running; page aborted sessions on unknown outcomes. Fixed in TinySiteWui PR 3 (stacked on PR 1, the unmerged branch production runs).
- UrlResolver (upstream): (a) any sandboxed call deadline retires the whole SJVM, making objects returned by sibling calls unusable (
Cannot invoke method on closed instance proxy: filedrop/api/DropImpl.getId()) — a successful commit can answer 502 — PR 1226, design open; (b) StreamAwareServiceHandler shutdown not terminal and close failures swallowed — PR 1227, done, published 0.0.1306.
Ruled out: host capacity (host MemAvailable 13.7 GB; SFS heap analysis shows leak/churn, not under-provisioning); filedrop expired-drop purge as the delete-storm source (deletes were non-recursive v2_delete from the read path, not deleteRecursively).
Incident caused during this effort (resolved) — read before touching CN
Restarting CN from a script started via HardwareControlFabric launch put CN and all its app JVMs into hcf-daemon.service's cgroup (MemoryMax=512M, CPUQuota=120%). CN's memory guard then refused/killed containers ("RAM avail 0MB") → kotlin.directory, buildtest, kotlin.build, handoff, tinysite, filedrop etc. down ~15:07–15:48Z 2026-10-08. Rolling back the jar did not help. Fixed by relaunching CN in its own transient systemd unit. Recorded: https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-08-1743-restarting-containernursery-from-a-hardwarecontrolfabric.md. Never (re)start CN or app services directly from an HCF-launched shell; use systemd-run (recipe in Operational knowledge).
Decision needed: UrlResolver PR 1226 design
Problem: in UrlResolver 0.0.1302, when one sandboxed service-proxy call hits the SJVM per-call deadline (30 s), invokeSandboxMethod retires the whole SJVM (SandboxedProxyGenerator.kt timeout branch ~1396-1404, from https://github.com/CodexCoder21Organization/UrlResolver/pull/713), and every object previously returned by that sandbox (e.g. TinySite's DropImpl/DroppedFile) then throws "Cannot invoke method on closed instance proxy". Reproducer: tests/testReturnedObjectSurvivesSiblingCallTimeout.kts on the PR branch (fails on main 4fe0615 with the production message); TinySite-side reproducer branch https://github.com/CodexCoder21Organization/TinySiteWui/tree/w3b/slow-storage-call-reproducer.
PR 1226's current implementation (head fc51e8a): if the timed-out call is inside ServiceBridge.rpc, interrupt only that call and keep the sandbox; otherwise retire. Both adversarial reviews (findings: reviews/out/ur1226-sol-findings.md, reviews/out/ur1226-opus-findings.md in the artifacts branch) showed holes: deadline firing in the post-RPC tail (effects, stream resolution, result marshaling) or a swallowed interrupt lets the abandoned guest keep computing in the live sandbox (probe: sandboxActiveAfterTimeout=true, worker at ~985 CPU-ms/s); effect-resume can send a new RPC after abandonment; the changed test leaves a CPU-burning guest thread running; the "no new RPC after abandonment" guard is untested. Also observed: retirement on main never actually stopped a runaway guest either (the interpreter has no cancellation points; SJVMImpl.close() only cancels a coroutine scope).
Lane analysis (from SJVM source, sandboxjvm main / 0.0.55 lineage): host exceptions thrown into the guest are catchable (SJVMThreadImpl.kt:466 converts to guest InternalError); the interpreter loop has no cancellation check; a host java.lang.Error would unwind uncatchably but skips guest finally/monitorexit, leaving monitors held (Monitor.kt has no per-thread release). So "abandoned call never does more work" (INV-A) and "guest locks/state stay consistent" (INV-B) cannot both be guaranteed inside UrlResolver alone.
Options:
- Retire for new calls only (orchestrator's recommendation, not yet chosen by the owner): keep main's "any deadline retires" rule, but retirement only (a) routes NEW calls to a fresh sandbox (LazyReconnectingMethodDispatcher already reconnects) and (b) refuses further host/remote calls from the retired SJVM; objects already returned stay readable on the retired SJVM until unreachable/idle-reaped (then the SJVM closes). Rationale: retirement never stopped runaway guests anyway, so keeping the retired SJVM's heap alive for already-returned value objects adds no new execution risk, needs no new time bound, and fixes the production symptom. Cost: retired SJVMs live as long as their returned proxies (needs a Cleaner/idle path); methods on returned objects that themselves need remote calls would fail after retirement.
- Bounded extra window, then retire (lane option 2): only the pure remote wait counts as "waiting on the service"; after the deadline every exit path throws into the guest,
enter()/effect-resume/stream callbacks refuse new host calls, and if the worker hasn't exited within one more method-deadline window the sandbox retires. Adds a new time bound (needs explicit owner approval under the timeout rules); a misbehaving guest can compute up to one window in a live sandbox.
- Upstream SJVM cancellation (lane option 1): per-thread cancellation checked at loop back-edges/calls + monitor-releasing unwind + handler detection in https://github.com/CodexCoder21Organization/sandboxjvm, then force-unwind when safe, retire otherwise. Correct and also fixes "retirement never stops a stuck guest", but large; guest-implemented locks (ReentrantLock) would still force retirement.
- Keep always-retire, fix only the client (TinySite copies fields out of returned objects immediately / tolerates failures). Contradicts the fix-upstream rule.
Whatever is chosen: the remaining asked-for work is ready (INV-C test hygiene: guest waits released in finally + assert worker exits; fail-first tests for post-RPC-tail deadline, abandoned-call RPC attempt, effect-resume after abandonment; regression test cleanup; README contract), then full suite, CI, publish next unused resolver version (0.0.1305 is declared on the branch but unpublished; 0.0.1306 is taken by PR 1227 — re-check kotlin.directory before publishing), then bump TinySiteWui and add the sibling site-creation case to its timed-out-call test. The idle-reaper variant (sandbox closed after 120 s idle while the caller holds returned objects) is a separate follow-up recorded in the PR.
Relevant PRs / refs (state at write time)
| Repo |
PR |
Branch @ head |
CI (kotlin.build (remote)) |
Review state |
Next |
| ContainerNursery |
https://github.com/CodexCoder21Organization/ContainerNursery/pull/632 |
merged 2026-09-27 |
green |
— |
deployed 2026-10-08 (see Deployed) |
| ContainerNursery |
https://github.com/CodexCoder21Organization/ContainerNursery/pull/652 |
test/url-facade-many-eof-reads-close-backend-connections @ 1fcb0f664217e7f716094ab6387e8fdb7aee55b7 |
green |
sol CLEAN + Opus CLEAN (round 2) |
MERGED 2026-10-08T18:14Z (merge commit 09e352ffeaa3df11de38d11d34c4eb027b3e574b); test-only, nothing to deploy |
| SimpleFileSystemServiceServer |
https://github.com/CodexCoder21Organization/SimpleFileSystemServiceServer/pull/25 |
fix/stream-release-bump @ ac8ba701fe343c10a55a117c0d986a2387ea3107 |
green (31/31 local) |
round 2: Opus CLEAN; sol FINDINGS(2 major: late stream registration after host close; swallowed stream-close failures) — both are UrlResolver defects now fixed by PR 1227/0.0.1306 |
bump resolver 0.0.1302→0.0.1306 (+POM alignment), add host-close-race + close-failure-retry tests, re-review (sol+Opus), merge, deploy SFS |
| SimpleFileSystemServiceServer |
https://github.com/CodexCoder21Organization/SimpleFileSystemServiceServer/pull/23 |
bump-embedded-0.0.26-upgrade-compatibility @ b5e408b3e2f231e0a43fd63ee930706713574288 |
red: all 42 tests passed but testDurableBackendLoopback exceeded @Timeout(60) at 62.5 s under CI CPU over-admission (buildtest log: CPU admission "OBSERVE mode" admitted 5573 millicores on a 2-vCPU droplet); evidence comment https://github.com/CodexCoder21Organization/SimpleFileSystemServiceServer/pull/23#issuecomment-6062919496 |
not reviewed this effort |
do NOT raise the timeout; reduce CPU cost of the PR's new upgrade tests or fix buildtest CPU admission; version collision with PR 25 (both claim 0.1.1x) → rebase whichever lands second |
| SimpleFileSystemServiceServer |
https://github.com/CodexCoder21Organization/SimpleFileSystemServiceServer/pull/24 |
feat/closeable-simplefilesystem-manager @ 26b0bd554be8c69d77fe4c835e61f5ef0e3bc0cd |
green |
not reviewed this effort |
needed by FiledropServiceServer PR 26 (requires publishing simplefilesystemservice-client 0.1.7) |
| FiledropEmbedded |
https://github.com/CodexCoder21Organization/FiledropEmbedded/pull/14 |
fix/reads-do-not-delete-session-markers @ 9c26a4a788c47182996280f69be1ed60869187fa |
green (88/88 local) |
round 3: Opus FINDINGS(2 minor: revision guard upsertNewerListingEntry ~1078-1083 untested/never fires; README overpromises listing freshness after non-retried uploadFile/deleteFile/createDrop listing-write failure); sol FINDINGS(2 major: listing rebuild can overwrite a newer projection across processes; overlapping 0.1.11/0.1.16 writers during upgrade invalidate revision ordering) — reviews/out/fdc-*-findings.md |
fix round 4, publish next unused version (0.1.14/0.1.15/0.1.16 claimed), repoint PR 27, re-review, merge |
| FiledropServiceServer |
https://github.com/CodexCoder21Organization/FiledropServiceServer/pull/27 |
fix/pin-filedrop-embedded-0.1.14 @ 48771cf2a10a104345e700a62c0cc259ac67b5bb |
green (24/24) |
reviewed with PR 14 (test fixture serialization judged a correct fixture fix) |
repoint to final embedded version; merge after PR 14; deploy filedrop-server |
| FiledropServiceServer |
https://github.com/CodexCoder21Organization/FiledropServiceServer/pull/26 |
fix/filedrop-init-retry-after-failure @ 0e2a3ad60ea92359be46c26110e6d96c2abdc457 |
red (testE2eSjvmColdStartInitializesAfterProtocolProbe, also flaky on main) |
not reviewed |
production-relevant (see Open production issue); blocked on simplefilesystemservice-client 0.1.7 (SFS PR 24) |
| TinySiteWui |
https://github.com/CodexCoder21Organization/TinySiteWui/pull/1 |
tinysite-initial @ 80225f405bc4e28fb60690504d724df5ad407503 |
green (14/14) |
round 2 findings below (shared with PR 3) |
production base (main is a skeleton); old handoff https://www.handoff.wasmserver.com → hf-2026-10-06-merge-the-tinysite-pull-request-once-the-owner-approves tracks it |
| TinySiteWui |
https://github.com/CodexCoder21Organization/TinySiteWui/pull/3 |
fix/storage-call-failure-does-not-fail-concurrent-requests @ 31b10f24ef51c4e5cbbbecc5c17afa1aaf2491e4 (stacked on PR 1) |
green (18/18) |
round 2: Opus FINDINGS(1 major: relay RELAY_FORWARD_FAILED after dispatch classified as service-answered → plain 502 + page abort; minors: PR 1's "refused 400" happens after commit; untested NOT_SENT branch; slow page SHA-256 ~27 MB/s; 502/504 without JSON body treated as definitive); sol FINDINGS(1 major: commit response truncated after headers still triggers abort; 1 minor: confirmation can match another site's older file) — reviews/out/tsb-*-findings.md |
fix round 3; also the RELAY_FORWARD_FAILED isServerError classification should be fixed upstream in UrlResolver; then bump resolver once PR 1226 is published |
| UrlResolver |
https://github.com/CodexCoder21Organization/UrlResolver/pull/1226 |
fix/returned-object-survives-sibling-timeout @ fc51e8aa63ee9bce304dd0581ac3ef50cef4bd46 |
red: run https://buildtest.kotlin.build/run?id=d6c2e103 1885/1892 (7 failed — not yet triaged); bld-build failed on testDeadConfiguredBootstrapDialsBackOffExponentially (pre-existing intermittent) |
FINDINGS from both reviewers (round 1) |
owner design decision, then implement + tests, triage the 7 CI failures, re-review, publish |
| UrlResolver |
https://github.com/CodexCoder21Organization/UrlResolver/pull/1227 |
fix/stream-handler-terminal-shutdown-and-cleanup-failures @ 0d832e1485d482f5da58668e62977ddd15c24b0f |
green (1897/1897) |
not yet reviewed |
sol + Opus review, merge; foundation.url:resolver:0.0.1306 already published from this head |
Branches pushed for this handoff (no PR)
| Repo |
Branch |
Remote head SHA |
PR |
What is on it |
State |
| CodexCoder21Organization/PlanRepository |
handoff-artifacts/2026-10-08-tinysite-outage |
8bad7f571e83b0f20fef397472d4d48033c13984 |
no PR |
handoffs/artifacts/2026-10-08-tinysite-outage/: investigation scripts + remote outputs (jstack summaries, ss/cgroup/memory snapshots, CN live histo), gzipped service logs, every adversarial review prompt + findings + final verdicts, reviewer probe tests and mutation patches, lane notes/findings for all lanes, delegation briefs, prior root-cause findings (agentA SFS heap, agentB filedrop, agentC CN leak + 40-read test) |
evidence only; raw codex transcripts and duplicate logs omitted (large, no extra content) |
| CodexCoder21Organization/Good-UI-Designs |
wip/dark-landing-sazabi-handoff-2026-10-09 |
1659dbe69853c082c55c389152c02d5c28641d39 |
no PR |
dark-landing-sazabi/ (captured site, screenshot, README) + _wip-sazabi-capture-tools/ (record/build/serve/verify scripts; delete before PR) |
sub-path HTTP copy verified against live site; not owner-approved; not deployed to TinySite; raw 36 MB recording not committed (regenerate with record.js) |
| CodexCoder21Organization/TinySiteWui |
w3b/slow-storage-call-reproducer |
973ea0c2123c9b7e5dfc2368bdce50c7f6c9c456 |
no PR |
reproducer for the UrlResolver returned-object defect |
reproducer only |
All lane clones were swept (git status, git log --branches --not --remotes, stashes): no unpushed commits; the only uncommitted diffs were reviewers' deliberate mutation/probe worktrees, saved as patches under reviewer-mutations/ in the artifacts branch.
Published artifacts (verified on kotlin.directory by SHA-256)
filedrop.embedded:filedrop-embedded:0.1.14 — 8ccf122e7c67e1da704ae14aea2eb57a7a5ec80bc4e1b40355cd97c95c194e17 (round-1 head 981fd49; superseded, contains round-1 defects — do not use)
filedrop.embedded:filedrop-embedded:0.1.15 — a1059bfa0a33d81f5d481037d5fa79f84ff55a7c91914d907fa8aa9d96a252b8 (round-2 head 7320540; superseded)
filedrop.embedded:filedrop-embedded:0.1.16 — cd81f30bf20a1e455fcc84005dd779f35bf20185af0a83c9da435fe38a9b6b37 (round-3 head 9c26a4a; has open round-3 review findings)
foundation.url:resolver:0.0.1306 — 0d134fcb961e8ef1030f18d4a409d63fcfb20ccde31be2f6c33efe8e72d73dbb (UrlResolver PR 1227 head 0d832e1; PR not yet reviewed/merged)
foundation.url:resolver:0.0.1305 — not published (declared on PR 1226's branch)
Deployed — what is live (verified 2026-10-09 01:59Z via HCF)
- ContainerNursery pid 1368870, transient systemd unit
containernursery-dab37c78-20261008T154648Z.service (KillMode=process, no Restart=), cmd java -Xmx768m -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/root/ContainerNursery/heapdumps/ -jar bin/container-nursery-dab37c78-20261008T1454Z.jar config.json (cwd /root/ContainerNursery), jar SHA-256 e3d7c099b05d4bfce1ef595fb5e83d4de651300b1024ae04f112230c1960c123, built from ContainerNursery main dab37c78 — merged (main has since advanced with PR 652, test-only). Previous jar bin/container-nursery-a2fc6415-20260926T235430Z.jar kept for rollback (but a rollback must also use systemd-run). Root crontab now has no cn-watchdog line. Audit lines appended to /root/ContainerNursery/bin/deploy-audit.log.
- Unchanged by this effort (all still the pre-fix builds): tinysite-wui.jar SHA-256 e24bd6161d4ebf11436e87d0e59e0e4cfe152d491e0357c8961ed2362b3d160a (built from TinySiteWui
tinysite-initial ~dcae604, not on main); filedrop-server.jar 30403d0a16b740e41a8d0d38adc03d4be39698f7d5eee2c4d84468d727fd7f1b (FiledropEmbedded 0.1.11); simplefilesystemservice-server-vnext.jar a78ff4df6437940c68ff7c79c10c69ef77703fece5dd5aaf6d9ff794394edfb7 (resolver 0.0.1095).
Open production issue at write time
https://filedrop.wasmserver.com/ returns HTTP 500 (2026-10-09 01:59Z): filedrop-server's init v2_openFilesystem to simple-filesystem-vnext failed with AmbiguousRpcRequestException ... transport failed after the request write began ... Stream closed while reading message data (read 0 of 4 bytes) at ~01:57Z, and every later listDrops fails with the cached error. Timeline: SFS cold-started 8 times since 23:02Z (idle-stopped/started); at 01:56:07 filedrop started before SFS was ready (Attempt 1/6 failed: All 1 peers ... failed bytecode fetch), connected 01:57:14, then the open failed. FiledropServiceServer PR 26 ("retry filedrop initialization after a failed attempt instead of caching the failure forever") fixes the caching half; why the SFS transport closed mid-request is not yet diagnosed — investigate (investigate-broken-service skill) before restarting filedrop. tinysite.wasmserver.com /health was 200 at the same time.
Next steps
- Get the owner's choice for the PR 1226 design (options above). Then implement on PR 1226 with fail-first tests per the review findings, triage the 7
kotlin.build (remote) failures on run d6c2e103, full local suite, CI green, publish the next unused resolver version (verify on kotlin.directory first; 0.0.1306 is taken).
- Review and land UrlResolver PR 1227 (sol + Opus on head 0d832e1, final review, merge).
- SFS PR 25: bump resolver to 0.0.1306 (align POM pins as before), add tests for the two round-2 sol findings (late registration after host close; close-failure retry), local suite, CI, sol + Opus re-review, merge, then deploy simple-filesystem-vnext (pull main, build fat jar, upload/update route via container-nursery-deploy skill; verify the route's process cgroup is CN's unit, not hcf).
- Filedrop: round-4 fixes on FiledropEmbedded PR 14 for the round-3 findings (cross-process listing rebuild overwriting a newer projection; mixed 0.1.11/0.1.16 writers during upgrade — define the upgrade/rollout story, e.g. revision-aware repair or a migration gate; revision guard test or removal; README accuracy), publish next unused version, repoint FiledropServiceServer PR 27, re-review, merge both, deploy filedrop-server. Resolve the open production 500 (step above) — land PR 26 too once SFS PR 24 + client 0.1.7 are published and its red test is root-caused.
- TinySite: round-3 fixes on PR 3/PR 1 for round-2 findings (truncated commit response → no abort; relay
RELAY_FORWARD_FAILED/RELAY_FORWARD_RETRYABLE classification — fix isServerError upstream in UrlResolver; PR 1's post-commit "400"; NOT_SENT test; non-JSON 502/504; cross-site confirmation), bump resolver once PR 1226 publishes and add the sibling site-creation timeout case, re-review, merge PR 1 then PR 3, deploy tinysite-wui (merge to main first; pull before build).
- After deploys: verify end-to-end — publish a multi-file site via
~/.claude/skills/good-ui-designs/scripts/tinysite-deploy.sh, confirm no 502/503, no v2_delete storms in SFS logs, __stream_close ≈ v2_readChunkStream, SFS thread count flat.
- Resume the sazabi.com capture (owner's original request): from the WIP branch, deploy
dark-landing-sazabi/ to TinySite (tinysite-deploy.sh <dir> --title dark-landing-sazabi), test the deployed site (WebGPU in headless needs the SwiftShader flags in gpu.js), give the owner the link and wait for approval; only then remove _wip-sazabi-capture-tools/, add the repo README entry, open the PR (never put the TinySite link in the PR).
- Follow-ups recorded, not started: UrlResolver idle reaper closing a sandbox while returned objects are held; CN heap back to 128m (analyze the ~180 MB live set: 85 MB
byte[]); buildtest repeated "chunked upload interrupted" / lost check-suite dispatch faults; buildtest CPU admission in OBSERVE mode over-admitting.
Operational knowledge
- HCF CLI:
coursier launch community.kotlin.hardwarecontrolfabric:cli:0.2.6 -r https://kotlin.directory -- -H 198.199.106.165 -c ~/.config/hardware-control-fabric/client_cert.pem -k ~/.config/hardware-control-fabric/fabric_key.pkcs8 --ca-cert ~/.config/hardware-control-fabric/server_cert.pem <cmd> (the skill's default cert filenames are stale). launch returns no stdout: write a script with write-file, run it with launch --cmd /bin/sh --args /tmp/x.sh, write output to a text file, read-file it. read-file mangles binary and refuses files over ~2 MB; write-file occasionally reports success without the file landing — re-check.
- CN restart recipe (from an HCF script): stop old CN (SIGTERM, wait), then
systemd-run --unit=containernursery-<tag>-<ts> -p KillMode=process -p WorkingDirectory=/root/ContainerNursery -p OOMScoreAdjust=-500 --setenv=CN_RESTART_HANDOFF=1 /bin/sh -c 'exec java -Xmx768m -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/root/ContainerNursery/heapdumps/ -jar bin/<jar> config.json >> /root/ContainerNursery/cn-<tag>-<ts>.log 2>&1'; health curl --resolve api.nursery.wasmserver.com:443:127.0.0.1 https://api.nursery.wasmserver.com/health; verify grep memory /proc/<pid>/cgroup. Build CN with scripts/build.bash containernursery.buildFatJar <out>.jar (no published artifact for main). The ssh-restart-containernursery tool hard-codes -Xmx128m — don't use it while 768m stands.
- CN container logs:
container-logs --route-key '<key>' where keys look like url:filedrop:, url:simple-filesystem-vnext:, https:tinysite.wasmserver.com:443; the API returns only the newest 10,000 lines and ignores --since/--until beyond that.
- Find a CN app process:
grep -l '<ENV>=<value>' /proc/[0-9]*/environ (e.g. FILESYSTEM_UUID=7c680d86 for filedrop) or by listening port from containers; CN launches jars from copied paths so name grep fails.
- Codex (only if a reviewer wants gpt-6.1-sol per the owner's merge rule):
codex exec --model gpt-6.1-sol --dangerously-bypass-approvals-and-sandbox -c model_reasoning_effort='"high"' -C <dir> -o <final.md> "$(cat prompt.md)" </dev/null; review prompts used are in the artifacts branch reviews/*.md (shared rules in reviews/review-common.md).
- Headless Chromium on the /code box (aarch64): Playwright
chromium_headless_shell-1228 + Debian bookworm arm64 libs extracted into ~/.cache/chromium-deps (apt lists are empty; fetch .debs from deb.debian.org Packages.gz); WebGPU via --enable-unsafe-webgpu --enable-features=Vulkan,WebGPUService --use-vulkan=swiftshader --use-webgpu-adapter=swiftshader --enable-unsafe-swiftshader --use-angle=swiftshader.
- Several CI runs failed on buildtest infrastructure (chunked upload interrupted, tarball download timeout, lost check-suite dispatch); each was re-requested once after confirming zero tests ran.