Harden the handoff service against the client-bytecode OutOfMemoryError cascade: bound UrlProtocol TcpServer outbox bytes, fix the HandoffEmbedded supervisor-snapshot timeout, and stop holding every handoff body in memory
RE-VERIFY before acting: everything below is a snapshot written 2026-10-09 00:55 UTC. Re-check the live service (handoff-cli health, ContainerNursery route url:handoff: container logs), the PR list of HandoffServiceServer, UrlProtocol and HandoffEmbedded, and current main of each before trusting any claim here.
Mission
On 2026-10-08/09 the production handoff service (url://handoff/, ContainerNursery route url:handoff:, jar handoff-service-server-03da8e78-20261007.jar, default -Xmx128m -XX:+ExitOnOutOfMemoryError) exited with java.lang.OutOfMemoryError: Java heap space three times (22:18:44, 23:56:42, 00:30:12 UTC), each process shorter-lived than the last (6.5 h, 1.6 h, 33 min). While down, orchestrators could not read, claim or heartbeat handoffs and the WUI returned 503. The primary cause (per-request base64 encoding of the 2.75 MB client bytecode, ~10 MB of temporaries per request, nothing cached, reconnect storms after each restart) is being fixed in HandoffServiceServer by a separate lane (see Relevant PRs). This handoff covers the three secondary defects the forensic analysis exposed, which the primary fix does not address.
What was found (evidence)
- Crash heap dump (
java_pid4076742.hprof, copy kept by the forensic lane) showed TcpServerOutboxAccounting.currentBytes = 21,969,390 = 6 × 3,661,565: six unsent frames of the same bytecode response queued, outbox past its allowance (22.0 MB against 18.3 MB) while new requests kept being admitted. UrlProtocol TcpServer admits by handler count (concurrentHandlerLimit 256) and does not bound response bytes in flight.
- Live thread dump:
getDashboardProjectionJson blocked >155 s in HandoffEmbedded.snapshotThreads, waiting on the AI CLI supervisor manager's getFleetSnapshotJson, itself stuck 762 s in a relay read. HandoffEmbedded.supervisorSnapshotTimeoutMs exists to cut this off and did not fire. WUI pages therefore 503 even with a healthy JVM (AmbiguousRpcRequestException ... getDashboardProjection ... read 0 of 4 bytes).
- Working set: 966 handoff bodies held in memory as UTF-16 strings, ~20.6 MB;
state.json grew from 2.9 MB (August) to 12.8 MB.
- Host is not the cause: load 33–42 on 8 vCPU, 3.5 GB RAM available, disk 82%, no kernel OOM kills since Oct 2.
Next steps
- UrlProtocol
TcpServer: add admission by outbox bytes (refuse or defer new handler admission while queued response bytes exceed the allowance), with a fail-first public-API test that sends N concurrent large responses to a slow reader and asserts the outbox never exceeds its allowance and the N+1th request is deferred, not admitted. Publish a new UrlProtocol version only after review; bump HandoffServiceServer's pin afterwards.
- HandoffEmbedded: write a fail-first test in which the supervisor snapshot provider blocks indefinitely and assert
getDashboardProjectionJson returns within supervisorSnapshotTimeoutMs (use the Clock abstraction, no real sleeps); find why the existing timeout does not fire (the forensic thread dump suggests the wait is inside a relay read that the timeout does not cover) and fix it.
- HandoffEmbedded: stop holding every handoff body resident; load bodies on demand from the store (or keep only summaries resident), with a test that memory for N bodies does not scale with body size when bodies are not read.
- After the primary HandoffServiceServer fix and these land: one production deploy of the handoff service (operator approval required), then verify no OOM exits for 24 h via ContainerNursery container logs.
Relevant PRs / refs
- Primary fix lane (HandoffServiceServer, encode the bytecode response once): PR to be opened by lane HSVC2 on 2026-10-09; check
gh pr list --repo CodexCoder21Organization/HandoffServiceServer --state open.
- Forensic notes: session scratchpad
hq/out/HSVC1-findings.md (not durable); the challenge record https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-09-0040-handoff-service-url-handoff-unreachable-from-the-code.md describes the outage as observed by clients.
- Deployed but not merged: nothing; production runs commit 03da8e7 of HandoffServiceServer (on main).
Operational knowledge
handoff-cli 0.0.8 failures during the outage looked like AmbiguousRpcRequestException: RPC request 'claimHandoff' ... transport failed after the request write began and All 0 peers for service 'handoff' failed bytecode fetch; a claim may have landed despite the error, so always re-read handoff-cli claims before assuming.
- Peer 12D3KooWLMyXNfwhcX1YsiNx3hnjk3GGSfsU1fydRa8bzrE6scMT at 198.199.106.165:35000 is ContainerNursery's UrlFacade, not the handoff JVM;
[UrlFacade] RPC forward failed: null in CN logs is the facade losing the local JVM.