Repository · handoffs
id: hf-2026-10-09-harden-the-handoff-service-after-the-client-bytecode-oom-cascade-bound-urlprotocol-tcpserver-outbox-bytes-fix-the-handoffembedded-snapshot-timeout-stop-holding-every-body-in-memory url: url://handoff/handoffs/hf-2026-10-09-harden-the-handoff-service-after-the-client-bytecode-oom-cascade-bound-urlprotocol-tcpserver-outbox-bytes-fix-the-handoffembedded-snapshot-timeout-stop-holding-every-body-in-memory title: Harden the handoff service after the client-bytecode OOM cascade: bound UrlProtocol TcpServer outbox bytes, fix the HandoffEmbedded snapshot timeout, stop holding every body in memory summary: Three secondary defects from the 2026-10-08/09 handoff-service OutOfMemoryError cascade remain after the primary per-request bytecode-encoding fix: UrlProtocol TcpServer admits requests by count with no bound on queued response bytes (outbox hit 22.0 MB against 18.3 MB); HandoffEmbedded's supervisor-snapshot timeout did not fire while a relay read blocked the dashboard for 762 s, so the WUI 503s on a healthy JVM; all 966 handoff bodies (~20.6 MB UTF-16) stay resident. Each needs a fail-first test and an upstream fix, then one operator-approved deploy. created: 2026-10-09T01:03:48.411Z completed: null dependencies:
Harden the handoff service against the client-bytecode OutOfMemoryError cascade: bound UrlProtocol TcpServer outbox bytes, fix the HandoffEmbedded supervisor-snapshot timeout, and stop holding every handoff body in memory
RE-VERIFY before acting: everything below is a snapshot written 2026-10-09 00:55 UTC. Re-check the live service (handoff-cli health, ContainerNursery route url:handoff: container logs), the PR list of HandoffServiceServer, UrlProtocol and HandoffEmbedded, and current main of each before trusting any claim here.
Mission
On 2026-10-08/09 the production handoff service (url://handoff/, ContainerNursery route url:handoff:, jar handoff-service-server-03da8e78-20261007.jar, default -Xmx128m -XX:+ExitOnOutOfMemoryError) exited with java.lang.OutOfMemoryError: Java heap space three times (22:18:44, 23:56:42, 00:30:12 UTC), each process shorter-lived than the last (6.5 h, 1.6 h, 33 min). While down, orchestrators could not read, claim or heartbeat handoffs and the WUI returned 503. The primary cause (per-request base64 encoding of the 2.75 MB client bytecode, ~10 MB of temporaries per request, nothing cached, reconnect storms after each restart) is being fixed in HandoffServiceServer by a separate lane (see Relevant PRs). This handoff covers the three secondary defects the forensic analysis exposed, which the primary fix does not address.
What was found (evidence)
- Crash heap dump (
java_pid4076742.hprof, copy kept by the forensic lane) showedTcpServerOutboxAccounting.currentBytes= 21,969,390 = 6 × 3,661,565: six unsent frames of the same bytecode response queued, outbox past its allowance (22.0 MB against 18.3 MB) while new requests kept being admitted. UrlProtocolTcpServeradmits by handler count (concurrentHandlerLimit256) and does not bound response bytes in flight. - Live thread dump:
getDashboardProjectionJsonblocked >155 s inHandoffEmbedded.snapshotThreads, waiting on the AI CLI supervisor manager'sgetFleetSnapshotJson, itself stuck 762 s in a relay read.HandoffEmbedded.supervisorSnapshotTimeoutMsexists to cut this off and did not fire. WUI pages therefore 503 even with a healthy JVM (AmbiguousRpcRequestException ... getDashboardProjection ... read 0 of 4 bytes). - Working set: 966 handoff bodies held in memory as UTF-16 strings, ~20.6 MB;
state.jsongrew from 2.9 MB (August) to 12.8 MB. - Host is not the cause: load 33–42 on 8 vCPU, 3.5 GB RAM available, disk 82%, no kernel OOM kills since Oct 2.
Next steps
- UrlProtocol
TcpServer: add admission by outbox bytes (refuse or defer new handler admission while queued response bytes exceed the allowance), with a fail-first public-API test that sends N concurrent large responses to a slow reader and asserts the outbox never exceeds its allowance and the N+1th request is deferred, not admitted. Publish a new UrlProtocol version only after review; bump HandoffServiceServer's pin afterwards. - HandoffEmbedded: write a fail-first test in which the supervisor snapshot provider blocks indefinitely and assert
getDashboardProjectionJsonreturns withinsupervisorSnapshotTimeoutMs(use the Clock abstraction, no real sleeps); find why the existing timeout does not fire (the forensic thread dump suggests the wait is inside a relay read that the timeout does not cover) and fix it. - HandoffEmbedded: stop holding every handoff body resident; load bodies on demand from the store (or keep only summaries resident), with a test that memory for N bodies does not scale with body size when bodies are not read.
- After the primary HandoffServiceServer fix and these land: one production deploy of the handoff service (operator approval required), then verify no OOM exits for 24 h via ContainerNursery container logs.
Relevant PRs / refs
- Primary fix lane (HandoffServiceServer, encode the bytecode response once): PR to be opened by lane HSVC2 on 2026-10-09; check
gh pr list --repo CodexCoder21Organization/HandoffServiceServer --state open. - Forensic notes: session scratchpad
hq/out/HSVC1-findings.md(not durable); the challenge record https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-09-0040-handoff-service-url-handoff-unreachable-from-the-code.md describes the outage as observed by clients. - Deployed but not merged: nothing; production runs commit 03da8e7 of HandoffServiceServer (on main).
Operational knowledge
handoff-cli0.0.8 failures during the outage looked likeAmbiguousRpcRequestException: RPC request 'claimHandoff' ... transport failed after the request write beganandAll 0 peers for service 'handoff' failed bytecode fetch; a claim may have landed despite the error, so always re-readhandoff-cli claimsbefore assuming.- Peer 12D3KooWLMyXNfwhcX1YsiNx3hnjk3GGSfsU1fydRa8bzrE6scMT at 198.199.106.165:35000 is ContainerNursery's UrlFacade, not the handoff JVM;
[UrlFacade] RPC forward failed: nullin CN logs is the facade losing the local JVM.