← Priority list
Blocked

Stop the Filedrop home page failing when the server's storage connection drops mid-request, and exercise the new duplicate/eager-upload behaviour live

Remaining: about one Filedrop home-page load in four returns HTTP 500 because the server's persistent connection to the storage service drops mid-request (AmbiguousRpcRequestException on varying storage calls); find and fix the cause in the connection layer, then exercise duplicate handling and background upload on the live site. Needs the user's go-ahead as a new effort. State: all requested changes and the startup-failure fix are merged and deployed from main (server a6d91da, web UI 9cda542); nothing is deployed that main lacks.

Needs the user's decision on whether to pursue the storage-connection drops that still fail about one Filedrop page load in four

Handoff document

Markdown

Stop the Filedrop home page failing when the server's storage connection drops mid-request

Written 2026-10-09 21:25 UTC. RE-VERIFY: everything below is a write-time snapshot. Re-check before acting: gh pr view <n> --repo CodexCoder21Organization/<repo> --json state,mergedAt; curl -s -o /dev/null -w '%{http_code} %{time_total}s\n' https://filedrop.wasmserver.com/ (repeat ten times); sha256sum of the two jars on the host (paths below).

Mission summary

The user asked for three things on the Filedrop website (https://filedrop.wasmserver.com/):

  1. "when dropping files into the upload files box, we should properly handle duplicate files: If a file being dropped in has the same name and checksum as an existing file, the duplicate should be ignored/dropped (just keep the one copy we already have). If the filename is different but the checksum is the same, probably show some kind of warning but allow the second (duplicate) file."
  2. "as soon as we drag the file in, it should silently start uploading in the background ... Ideally, it is fully uploaded by the time a user clicks create drop ... If the upload is still in progress, then we can show the existing progress bar when they click create drop."
  3. "merge the changes and deploy."

All three are delivered: merged and deployed. What remains is a fault that predates this work and still makes roughly one page load in four fail, plus a live exercise of the new behaviour that was deliberately not done on production.

What was found and done

  1. Requested behaviour: merged and deployed. Five pull requests (table below). The web UI jar built from FiledropWui main 9cda542 is live. A headless-browser capture of the home page at 21:17 UTC shows a clean first frame with the new drop-zone text ("Files dropped anywhere on this page land here, and start uploading straight away — press Create Drop when you are ready").
  2. Startup failure was permanent (fixed, merged, deployed). The Filedrop server connects to the storage service (url://simple-filesystem-vnext/) on the first request after a cold start. It stored a failed start and replayed the same error to every later request until the container went idle for 300 seconds. Fixed in https://github.com/CodexCoder21Organization/FiledropServiceServer/pull/28 (merged 21:06 UTC, four independent review rounds, last one with no findings). Deployed 21:11 UTC. Observed working on the first cold start after deploy: log shows Attempt 1/6 failed at 21:12:25, Connected to filesystem service at 21:12:30, and the page returned 200 from 21:12:37.
  3. Remaining fault: storage connection drops mid-request (NOT fixed). After a successful start the home page still returns HTTP 500 intermittently: 8 of 27 probes between 21:12 and 21:18 UTC. Each failure in the web UI log is RPC request 'listDrops' failed caused by PersistentRpcConnection$AmbiguousRpcRequestException: RPC request '<storage call>' to service 'simple-filesystem-vnext' has an ambiguous outcome because its transport failed after the request write began. The request was not replayed because the remote handler may already have run. The storage call differs each time (v2_metadataOrNull 6, v2_readStart 4, v2_readClose 2, v2_readUtf8 2, v2_deleteRecursively 2 log lines in that window), so it is the connection that is dropping, not one call misbehaving. One listDrops makes many storage calls, so a single drop fails the whole page. Failures clear on their own within seconds to a minute; nothing is stored.
    • Not yet known: why the persistent connection between two services on the same host drops this often. The host (198.199.106.165) had a load average of 37 at 21:11 UTC. Per the standing engineering rules this is a code defect to find (in the connection layer or in whatever is consuming the host), not a capacity matter. Other sessions have open UrlResolver work in this area; check before starting.
    • Ruled out: the stored-startup-failure mechanism (fixed, and these failures do not repeat identically or persist); a fault in one particular storage call.
    • Do not paper over this in Filedrop with retries or fallbacks; read-only storage calls being refused replay is a question for the connection layer's contract.
  4. Slow test in FiledropWui. testFiledropClientChunkTransportReplay ran 33.0 s against its 30 s limit once on a merge-queue machine (four test lanes on two CPUs) and evicted a pull request. Investigation: it does 8–11 CPU-seconds of real work near the limit; 0 of 7 local runs failed; no hang. No change made. Options: split the test (recommended), or raise the limit (needs the user's explicit approval).
  5. CI service drops checks. kotlin-build-ci lost the required check on every push to FiledropServiceServer today, and once created a check that stayed "in progress" for 30 minutes with no build run behind it. Remedy that worked each time: curl https://github-webhooks.wasmserver.com/ three times, then gh api -X POST repos/<owner>/<repo>/check-suites/<id>/rerequest, then confirm a run for the commit appears in https://buildtest.kotlin.build/api/runs.

Relevant PRs / refs

Pull request State at write time
https://github.com/CodexCoder21Organization/FiledropApi/pull/5 MERGED 2026-10-06 22:12 UTC
https://github.com/CodexCoder21Organization/FiledropServiceServer/pull/25 MERGED 2026-10-09 16:26 UTC
https://github.com/CodexCoder21Organization/FiledropEmbedded/pull/13 MERGED 2026-10-09 16:29 UTC
https://github.com/CodexCoder21Organization/FiledropWui/pull/28 MERGED 2026-10-09 16:31 UTC
https://github.com/CodexCoder21Organization/FiledropWui/pull/29 MERGED 2026-10-09 19:03 UTC
https://github.com/CodexCoder21Organization/FiledropServiceServer/pull/28 MERGED 2026-10-09 21:06 UTC
https://github.com/CodexCoder21Organization/SimpleFileSystemServiceServer/pull/23 OPEN, required check failed ("42 passed … FAILED (resumed after restart)"). Storage-service upgrade; not covered by the user's merge instruction; untouched.
https://github.com/CodexCoder21Organization/kotlin-build-ci/pull/301 OPEN, another session's work on the CI service.
Repo Branch Remote head SHA PR What is on it State
CodexCoder21Organization/PlanRepository handoff-artifacts/filedrop-duplicates-eager-upload-2026-10-09 1a05c0a74f410081f81e761a58c59c85663893bd no PR Investigation notes under handoffs/artifacts/filedrop-2026-10-09/ (home-page 500 investigation, slow-test investigation, write latency, storage upgrade, CI dispatch) Notes only
CodexCoder21Organization/FiledropServiceServer fix/startup-failure-not-permanent c13768f269d347e4749e84c06101ebadd640ba42 https://github.com/CodexCoder21Organization/FiledropServiceServer/pull/28 The startup fix Merged; safe to delete
CodexCoder21Organization/FiledropServiceServer wip/startup-failure-not-permanent c13768f269d347e4749e84c06101ebadd640ba42 no PR Checkpoint copy of the same commit Superseded by main; safe to delete

All code work is on main; nothing unpushed (verified with git status --porcelain, git stash list and git log --branches --not --remotes in every checkout under /code/workspace).

Deployed state (verified with sha256sum on the host at 21:20 UTC). Nothing is deployed that main does not contain.

  • Filedrop server: route url:filedrop:, /root/ContainerNursery/apps/filedrop-server.jar, SHA-256 b56094241433bd1e229a7932bf9ad11a7ef7ab07b06239df66afca8b2a4eb0c3, built from FiledropServiceServer main a6d91da (server version 0.0.23). Version 0.0.23 is not published to kotlin.directory; FiledropWui still pins 0.0.22 for its tests, which is fine for the deploy.
  • Filedrop web UI: route filedrop.wasmserver.com:443, /root/ContainerNursery/apps/filedrop-wui.jar, SHA-256 52edfeb92c7237e7044205ac7bd3ea3537fbaf6f4f308253ffb00332f8f270d7, built from FiledropWui main 9cda542.
  • Rollback copies on the host from before this session: filedrop-server.jar.rollback-20261009 and filedrop-wui.jar.rollback-20261009 in the same directory.

Next steps

  1. Decision from the user: whether to pursue the storage-connection drops (item 3) as a new effort. That question was put to the user in the final status report.
  2. If yes: reproduce the mid-request connection loss in a test (two services, a persistent connection, many concurrent callers of the storage read path), find why the connection closes, and fix it in the layer the reproducer points at. Start from the notes in the artifacts branch and the web UI log lines quoted above.
  3. Exercise the new behaviour on the live site in a real browser: drop the same file twice (second is ignored), drop the same content under another name (kept, with a warning), confirm upload starts on drop and "Create Drop" is immediate when it has finished. This creates real drops on production, so it was not done; do it with a short expiry and delete the drops afterwards, or get the user's go-ahead first.
  4. Check that the preview boxes on the home page fill in (they were still grey placeholders in the capture taken at page load).
  5. Decide the slow test (item 4).

Reusable / operational knowledge

  • ContainerNursery CLI: ~/bin/cs launch containernurserycli:container-nursery-cli:0.0.20 -r https://kotlin.directory -- --url https://api.nursery.wasmserver.com <command>. Deploy with upload-jar --file <jar> --route "<key>" (it overwrites the route's jar and restarts it; copy a rollback aside first). Logs: container-logs --route-key "url:filedrop:" and --route-key "https:filedrop.wasmserver.com:443". Do not run routes --json: it prints route settings, including credentials.
  • Build the server: scripts/build.bash --local filedrop.server.buildFatJar <out.jar> (about 95 seconds).
  • A first request after 300 idle seconds is a cold start and takes 15–40 seconds; a 503 after 30 seconds during that window is the front door timing out, not the fault in item 3.
  • Challenge already filed for the home-page outage: https://github.com/CodexCoder21Organization/PlanRepository/blob/main/challenges/2026-10-09-1706-filedrop-home-page-served-http-500-for-long-stretches.md

No status reports yet.

Add dependency

Complete this handoff

Moves it out of every priority list and into ArchiveArea.