← Challenges

Repository · challenges

kotlin.directory maven-server intermittently parks requests >30s under CI load, surfacing as 503s

View on GitHub ↗

kotlin.directory maven-server intermittently parks requests >30s under CI load, surfacing as 503s

Reported (UTC): 2026-07-15 15:09

kotlin.directory maven-server intermittently parks requests >30s under CI load, surfacing as 503s to coursier on buildtest shards

What was being attempted: driving the buildtest CI flakes/parallelism handoff to completion; investigating why builds miss the 10-minute time-to-first-test SLO. Shard build.logs showed 'Server returned HTTP response code: 503 for URL: https://kotlin.directory/...'.

What went wrong: ContainerNursery's log (host 198.199.106.165, /root/ContainerNursery/cn-stdout-restart-20260715.log) recorded 6,028 'HTTP_ERROR host=kotlin.directory ... error=SocketTimeoutException: Read timed out duration���30000ms' entries in ~96 minutes (2026-07-15 13:26���15:02 UTC, steady 13���93/minute), mostly HEAD requests for .pom/.sha1 paths, while shard droplets drove ~100���600 req/s of cold-cache coursier resolution. This is NOT constant overload: direct backend probes (curl -I http://localhost:40189/...) answer in 6ms while the errors are ongoing, and the maven-server process sits at ~15% CPU ��� intermittent whole-request parks (~0.2% of requests), always pinned at ContainerNursery's 30s proxy read-timeout. The server is directory.kotlin.www.server (Jetty 11.0.15, JDK 21) deployed as /root/maven-server.jar behind route https:kotlin.directory:443 with memoryLimitMb=512 and JAVA_TOOL_OPTIONS -XX:+UseSerialGC -Xss512k. Prime suspects: SerialGC full-GC pauses on the small heap, Jetty thread-pool saturation bursts on blocking file I/O, or a shared lock on a hot path.

Impact: every cold-cache shard's coursier resolution intermittently fails/retries; kompile 0.0.122's progress-aware backstop converts the failures into added minutes per build, blowing the 10-minute time-to-first-test SLO on large suites. Before kompile 0.0.122 these surfaced as deterministic-looking 'Failed to resolve kompile:build-kotlin-jvm:0.0.23' hard failures.

Workaround used: re-baking the buildtest warm droplet image with a current-artifact coursier cache (removes the stampede load source) ��� load reduction, not a fix.

Suggested durable fix: root-cause and fix upstream in directory.kotlin.www.server via a hermetic concurrent HEAD/GET load test at production JVM settings that reproduces the >30s park, then fix the mechanism (a codex gpt-5.6-sol TDD investigation is in flight as of this filing). Secondary observations for the ContainerNursery owner: the request-logs API returned '0/0 lines' for this route while the main log recorded the errors, and DEBUG-level per-request logging writes ~2000 lines/second under this load, making CLI log forensics impractical.


Production verification — 2026-07-17

Status: STILL EXISTS. A production challenge merged at 03:15 UTC after the original audit reproduced this failure mode under concurrency: multiple Coursier TCP connections to 198.199.106.165:443 remained in SYN_SENT through the 60-second attempt budget while isolated HTTPS probes succeeded. That evidence invalidates the audit's sequential 20/20 probe as proof of recovery. This record is restored because concurrent repository access—the workload that CI and Coursier actually generate—remains unreliable; diagnosis should focus on the front-door listener, SYN backlog, firewall/conntrack state, and concurrent accept capacity.