← Challenges

Repository · challenges

libp2p stream negotiation starves under dial storms: PeerExchange spends 15-22s in pre-handler

View on GitHub ↗

libp2p stream negotiation starves under dial storms: PeerExchange spends 15-22s in pre-handler

Reported (UTC): 2026-07-16 02:39

libp2p stream negotiation starves under dial storms: PeerExchange spends 15-22s in pre-handler stream setup then Negotiator ResponderHandler exceptions ��� the last residual behind the resync-under-saturation flake family

What was being attempted: Closing the 20/20 N=64 gate for stressTestBootstrapServiceResyncDoesNotGiveUpUnderSaturation on UrlResolver main (post-PR 776, protocol 0.0.364 which includes both the PeerRegistry service-index fix and the new bounded peer-exchange responses from https://github.com/CodexCoder21Organization/UrlProtocol/pull/364).

What went wrong: Even with all three shipped/staged fixes (service-index O(N��) via https://github.com/CodexCoder21Organization/UrlResolver/pull/776, bounded peer-exchange responses via UrlProtocol PR 364 / protocol 0.0.364, and a dial-aware resync deadline), the N=64 gate still fails: one of 64 services never reached the relay within 75 seconds. Forensics (full log preserved at /tmp/urlresolver-post776-n64-gate/run-01.log on the /code box): multiple PeerExchange attempts spent 15-22 seconds in pre-handler connection/stream setup and then failed with Negotiator$ResponderHandler exceptions; the relay retained only the provider's join alias, proving the authoritative service record never reached onPeersReceived. The addPeer O(N��) from https://github.com/CodexCoder21Organization/UrlResolver/pull/777 was explicitly ruled out (0.0.364 contains the stronger merged service-index implementation and its large-registry regression passes).

Impact: The registration-resync flake family (the tests that evicted a green PR from the UrlResolver merge queue 12 times over 2026-07-14/15) cannot be fully de-flaked until stream negotiation stays responsive under concurrent dial load. Every saturation-family stress test inherits a residual failure rate.

Workaround used: None shipped. An experimental event-loop isolation patch (separating dial-storm work from the event loops serving established connections' stream negotiations) improved the gate to 7/8 and is preserved in git stash@{0} of /code/workspace/UrlResolver-ab ��� suggestive evidence for the mechanism but not reliable enough to ship. A clean downstream adoption branch (commit 08135478, pinned to protocol 0.0.364, post-776 main) is ready to become a PR once this defect is fixed.

Suggested durable fix: In the libp2p/Netty layer (community.kotlin.libp2p fork and/or UrlProtocol's Libp2pHostFactory), isolate outbound dial initiation and connection-scan work from the event loops that service established connections' stream negotiation, or bound negotiator work per event-loop tick, so a dial storm cannot delay ResponderHandler processing by tens of seconds. Reproduce with the N=64 amplifier (52-64 providers through one relay); acceptance = 20/20 on that gate. Track alongside https://github.com/CodexCoder21Organization/UrlResolver/issues/682.


Production verification — 2026-07-17

Status: STILL EXISTS. Live GitHub verification found 2 referenced tracker(s) still open: https://github.com/CodexCoder21Organization/UrlResolver/issues/682, https://github.com/CodexCoder21Organization/UrlResolver/pull/777. A live ProductionHealth connection also emitted repeated NothingToCompleteException gossip failures, while the stopped HardwareControlFabric daemon log ends with Netty ByteBuf leak reports.

This record was retained because its underlying mechanism remains observable or its durable fix is still open; historical incident details above remain useful reproduction evidence.