← Workstreams

Workstream: NetLab

Status: Planned · Component: Maximize developer productivity

Goal

Make the hard-to-test network scenarios routine to test. A developer should be able to stand up a realistic, isolated network — nodes behind NATs, multiple segments, a slow or flaky link, a partition — and run their application inside it as an ordinary, reproducible test. The goal is to remove "we can't really test that without real infrastructure" as an excuse, especially for the P2P fabric the whole platform stands on.

The motivating cases:

  • P2P gossip across NATs. UrlResolver must work when peers are behind NAT gateways and have to discover each other and hole-punch. That is impossible to exercise on a single flat LAN.
  • Adverse links. Slow connections, packet loss, flaky/intermittent links, and partitions — the conditions where distributed systems actually break.
  • Topology-dependent behavior. Multiple network segments, custom routing, STUN/TURN paths, and isolation guarantees.

Current state

This exists as the NetLab project — a network-topology lab that builds isolated virtual networks with Docker and exposes the same TopologyService API over url:// regardless of backend:

Layer Repository What it provides
Api NetLabApi TopologyService interface and topology types (Topology, SwitchConfig, GatewayConfig, HostConfig, …)
Parser NetlabTopologyParser Parses JSON topology files into typed objects
Embedded NetLabEmbedded TopologyManager — creates Docker networks (switches), NAT gateway containers (iptables), and host containers; idempotent apply
Worker NetLabWorkerServer Wraps Embedded behind url://netlab-worker-{uuid}/
Manager NetLabManagerServer Provisions on-demand DigitalOcean droplets, exposes url://netlab-hosted/
CLI NetLabCLI Lifecycle, apply, logs, health, demo; --json

Topology model: switches are Docker bridge networks that are always internal — a topology is isolated from the real internet by default (the internal flag defaults to true, "internal": false is rejected, and a host with no networks gets --network none); the only opt-in to real internet access is an internet_gateways node bridging one switch to a non-internal uplink. gateways perform NAT between two topology switches via the netlab-gateway image, and hosts are containers attached to one or more switches with custom routing. It runs in local mode (your Docker daemon) or managed multi-node mode (each environment on its own droplet), with the identical API either way. So the machinery exists; this workstream is about making it the default, low-friction way to write network tests. In managed mode each per-test droplet's lifecycle is today the client's responsibility (acquire → delete), with a time-based reaper as the only backstop — so a client that dies without calling delete() leaks a droplet until a timeout notices; making NetLab a first-class runner in the Execution Environments workstream moves that lifecycle to a scheduler that owns it (create-on-dispatch, reclaim-on-death), so an abandoned environment is detected and reclaimed rather than left to a timeout. The activity-independent ENV_MAX_LIFETIME_MINUTES cap in NetLabManagerServer is interim defense-in-depth, not that systemic fix.

Driving apps inside a topology. Hosts run arbitrary images (any registry image, or an OCI tarball via ociTarball), and the worker already exposes the surface a test needs: uploadFile stages a file (a built JAR, an extension bundle, …) into every host at /jars/…, exec runs a command in a host and returns its exit code (so the exit code is the assertion), and logs pulls back combined stdout/stderr. In-topology url:// discovery needs no public DHT: a node inside the topology serves as the bootstrap/root and its address is injected directly into the other hosts as environment variables — exactly the pattern UrlResolver's own NetlabHarness and urlresolver-test.json use (BOOTSTRAP_IP / BOOTSTRAP_PORT / BOOTSTRAP_SEED, with the peer-id derived offline via peerIdFromSeed), because the test sets up the topology and therefore knows every address up front. UrlResolver additionally has tested circuit-relay fallback for peers that can't be reached directly behind a NAT.

Why it accelerates developers

  • It removes a whole category of "untestable." The bugs that only appear behind a NAT or on a lossy link are exactly the ones that reach production today; NetLab brings them into the test suite.
  • It protects the foundation. UrlResolver is load-bearing for every url:// service; reliable NAT/gossip tests guard the layer everything else depends on.
  • Same API, local or hosted — fast local iteration on Docker, scaled-out isolation on droplets, with no test rewrite.

Pluggable node types

The topology model above assumes a single kind of node — a Docker container — but that is an implementation choice, not a fundamental one. Abstractly a node needs only a set of network interfaces attached to switches and something running behind them; separating how a node is executed from how its packets move turns the node into a pluggable backend and lets one topology mix and match kinds of node. Five are in view:

Node type Weight Datapath Notes
Docker container medium real-kernel today's only type
VM heavy real-kernel full guest; differs from Docker mainly in lifecycle
Firecracker microVM medium real-kernel needs /dev/kvm, so managed mode needs nested-virt-capable hosts
Manual node feather in-process test code drives the node directly — send/receive packets, open a TCP connection for streams; fully in-process, no Docker
Java-rewrite node feather in-process an unmodified user JAR loaded in-process whose java.net/NIO calls are rewritten to route through the fabric — the code believes it is on a real VM

Two datapath regimes. What makes mix-and-match coherent is splitting nodes by where packets actually flow:

  • Real-kernel datapath — Docker, VMs, firecracker. All are "real" nodes with a tap/veth on a real Linux bridge; packets traverse the kernel stack and impairment is real tc/netem. They differ almost entirely in lifecycle/provisioning (docker run vs. booting a microVM with a kernel + rootfs), so VM and firecracker are largely a new lifecycle provider on the existing switch/NAT/impairment machinery, not a new fabric.
  • In-process datapath — manual and java-rewrite. "Packets" are objects in the JVM and the switch/router/NAT/firewall/impairment is a pure-Kotlin software-defined network: no Docker, no kernel, deterministic and microsecond-fast. This is the high-leverage new capability, and both lightweight node types are front-ends onto it.

Another forcing function — Docker's shared address pool. The real-kernel regime carries a hidden hazard the in-process regime does not: Docker's IPAM is a per-daemon allocator that refuses to create two networks with overlapping subnets — even when they are fully isolated internal bridges that could never exchange a packet. It is an allocation-bookkeeping rule, not a data-plane check, so isolation working perfectly does not save you. Because the kotlin.build runner shards many tests onto one droplet and runs them concurrently against that droplet's single Docker daemon, two tests that happen to hard-code the same /24 race on docker network create, and the loser fails at allocation time with invalid pool request: Pool overlaps with other one on this address space (exit 125) — a non-deterministic flake that depends only on shard layout, not on anything the test does. This bit NetLabEmbedded in 2026 (the impairment and isolated-networks integration tests both claimed 172.39.0.0/24); the immediate fix gave every Docker test a unique subnet plus a fake-daemon regression guard (NetLabEmbedded #52), but the root cause is structural — a single shared daemon with one global address space is a hermeticity hazard (see the testing standards). The in-process datapath removes the whole class by construction: each topology's fabric is its own address space (the same per-node-classloader isolation that lets N peers reuse IPs in one JVM), so identical subnets across concurrent topologies simply cannot collide and there is no shared daemon to contend on. A real-kernel topology inherits the same immunity only to the extent addressing moves into NetLab's own switch/router (containers on --network none with a TAP into the fabric) rather than being delegated to Docker's IPAM — one more reason the userspace fabric, not the native mechanism, should be the source of truth.

Wire-faithful, by necessity. The in-process datapath carries real IP/TCP/UDP packets, not a reified "reliable stream" — for two reasons that turn out to be one:

  • IP-level firewall rules. A rule like drop tcp/443 from subnet X has nothing to match unless a real 5-tuple is on the wire; packet-level filtering requires packet-level framing.
  • Cross-regime mixing (a manual node sharing a switch with a Docker container) bridges the in-process fabric to the kernel through a TAP device, where the peer is a real Linux TCP stack that will only complete a handshake / ack / retransmit with something speaking real TCP. Crossing the boundary therefore forces real endpoint behavior, not just real framing — and sets the floor on TCP fidelity, since interoperating with Linux is what stress-tests any simplification.

There is a subtler forcing function: once a firewall can drop a packet mid-stream, "what happens to the byte stream" only has a correct answer — retransmit, slow down, survive — if something below the socket performs real segmentation and retransmission. That is a TCP stack; no shortcut keeps firewall/impairment semantics correct without one.

The spine: one clock-driven userspace TCP/IP stack. Everything in-process attaches to a single IP+UDP+TCP stack:

  • Manual node = the stack's API exposed directly to test code (raw IP at the bottom; open-a-TCP-connection-for-streams at the top).
  • Java-rewrite node = the same stack's API exposed to a loaded JAR via shadow classes (below).
  • Cross-regime = the stack's raw-IP bottom bridged to the kernel over TAP.

So manual and java-rewrite are not two features but one stack with two front-ends, with the whole datapath — switching, routing, NAT, firewall, and impairment (delay, loss, drop, reorder) — living in the fabric beneath every node type. Wherever feasible that datapath logic should be NetLab's own implementation applied uniformly across both regimes rather than delegated to native kernel mechanisms (Linux bridges, iptables/nftables, tc/netem) — on the real-kernel datapath, by routing traffic through the same userspace switch/router over the TAP bridge where the throughput cost is acceptable. A single implementation means a topology behaves identically whether it runs in-process or in-kernel — the switch, the router, the firewall rule, and the dropped packet all behave the same way regardless of backend — and shedding the Linux-only dependencies (bridge, iptables, tc) lets local mode run on Windows and macOS; the native mechanisms stay available as an optional in-kernel acceleration, not the source of truth. Driving the stack's timers (RTO, delayed-ACK, …) off NetLab's injected Clock rather than wall-clock makes impaired-link tests deterministic — the answer to the determinism open question below that a real kernel stack can never give. Selecting or building this stack is the critical-path, highest-risk component; a purpose-built clock-driven stack is preferable to embedding a C/Go stack (lwIP, gVisor netstack) whose per-node state is hard to isolate in one JVM and whose timers assume real time. Treat fidelity as a dial: framing is always real, while TCP endpoint behavior can start minimal (connection setup/teardown, in-order delivery, retransmit, a simple window — no SACK/ECN/window-scaling) and grow, with cross-regime setting the minimum bar.

Java-rewrite interception: shadow classes, not SPI. Constrain java-rewrite nodes to JVM apps using java.net/NIO (explicitly excluding native-transport Netty / JNI sockets), and intercept by bytecode type-remapping to shadow classes inside a per-node classloader: enumerate the finite public network surface (java.net.{Socket, ServerSocket, DatagramSocket, InetAddress, …}, java.nio.channels.{SocketChannel, ServerSocketChannel, DatagramChannel, Selector, SelectorProvider}), provide drop-in shadow types backed by the stack, and have an ASM Remapper rewrite every reference in the node's classloader — app and bundled libraries such as Netty — from the real type to the shadow type. This is deliberately not the JDK SelectorProvider/SocketImplFactory SPI route, which is a partial seam that fails by silently leaking traffic to the real network: a blocking java.net.Socket bypasses a custom provider (it delegates to sun.nio.ch.NioSocketImpl/sun.nio.ch.Net, not openSocketChannel()), DNS resolution escapes, the legacy datagram factory is deprecated/removed, a custom Selector is fragile, and the provider is a process-wide singleton that needs classloader isolation anyway. For a network simulator a silent escape is the worst possible failure mode. Remapping makes the seam total and explicit, and per-node classloaders give the isolation that lets N peers with N IPs share one JVM. Hard rule: never fall through to the real network — any un-shadowed network API throws loudly (e.g. UnsupportedOperationException("netlab: unsupported network API …")), turning "a path was missed and traffic leaked" (silent, found in production) into "this app uses an API we have not shadowed yet" (loud, found at test time).

Where each regime runs. TAP and every kernel-datapath node need CAP_NET_ADMIN and a real kernel, so cross-regime and Docker/VM/firecracker topologies run only on a privileged host (managed droplet or privileged local). Pure in-process topologies — manual + java-rewrite only — run anywhere, including the restricted test sandbox, fast and deterministic. This maps directly onto the "where do NetLab tests run in CI" open question: pure in-process is the default, runs-in-the-sandbox path, and only kernel/cross-regime topologies need a privileged host.

Sequencing — de-risk the spine first. Because the userspace stack is the whole ballgame, build it as a vertical slice before anything else: IP + UDP + minimal-TCP, the manual-node front-end, and a software switch with drop/latency rules, proven by a hand-driven two-node firewall-and-loss test. Only once the stack is right does java-rewrite become "just" the shadow-class front-end and cross-regime "just" the TAP adapter; if the stack is wrong, that surfaces in a few hundred lines of manual-node test rather than after the rewriter and TAP bridge are built.

Plan / roadmap

These are proposed and need sharpening before work starts.

  • [ ] Pluggable node types. Make the node a pluggable backend so one topology can mix and match Docker with VMs, firecracker, and two lightweight in-process types — manual nodes (test code drives packets and TCP streams directly) and java-rewrite nodes (an unmodified java.net/NIO JAR loaded in-process, its socket calls bytecode-remapped through the fabric). The full design — two datapath regimes, a wire-faithful clock-driven userspace stack as the shared spine, shadow-class interception, and the cross-regime TAP bridge — is written up under Pluggable node types; de-risk it by building the in-process stack + manual-node slice first. The manual-node API has a flagship consumer waiting: the Pluggable Network Layer workstream plans a NetLabFabricNetworkProvider that runs UrlProtocol's real gossip/peer-management code over this stack, giving wire-faithful thousand-node gossip tests (its tier-1 in-memory provider does not depend on this stack and lands independently).
  • [ ] A test harness, not just a CLI. A small library so a tests/*.kts or JUnit test can declare a topology, apply it, run the app inside it, and assert — without shelling out. (Mind the sandbox constraints in the testing standards; tests needing Docker run on a remote droplet, never host Docker.)
  • [ ] Impairment primitives. The imperative building blocks already exist on TopologyService — setLatency, setPacketLossRate, setRateLimit, and clearImpairment apply real tc/netem qdiscs to a host interface at runtime. What remains is making "slow and flaky" a declarative topology setting (applied automatically on apply, not a manual post-apply call) and adding first-class partition controls, so adverse links are part of the topology definition, not a manual hack. Wherever feasible the impairment logic should be NetLab's own implementation applied to both datapath regimes (see Pluggable node types) rather than native tc/netem — for behavior that is identical across in-process and in-kernel runs, and to drop the Linux-only tc dependency so local mode is portable to Windows/macOS.
  • [ ] Canonical UrlResolver scenarios. Ship reusable topologies for the cases that matter — two peers behind separate NATs, relay-required paths, partition-and-heal — and wire at least one into UrlResolver's own CI. (These build on the bootstrap-by-direct-address-injection pattern above, so no scenario depends on reaching the public mesh.)
  • [ ] Application-level e2e behind a NAT. Go beyond raw url:// reachability to whole-application flows. The motivating case is W3Wallet: one host running a browser + the W3Wallet extension + the W3Wallet daemon in a single container (so the extension reaches the daemon over localhost) placed behind a NAT, talking to a public host that runs a web app and the W3Wallet permissions server (PMS) — which doubles as the in-topology url:// bootstrap node, its address injected into the NAT'd host per the pattern above. Playwright drives the browser via exec(), and the run's exit code is the assertion. Because the client behind the NAT initiates the outbound connection to the public PMS, NAT traversal is the easy (outbound) direction; the reverse path leans on UrlResolver's relay fallback. The same topology, with the public host's image/command swapped, yields two variants — the JavaScript demo (a static page driving window.w3wallet client-side, so it has no server-side capability check) and the Java-backend demo (a JVM server that reaches back into the NAT'd daemon over url:// to verify capabilities — the variant that actually exercises cross-NAT capability enforcement). Sequencing matters: the PMS is itself still aspirational, so the first realizable build target is the daemon-enforced path — the public host runs only the web app, and the daemon on the NAT'd node is the authority home reached over url://; the PMS-on-the-public-host version follows once the PMS and token-capability work land. Ship it as a reusable topology plus a maintained browser-host image (Chromium + extension + Playwright + a virtual display, in its own repo and registered like netlab-gateway) — the application-level companion to netlab-gateway. A working first cut already exists — W3WalletTests/browser-nat-e2e places Chromium + the extension + the daemon in one NAT'd container and drives it with Playwright against a public web host, in both the JavaScript-demo and Java-backend-demo variants. (The older W3WalletTestHarness netlab.ts buildTopology() is a separate, curl-only, JVM-only NAT helper — not this browser suite.) It is the daemon-enforced path as expected, and most of the hardening that makes it trustworthy has since landed: it is now fully hermetic — both switches are internal: true and it runs its own in-topology libp2p relay (NetlabHarness ROLE=serve with a fixed KEY_SEED, peer-id derived offline and injected as BOOTSTRAP_IP/PORT/SEED), so there is no internet gateway and no path to the public mesh (W3WalletTests #21); the pre-baked browser-host image (ghcr.io/codexcoder21organization/w3wallet-browser-host — JRE + Playwright + Xvfb + the extension dist baked in) is wired in as an opt-in via W3W_BROWSER_HOST_IMAGE (the default still stages a JRE + Playwright node_modules into the stock playwright:jammy image at runtime); and Playwright failure traces are now pulled back out on failure via downloadFile (see the artifact-retrieval item below — W3WalletTests #22). Both variants run in CI against url://netlab-hosted/ on netlab-cli 0.0.26 (which now exposes exec + download-file). What remains is refinement, not the isolation gap: the browser-host's container command still is the test (printing E2E_RESULT=PASS|FAIL, polled via logs) rather than the exec()-driven exit-code form — now unblocked by netlab-cli's exec, but not yet adopted; assertions are still shallow (coin count only); and the JS variant's wallet path is entirely localhost — so the Java-backend variant carries the real cross-NAT coverage.
  • [x] Host→client artifact retrieval — DONE. TopologyService could stage files into a topology (uploadFile) but not pull them out, so browser e2e couldn't extract failure artifacts (Playwright traces, screenshots, the HTML report). Shipped downloadFile(topologyName, hostName, sourcePath): InputStream — symmetric to uploadFile, reading a host file via docker exec <c> base64 <path> (executor-agnostic, so it works over the managed-mode SSH path too): netlab-api 0.0.11 (interface), netlab 0.0.19 (TopologyManager impl), exposed over url:// by the worker (service-server 0.0.27) and forwarded by the hosted manager (manager-server 0.0.35), plus a netlab-cli download-file <env> <host> <remote-src> <local-dest> command (netlab-cli 0.0.22). Now wired into browser-nat-e2e (W3WalletTests #22): each runner captures a Playwright trace per browser session and the driver pulls the tarball out of the browser-host on failure (best-effort, before teardown), which CI uploads as a workflow artifact.
  • [ ] Faster managed startup. Managed mode takes ~3–5 minutes to provision a droplet; investigate pre-baked snapshots / pooling so network tests are cheap enough to run often.
  • [ ] Teardown guarantees. Ensure environments (and droplets) are always reclaimed even when a test fails, to avoid leaking Docker resources or droplet quota.

Open questions

  • Where do network tests run in CI? Host Docker is off-limits in the sandbox; managed droplets cost minutes. What is the default execution path for a NetLab-backed test, and how do we keep it fast? The pluggable node types direction answers this for a large class of tests — pure in-process topologies (manual + java-rewrite nodes) need neither Docker nor kernel privileges, so they run directly in the sandbox, leaving only kernel/cross-regime topologies on a privileged host. For NetLab-backed tests specifically, the committed answer is Pluggable Execution Environments phase E: a @ExecutionEnvironment("url://netlab-hosted/...") test is routed to a NetLab runner whose environment lifecycle is owner-reclaimed — a NetLab droplet never outlives its owning test.
  • Determinism. How reproducible are impaired-link tests (loss/latency are stochastic)? Do we need seeded/deterministic netem, or statistical assertions with budgets? The in-process datapath resolves this cleanly for topologies built from manual/java-rewrite nodes — a clock-driven userspace stack makes loss, latency, and retransmission fully deterministic — so the question remains open only for the real-kernel (tc/netem) regime.
  • Scope vs. the agent story. Should agent swarms be able to spin up NetLab environments as part of autonomous testing? If so this also touches component 2.
  • Browser-host image upkeep. A Chromium-+-extension-+-Playwright host image is heavier to maintain than netlab-gateway (browser/driver version drift, headless/virtual display). Who owns it, and should it be baked into the managed snapshot so application-level e2e startup stays cheap?

Graduation

NetLab already has a Documentation Repository project page. This workstream graduates when a reusable test harness with impairment primitives is in use and at least one foundational service (UrlResolver) gates merges on a NetLab-backed test; at that point it is no longer tracked as a workstream here.