← Workstreams

Workstream: Pluggable Execution Environments

Status: Partially delivered — default BuildTest dispatch, annotation/routing, and manager-owned NetLab execution-environment routing delivered 2026-07-31; compile-once shared-cache execution and remaining trust/gating work are still planned · Component: Maximize developer productivity

Goal

A test should be able to declare where it runs — @ExecutionEnvironment("url://my-runner/run?gpu=true") — and have the build system route that test to the named runner service. The official url://buildtest/ droplet pool becomes the default runner, used only when a test names no environment of its own. Underneath, the compilation that produced the test's binaries happens once: any runner executes a test by pulling the already-compiled classpath from the shared build cache instead of rebuilding it.

The result reframes test execution the same way url:// already reframed every other capability: as a pluggable, horizontally-scalable service plane. The build system's job collapses to compile once, then schedule each test onto the right runner — and an entire class of "impossible to test here" scenarios (a GPU, a NetLab topology, a specific OS or piece of hardware) becomes a one-line annotation pointing at a runner that has it.

This is the test-execution counterpart to the shared build cache: that milestone makes a compiled rule reusable across machines; this workstream makes test execution itself a distributed, pluggable service that consumes those reusable artifacts.

Current state

The full compile-once/run-anywhere design is still future work, but the first routing layer is now delivered. BuildTest remains the default runner, with dynamic lease/memory-aware dispatch, and annotated tests can route to a named execution environment. The named NetLab route is manager-owned and releases its exact environment lease on terminal paths; the default and named paths still compile in the current BuildTest flow until the shared-cache phases land.

Piece Role today Limitation this workstream removes
ExecutionEnvironment The interface "responsible for executing a compiled build rule (or unit test)". The default implementation runs a build rule in-process or in a child JVM over a local TCP loopback. The marker annotation and URL metadata are now available through execution-environment interfaces PR 18. The full TestRunnerApi registry and compile-once remote runner contract are still planned.
kompile.Workspace.execute(...) in kompile-core One fused call: compile + resolve ("phase 1"), then run the selected tests in up to four child JVMs ("phase 2"). Tests do not go through ExecutionEnvironment — they use a separate JvmUnitTestLauncher. The test list is fixed up front, there is no compile-once / run-later split, and each execute() re-resolves dependencies. Build-once-run-anywhere is impossible without splitting this in two.
The build cache BuildCache = a SHA-256 ContentAddressedStore plus two BuildCacheIndexes. Local-file backend only. Every machine that hasn't built a rule rebuilds it — see the shared build cache milestone, which this workstream depends on.
url://buildtest/ (BuildTestServerService / BuildTestEmbedded) The default remote build/test service. Its current dispatch includes memory-aware admission and dynamic lease-based work distribution; it still uploads and recompiles the archive for the current run. The NetLab-routing release is buildtest.embedded:buildtest-embedded:0.0.61403; manager releases are netlabmanager:netlab-manager-server:0.0.76 and netlabmanager:netlab-manager-client:0.0.13. Compile-once materialization from the shared cache, fully stateless runners, and the remaining tail-rescue/re-lease graduation tests are still planned.
@ExecutionEnvironment annotation A marker plus URL metadata is defined in execution-environment interfaces PR 18, parsed during discovery by kompile-buildscript PR 69, and routed by BuildTestEmbedded PR 446. The first named route is NetLab; the full generic TestRunnerApi contract and trust model remain planned.

What makes this genuinely hard. Reading the source closely surfaces three constraints the naive "ship a classpath to a stateless runner" framing hides — they shape the roadmap below:

  • A test is not inherently stateless. Test execution binds a live LocalBuildWorkspace into the result-capturing bridge. When a test references a workspace-built artifact (@file:WithArtifact("my.module:lib:1.0") produced by a build rule), it is compiled against a stub jar whose first use triggers a synchronous buildLocalRule() → workspace.buildInternal() callback at run time to build the real artifact on demand. With no workspace bound, such a test fails outright — which is exactly why every buildtest droplet today is a full workspace. Making a runner stateless therefore requires eagerly building each test's workspace-local artifact closure during preparation and putting the real jars on the classpath, so buildLocalRule never fires at run time — trading some possibly-unused builds for true portability. (A live back-channel from runner to builder is the alternative, but it defeats "stateless" and reintroduces the coupling.)
  • The classpath is heterogeneous, not a bag of hashes. Its entries are (1) workspace-built artifacts — which genuinely belong in the shared build cache, with their POMs so transitives can be re-resolved — and (2) ordinary Maven dependencies (including the Kotlin stdlib), which are already content-addressed by Coursier's own cache and are far cheaper to re-resolve locally by coordinate on the runner than to push hundreds of transitive jars through the shared store. A prepared-test descriptor must carry a structured manifest (fetch-by-hash vs. re-resolve-by-coordinate), not a flat list of refs.
  • Today's classpath is unshippable as-is, and cross-phase state lives outside the per-test descriptor. The classpath holds a temporary compile-output directory, absolute cache paths, and a bootstrap-runner jar located via a system property — none portable. And state that flows compile→run (Maven-resolution memos, deferred effect ordering, the workspace binding above) is not captured in the per-test descriptor. The two-phase split is behavior-preserving locally, but the descriptor only becomes serializable once preparation explicitly materializes all of this.

Why it accelerates developers

  • Build once, run everywhere — not once per droplet. In a world of agent swarms and sharded CI, recompiling the same workspace on every executor is wasted work multiplied by the fleet. Compiling once and having every runner pull the classpath from the shared cache is the same compounding the shared build cache is after, extended from builds to test runs.
  • Work-stealing kills tail latency. A central queue that runners pull from self-balances: a fast or idle runner steals more, a slow test never strands a static shard, and a dead runner's unfinished tests are simply re-handed-out. No historical-duration model to maintain, and no orchestrator-side name reconstruction to get wrong.
  • Pluggable runners make the untestable testable. @ExecutionEnvironment("url://gpu-runner/..."), @ExecutionEnvironment("url://netlab-hosted/..."), or a runner pinned to a specific OS or device turns "can't run this in CI" into one annotation — while the common case stays zero-config on the default pool. This is also the sanctioned way to satisfy a test whose dependency the default environment lacks: TESTING.md forbids skipping such a test (a test needing Docker, a GPU, or a NetLab topology must fail where that capability is absent, never quietly pass), so instead of being gated out the test is routed to a runner that has the dependency and runs for real.
  • It is a toolchain-level win. Because everything is built and tested with kompile, faster and more flexible execution is felt by every other workstream and every agent.

Plan / roadmap

These milestones were proposed; the checked routing and current-dispatch items below are delivered, while the unchecked compile-once, generic-runner, trust, and full-test milestones remain future work. They are ordered so each is independently valuable and keeps CI green; the hard upstream refactor (A) is behavior-preserving before any distributed moving parts appear.

Tests are the final step of every phase, and the same standards apply throughout, per the TESTING.md guidelines: written test-first (red before green); exercising only the public API against real dependencies started in-process (never mocks or fakes) on hermetic, isolated url:// nodes (bootstrapPeers = emptyList(), OS-assigned ports, every resource closed in finally); wired across repositories via @file:WithArtifact in both directions — including real downstream consumers such as kompile-cli --remote and BuildTestCli; using real in-memory Api implementations (each in its own repo) to inject faults rather than faking symptoms or reflecting into internals; observing behaviour only through the public API — e.g. proving no recompilation by giving a runner node no source, and asserting full error text; deterministic via ManualClock; stressing in parallel with bounded concurrency; carrying measured @Timeout / @BillingQuotaExecutionLimit; and proving bounded memory under a bounded workload. One self-contained .kts per test under tests/, with shared harness code in its own Maven repo, never a tests/ helper.

A. Decouple build from execution in kompile-core

  • [ ] Split the fused Workspace.execute() into two phases. Introduce prepareTests(...) (compile + resolve + discover, emitting one prepared-test descriptor per test) and runPreparedTest(descriptor) (execute a single already-compiled test in an isolated JVM). The existing execute() becomes prepareTests().forEach { runPreparedTest(it) }. kompile already holds the pieces internally (compilation produces a per-test descriptor; execution runs them one at a time), but "behavior-preserving" is not free: the refactor must keep the cross-phase state that lives outside the descriptor today — the Maven batched/unified resolution memos, the deferred effect-emission ordering (TestDiscoveredEffect then completion-ordered TestCompletedEffect), and the live workspace binding used for stub-jar dispatch. At this stage the descriptor is an in-process handle, not yet serializable.
  • [ ] Give tests their own executor interface; share only the infrastructure around it. The existing ExecutionEnvironment is, in effect, already the build-rule executor — shaped for them: a compilationUnitFragment and arg values in, an output artifact out, the full privileged workspace bridge (visitMethodInsn, buildLocalArtifact, dynamic buildLocalRule) available throughout. A test wants a different shape entirely — a materialized classpath plus timeout/memory in, a pass/fail/duration/logs result out, and (post eager-build) no workspace back-channel. Introduce a distinct test executor interface rather than forcing tests onto the build-rule signature. What is genuinely shared stays generic below both executors: the @ExecutionEnvironment annotation, the routing/scheduling, the shared build cache + materialization, and the url:// runner-service envelope (a service advertises which kind(s) it serves). That shared layer — not a common execute() — is what makes adding build-rule routing later cheap, while a test-runner author never sees a build-rule concept. Keep the genericity at the queue/scheduler boundary (a PreparedAction the dispatcher routes and requeues), not in the executor.
  • [ ] Make the descriptor self-describing. Replace the classpath's temp dirs / absolute paths / property-located bootstrap jar with a structured form preparation can later materialize anywhere (the actual content-ref-vs-coordinate manifest is built in phase B). This is the seam that turns the in-process handle into something shippable.
  • [ ] Test the phase end-to-end — a minimum bar, not a ceiling. Write these failing-first against real dependencies; they are the floor, not the target. Before the phase is done, ultrathink whether the suite is comprehensive and evaluate every change in this phase for full coverage of its edge cases, error conditions, boundary values, concurrency, and integration paths — adding tests beyond this list as needed.
    • prepareThenRunMatchesFusedExecute — a real multi-test workspace (a passing test, a failing test, a throwing test, and a no-package bare-name test) run both via the existing fused execute(["."]) and via prepareTests + runPreparedTest; assert identical statuses, pass/fail counts, and the same emitted effect set. Red until the split exists.
    • preparedRunPreservesEffectOrdering — assert every TestDiscoveredEffect is emitted before any TestCompletedEffect, guarding the cross-phase effect-ordering hazard the source review surfaced.
    • preparedRunWithWorkspaceLocalArtifactDependency — a test that @file:WithArtifacts a workspace-built artifact passes under the split (workspace still bound in-process), proving the local refactor preserved stub-jar buildLocalRule dispatch.
    • preparedRunSharedBatchedResolutionMatchesFused — two test files sharing batched Maven coordinates produce identical results across both paths, guarding the resolution-memo hazard.

B. Build once, run anywhere

Depends on the shared build cache (a remote ContentAddressedStore over url://simple-filesystem/).

  • [ ] Eagerly build each test's workspace-local artifact closure during prepareTests. So the runtime buildLocalRule → buildInternal back-channel never fires on the runner, preparation builds the real artifacts a test depends on (not stubs) and publishes them — and their POMs — to the shared build cache. This is what lets a runner be stateless; it is the crux of this phase, not an afterthought.
  • [ ] Give the descriptor a structured classpath manifest, not a flat ref list. Distinguish (a) workspace-built artifacts → SHA-256 content refs fetched from the shared store, from (b) Maven dependencies → group:artifact:version coordinates re-resolved locally by Coursier on the runner. Pushing transitive Maven jars through the shared store is the opposite of an optimization; only workspace-built outputs (and their POMs) travel.
  • [ ] Materialize on the runner. runPreparedTest fetches the workspace-artifact refs from the shared store into a local L1 cache, re-resolves the Maven coordinates locally (parsing fetched POMs to re-resolve transitives), assembles the real classpath, and launches the JVM — never recompiling.
  • [ ] Test the phase end-to-end — a minimum bar, not a ceiling. Write these failing-first against real dependencies; they are the floor, not the target. Before the phase is done, ultrathink whether the suite is comprehensive and evaluate every change in this phase for full coverage of its edge cases, error conditions, boundary values, concurrency, and integration paths — adding tests beyond this list as needed.
    • secondNodeRunsTestWithoutRecompiling — node A prepares against a real in-process shared-cache service (the SimpleFileSystem artifact, isolated node); node B, with an empty local cache and no source, runs a descriptor by materializing from the cache and reports the correct status. The absence of source on B is the hermetic proof that no recompilation occurred. Red until the remote cache backend exists.
    • workspaceLocalArtifactDependencyRunsOnSourcelessNode — the crux: a test depending on a workspace-built artifact runs on a source-less, back-channel-less node B because preparation eager-built the closure and published the artifacts and their POMs. Red if eager-build is missing — B throws the "no workspace bound" error.
    • mavenDependencyReResolvedLocallyNotFetchedFromCache — a test depending only on ordinary Maven deps runs on B while the shared cache holds zero Maven jars, proving local Coursier re-resolution per the structured manifest (not pushing transitive jars through the store).
    • runnerFailsLoudlyWhenArtifactMissingFromCache — an in-memory cache configured to drop one hash → B fails with a descriptive full-text error, never a fabricated test result.
    • manyRunnersConcurrentlyMaterializeOverlappingClasspaths — N runners pull overlapping classpaths from one cache in parallel (bounded concurrency); assert all succeed, getOrCompute compute-once dedup holds under contention, and peak memory stays bounded.

C. The test-runner service contract and the default pool

  • [ ] Define a TestRunnerApi url:// contract (its own Api / Embedded / ServiceServer / CLI per the standard layered architecture): given a prepared-test descriptor (content-ref classpath, class, method, timeout, memory quota, pass-through parameters), materialize the classpath, run the test in an isolated JVM, and return / stream the result. This is the single interface every execution environment — official or third-party — implements.
  • [ ] Make url://buildtest/ the default TestRunnerApi implementation. Wrap the existing droplet pool behind the contract: a builder compiles once and publishes to the shared cache, and the droplets become runner nodes that pull prepared tests and their artifacts. The transport-layer precursor — getting the droplets reporting over url:// at all (the droplet-side build tool hosting a url:// service so the coordinator stops holding a persistent per-shard SSH connection, while keeping today's upload-and-build-on-the-droplet model so §B's statelessness is not yet required) — is its own nearer-term effort: BuildTest Runner Reporting. Note that workstream's 2026-07-25 revision: it retires the held connection, not SSH, and there is no separate droplet daemon — the build tool serves its own reporting. It establishes the url:// reporting transport; this bullet and the work-stealing scheduler below then build the stateless, pull-based runner on top of it.
  • [x] Replace static sharding with a work-stealing scheduler. Delivered 2026-07-31 through BuildTestEmbedded PR 441, BuildTestEmbedded PR 446, and BuildTestEmbedded PR 448. The current BuildTest dispatch surface uses shared lease-based work distribution and memory-aware admission; compile-once materialization and the full tail-rescue/re-lease end state remain separate unchecked items.
  • [x] Land the dispatch plane early, transport-agnostic — do not gate it on the url:// reporting swap. The current dispatch plane landed in BuildTestEmbedded PR 441, with named-route work in BuildTestEmbedded PR 446 and BuildTestEmbedded PR 448. The remaining TestRunnerApi/url:// reporting and compile-once contract are still planned.
  • [ ] In-run flake re-lease — the mechanism reruns ride on. Because the queue is live for the whole run, a failed test can be re-leased within the same run — preferentially to a different runner than the one it failed on — while every droplet still holds the compiled workspace: no new droplet, no recompile, seconds instead of a fresh submission round. The scheduler owns only the mechanism (re-enqueue with an attempt counter and a per-test attempt cap); the policy — which failures are believed flaky, how much rerun budget a build gets — belongs to Kompile Remote Build and arrives as parameters. Pass-on-runner-B-after-failing-on-runner-A is also a diagnostic signal in its own right: it separates unhealthy-droplet flakes from genuinely nondeterministic tests, and feeds the flake analytics accordingly.
  • [ ] Tail rescue test execution — speculative duplicate execution at the end of a run. When the queue is empty and a runner goes idle while another runner still holds in-flight work, the idle runner first adopts any leased-but-unstarted tests (the busy runner's lease is trimmed back to what it is actually executing). If nothing is adoptable — the tail is down to tests already executing on a slow or suspect runner — the scheduler duplicates the in-flight attempts onto the idle runner and races them: first terminal completion wins, the run finishes on the winner, and the loser is cancelled (or its late result recorded as superseded). This bounds the tail by the healthiest runner rather than the sickest, and rides out a wedged droplet or a droplet-specific flake without human intervention. Three rules keep it honest: (1) results are exactly-once — attempts are keyed (test, attemptId), the projection records one authoritative result per test, and superseded attempts are kept as attempt history (never silently dropped — they are flake-analytics gold), with the finalize-incomplete backstop made attempt-aware (a test with one dead attempt but another live one is not "incomplete"); (2) rescue is loud, never silent — every adoption and duplication is recorded on the run with its inputs (which runner was deemed slow/suspect and why, who won the race, by how much), surfaced in the WUI and rolled into per-droplet health analytics, because tail rescue mitigates an unhealthy runner or a straggling test and the underlying defect must stay visible for root-cause work, not be papered over; (3) speculation is safe by default and opt-out by exception — kompile tests are hermetic by the testing standards, so concurrent duplicate execution is safe for the overwhelming default, and the rare test that genuinely cannot run twice concurrently (e.g. it mutates a shared external resource, which the standards already discourage) can be marked speculation-ineligible and is then only ever re-run after its first attempt terminates, never raced.
  • [ ] Test the phase end-to-end — a minimum bar, not a ceiling. Write these failing-first against real dependencies; they are the floor, not the target. Before the phase is done, ultrathink whether the suite is comprehensive and evaluate every change in this phase for full coverage of its edge cases, error conditions, boundary values, concurrency, and integration paths — adding tests beyond this list as needed.
    • schedulerDistributesTestsAcrossRealRunners — K real runner url:// services plus the scheduler; a run-all of N tests yields N results with correct statuses and disjoint work (no test run twice, none dropped).
    • deadRunnerInFlightTestsRequeuedAndComplete — actually kill a real runner mid-run (close its node) while it holds in-flight tests; assert the scheduler requeues them and the run still completes with all N reported. A second variant kills every runner and asserts the finalize-incomplete backstop marks each unfinished test FAILED — never silently passed. Real teardown, not a flag.
    • noRunnerIdlesWhileQueueNonEmpty — with uneven test durations across K runners, assert the work-stealing invariant (no idle runner while work remains) through the scheduler's public status stream — the deterministic form of the balance claim, with no wall-clock assertion to flake on.
    • buildOnceManyRunnersNoRebuild — one builder publishes; K source-less runners execute pulled tests and the results aggregate; assert none recompiled and the counts are exact. The headline B+C integration.
    • schedulerBoundedMemoryUnderLargeRun — a large bounded N driven through the scheduler under a tight @BillingQuotaExecutionLimit heap completes without OOM, proving in-flight/queue tracking is bounded.
    • kompileCliRemoteRunStillPasses — the real kompile-cli --remote (and BuildTestCli) drive the scheduler-backed service end-to-end with correct counts, proving the existing consumer is unbroken.
    • idleRunnerAdoptsBusyRunnersUnstartedLease — with the queue empty, a runner going idle adopts another runner's leased-but-unstarted tests (observed via the dispatch log: the lease is trimmed and the tests complete on the adopter); no test waits on a busy runner while an idle one exists.
    • tailRescueRacesInFlightStragglerAndFirstCompletionWins — one test injected to hang on its original runner while the rest of the run drains; the idle runner receives a duplicate attempt, the run completes when the duplicate finishes (bounded by the healthy runner, asserted via ManualClock), exactly one authoritative result is recorded, and the superseded attempt is present in the attempt history — not dropped, not double-counted.
    • tailRescueSurvivesDeathOfOriginalRunner — the slow runner is actually killed after the rescue duplicate dispatches; the run still completes with correct counts from the rescue attempt.
    • rescueDecisionsAreExplainable — every adoption/duplication in the above scenarios is queryable on the run with the inputs that produced it (which runner, why deemed slow, race outcome) — full-text assertion of the recorded reason.
    • speculationIneligibleTestIsNeverRaced — a test marked speculation-ineligible is never concurrently duplicated: with its first attempt in flight, the idle runner does not receive a duplicate; after the first attempt terminates it may be re-leased sequentially.
    • inRunReLeaseRetriesFailedTestOnDifferentRunner — a fail-once-then-pass test injected on runner A is re-leased within the same run, lands on runner B (different-runner preference asserted), passes, and the run records both attempts with one authoritative passing result — no new build submission, no recompilation (runner B has no source).

D. @ExecutionEnvironment routing

  • [x] Define the single @ExecutionEnvironment("url://...") annotation. Delivered 2026-07-29 in execution-environment interfaces PR 18; it carries the named URL and query parameters.
  • [x] Parse it during discovery. Delivered 2026-07-29 in kompile-buildscript PR 69; discovery carries the annotation metadata into routing.
  • [x] Route in the scheduler. Delivered 2026-07-30/31 through BuildTestEmbedded PR 446 and BuildTestEmbedded PR 448; named execution-environment failures do not silently fall back to the default runner.
  • [ ] Test the phase end-to-end — a minimum bar, not a ceiling. Write these failing-first against real dependencies; they are the floor, not the target. Before the phase is done, ultrathink whether the suite is comprehensive and evaluate every change in this phase for full coverage of its edge cases, error conditions, boundary values, concurrency, and integration paths — adding tests beyond this list as needed.
    • annotatedTestRoutedToNamedRunner — a @ExecutionEnvironment("url://special/...") test runs on a real "special" runner while unannotated tests use the default pool, observed via runner-reported identity.
    • executionEnvironmentQueryParamsDeliveredToRunner — ?mode=x&n=3 arrives at the runner exactly (full-value assertion of the parameter map).
    • functionLevelEnvironmentOverridesFileLevel — a file-level default plus a function-level override route correctly, verifying PSI parsing through observable routing rather than parser internals.
    • unreachableExecutionEnvironmentFailsLoudly — an annotated URL that declines or does not exist yields a descriptive full-text infrastructure failure — never a silent pass, never a silent fallback to the default pool.
    • malformedExecutionEnvironmentUrlReportsDescriptiveError — a bad URL produces the exact documented error text.

E. NetLab: the first first-class execution environment

NetLab is the canonical "impossible to test otherwise" runner — a test runs inside a virtualized network topology — and it is now the first named @ExecutionEnvironment. BuildTestEmbedded PR 446 added named-runner routing, and PR 448 routes url://netlab-hosted/... through one manager-owned environment per dispatch and releases the exact lease on completion, returned test failure, interruption, runner failure, or service close. The manager-side ownership and reclamation seam landed in NetLabManagerServer PR 127. The full generic TestRunnerApi contract and compile-once runner model remain planned; this delivered route closes the previously documented client-cleanup gap for the BuildTest path.

  • [x] Make url://netlab-hosted/ a named execution-environment route. Delivered 2026-07-31 through BuildTestEmbedded PR 446 and BuildTestEmbedded PR 448. It is routable through the current BuildTest dispatch seam; the full TestRunnerApi wrapper remains planned.
  • [x] Give the BuildTest dispatch path ownership of the environment lifecycle. Delivered 2026-07-31 in BuildTestEmbedded PR 448, with the manager ownership/reclamation API in NetLabManagerServer PR 127. Completion, failure, interruption, runner failure, and service close release the exact lease.
  • [ ] Keep the absolute-lifetime cap as a backstop, not the fix. The activity-independent per-environment cap added to NetLabManagerServer (ENV_MAX_LIFETIME_MINUTES) stays as defense-in-depth so a leaked droplet can never reach the 1-hour DigitalOcean backstop, but it is explicitly not the systemic answer — ownership-based reclamation is.
  • [ ] Test the phase end-to-end — a minimum bar, not a ceiling. Write these failing-first against real dependencies; they are the floor, not the target. Before the phase is done, ultrathink whether the suite is comprehensive and evaluate every change in this phase for full coverage of its edge cases, error conditions, boundary values, concurrency, and integration paths — adding tests beyond this list as needed.
    • netlabEnvironmentTornDownWhenOwningTestCompletes — a routed NetLab test completes and its topology/droplet is reclaimed by the scheduler with no client-issued delete, observed via the manager's environment list going empty.
    • killedNetlabRunnerEnvironmentReclaimedNotLeaked — actually kill the owning test/runner mid-run; assert the scheduler reclaims the environment promptly (not after the idle-reaper timeout), proving detection rather than a timeout. Real teardown, not a flag.
    • netlabTestRunsInsideTopologyForReal — a @ExecutionEnvironment("url://netlab-hosted/...") test observably executes inside the virtualized topology (e.g. a NAT-crossing resolve that only succeeds under the applied topology), proving the runner is genuine and not a silent default-pool fallback.

F. Cross-cutting

  • [ ] Trust and gating. Decide which runners may gate a merge. An arbitrary @ExecutionEnvironment URL is fine for development, but a runner that simply returns PASSED cannot be allowed to gate CI — likely an allow-list of trusted runner identities (W3Wallet-attested) for gating, with arbitrary runners permitted for non-gating/experimental runs.
  • [ ] Capabilities, not ambient access. Reuse the W3Wallet capabilities raised in the Kotlin Build workstream: a scoped capability authorizes writing the shared cache, a capability authorizes a runner to read it, and capabilities are provisioned into a test so it reaches credential-bearing services without secrets in source — exercised on whichever runner the test lands on.
  • [ ] Artifact portability. Document and enforce the compatibility envelope for "build here, run there" — pure-JVM bytecode is portable; native dependencies, OS, and JVM-version assumptions are not. The prepared-test descriptor should carry enough metadata for a runner to reject an incompatible artifact rather than fail mysteriously.
  • [ ] Documentation. Keep the BuildTest and kompile-core READMEs and the Kotlin Build project page in sync as these land.
  • [ ] Test the phase end-to-end — a minimum bar, not a ceiling. Write these failing-first against real dependencies; they are the floor, not the target. Before the phase is done, ultrathink whether the suite is comprehensive and evaluate every change in this phase for full coverage of its edge cases, error conditions, boundary values, concurrency, and integration paths — adding tests beyond this list as needed.
    • untrustedRunnerResultDoesNotGateMerge — a result from a non-W3Wallet-attested runner is observably marked non-gating, while an attested runner's result gates; exercised against a real isolated W3Wallet.
    • testReceivesScopedCapabilityOnAnyRunner — a test reaches a real isolated credential-bearing service via an injected scoped capability, on whatever runner it lands, with no secret in source or process.
    • incompatibleArtifactRejectedWithDescriptiveError — a runner advertising a different platform rejects an incompatible artifact with a full-text error rather than a mysterious downstream failure.

Open questions

  • Result authenticity & gating. How does the scheduler trust a result from a runner it does not control? Is gating restricted to W3Wallet-attested runners, and what does a non-gating "ran somewhere untrusted" result look like in the BuildTest WUI?
  • Eager-build cost vs. lazy stubs. Building a test's full workspace-local artifact closure up front (to remove the runtime buildLocalRule back-channel) can build artifacts the test never exercises — today's stub-jar mechanism builds them lazily on first use. Is the portability worth the extra build work, and can the closure be pruned to only the artifacts a test statically references?
  • Parameterized workspace rules. The stub-jar bridge only dispatches parameterless build rules; a test depending on a parameterized workspace artifact already fails today. Does the eager-build path need to cover these, or is the limitation inherited and simply documented?
  • Dispatch granularity. One test per dispatch maximizes balance but pays a child-JVM start (~2–3s) each time; small per-runner batches amortize it. What buffer keeps a runner's JVM slots full without re-introducing static imbalance?
  • One annotation; kind-specific executor interfaces. @ExecutionEnvironment is a single marker — it says only "run this unit in the environment at this URL." Whether the annotated unit is a test or a build rule is decided by what it sits on, and that decides which interface kompile uses to talk to the environment. Those interfaces genuinely differ: a test takes a materialized classpath plus timeout/memory and returns pass/fail/duration/logs as a stateless leaf; a build rule takes a compilationUnitFragment, keeps a live workspace back-channel as it runs (visitMethodInsn bytecode interception, buildLocalArtifact, dynamic buildLocalRule), and returns an output artifact — closer to an environment-tagged generalization of today's workspace-bearing RemoteBuildWorkspace (serve the workspace, run there, return the artifact by hash). So the two share the annotation, the routing, the cache, and the url:// transport, but get separate executor interfaces (a runner advertises which kinds it implements), because it is common to implement a test environment yet never a build one. Test execution comes first; the build-rule backend is deferred. The discipline required now is only to keep test out of the shared routing/queue/cache — not to collapse both onto one fat execute(). A concrete instance of the URL-marker principle: today an ExecutionEnvironment such as LaunchDocker takes a docker image as a typed parameter; under this design that becomes a URL to a docker-capable execution service with the image carried as a query parameter — @ExecutionEnvironment("url://docker-runner/run?image=acme/builder:1.4") — so the annotation stays a pure URL marker and all environment configuration rides in the URL rather than in typed parameters.
  • Relationship to LambdaServer. LambdaServer already runs registered code on demand and its API implements kompile.Workspace. Is a LambdaServer function the natural substrate for a TestRunnerApi runner, rather than a bespoke service per environment?
  • Local mirror & eviction. Each runner caches fetched artifacts locally; how is that mirror bounded, and how does eviction interact with the shared cache's "a named result's files still exist" invariant (the same question the shared build cache raises)?

Related

  • Kotlin Build — the hard dependency: its shared build cache is what makes build-once-run-anywhere possible, and its url://buildtest/ remote service is reframed here as the default runner rather than the only one.
  • LambdaServer — on-demand remote execution whose API already implements kompile.Workspace; a leading candidate substrate for runner services.
  • NetLab — a NetLab-backed runner is the canonical "impossible to test otherwise" environment: @ExecutionEnvironment("url://netlab-hosted/...") runs a test inside a virtualized network topology. It is this workstream's first first-class runner (milestone E above), and scheduler-owned environment lifecycle is what lets an abandoned per-test topology be detected and reclaimed rather than left to a timeout.
  • W3Wallet — the capabilities that authorize writing/reading the shared cache, attest a trusted gating runner, and are provisioned into tests so they reach secured services without secrets in source.
  • Manager–Daemon Pairing — the registered-worker-pool pattern and enrollment handshake by which a user-contributed runner (a GPU box, a specialized-hardware host, a whitelisted machine) would register itself with the scheduler when the generic TestRunnerApi contract lands.
  • SimpleFileSystem — the remote filesystem the shared build cache (and therefore artifact transfer between builder and runners) is built on.

Graduation

The routing and current-dispatch portions are delivered: the default url://buildtest/ path has dynamic/memory-aware dispatch, and named url://netlab-hosted/... tests use manager-owned environments with terminal-path reclamation. This remains an active workstream until a second node demonstrably runs a test against artifacts a first node compiled with no recompilation; the default pool demonstrably re-leases failed tests within a live run and tail-rescues stragglers; and the generic TestRunnerApi, trust/gating, and end-to-end coverage are complete. Graduating folds these capabilities into the existing Kotlin Build and BuildTest documentation rather than creating a new project page.