Repository · workstreams
Workstream: Kompile Remote Build
Status: In progress — core orchestration, opt-in routing, reruns, optimizer surfaces, dashboards, and named execution-environment routing delivered 2026-07-31; exhaustive validation, server-side build-watchman integration, WUI follow-ups, and cross-cutting coverage delivered 2026-08-05; production CI cutover remains pending an explicit user decision, with a few related PRs riding out a shared-transport infrastructure incident · Component: Maximize developer productivity
Goal
Component 1 names the enemy directly: "every flaky build … is a tax paid on every change forever." This workstream is about making the organization's CI — the kotlin.build (remote) check that gates every pull request — fast, trustworthy, and self-healing, so a single flaky test never forces a human (or an agent) to rebuild and re-run an entire suite just to get a green check.
It covers the KompileRemoteBuild family — the KompileRemoteBuildServiceServer that hosts url://kompile-remote-build/, its WUI dashboard, and the Api / Embedded / CLI around them — together with a re-architecture of the kotlin-build-ci GitHub App to sit on top of that service rather than driving the build backend directly. (The GitHub App kotlin-build-ci and the analytics service family, formerly kompile-build-ci / KompileBuildCi*, used to look confusingly similar — closely enough that the pairing once tripped up an AI agent mid-conversation. The analytics family has since been renamed org-wide to KompileRemoteBuild*, with its protocol identifier renamed to url://kompile-remote-build/ in lockstep, specifically to make the two unmistakably distinct at a glance. The GitHub App keeps its original kotlin-build-ci name unchanged — it is a separate repo and is not part of this rename.)
Delivered state
The core implementation landed on 2026-07-31; exhaustive validation, server-side watching, WUI follow-ups, public-URL coverage, and reliability hardening landed on 2026-08-05. The production GitHub App still uses the direct url://buildtest/ path until an explicit user decision authorizes the production cutover.
- CI routing seam. kotlin-build-ci PR 201 merged on 2026-07-31.
CI_BUILD_ROUTE=kompile-remote-buildselects the hosted-service adapter; an unset value ordirect-buildtestkeeps the directurl://buildtest/path; invalid values fail loudly. This PR did not activate the route in production. - Hosted orchestration service. KompileRemoteBuildServiceServer PR 17 merged on 2026-07-31. It hosts the durable planner-gated FIFO engine behind
url://kompile-remote-build/and exposes CI lifecycle, result, rerun, queue, optimizer, and resource-stat surfaces. The published server artifact iskompile.remotebuild.server:kompile-remote-build-server:0.1.17. - Execution environments.
url://buildtest/remains the default execution environment. Namedurl://netlab-hosted/...routing and manager-owned lease reclamation landed through BuildTestEmbedded PR 446 and BuildTestEmbedded PR 448, with the NetLab ownership seam in NetLabManagerServer PR 127. The NetLab-routing release isbuildtest.embedded:buildtest-embedded:0.0.61403; manager releases arenetlabmanager:netlab-manager-server:0.0.76andnetlabmanager:netlab-manager-client:0.0.13. - Published supporting surfaces. The live artifact set includes
kompile.remotebuild.api:kompile-remote-build-api:0.0.7and0.0.8,kompile.remotebuild.embedded:kompile-remote-build-embedded:0.0.13,0.0.14, and0.0.19,community.kotlin.testbundlebuilder:community-kotlin-test-bundle-builder:0.0.11, and thebuildwatchman:build-watchman:0.0.14CLI. The service currently pins the0.0.8API and0.0.19Embedded releases; the0.0.13/0.0.7releases are the earlier direct-contract releases. - Related resolver work. UrlResolver PR 904 merged on 2026-08-01, making relay selection deterministic under contention while reserving lower-tier fallback budget; it supports the routing infrastructure but is not itself a KompileRemoteBuild surface.
- Exhaustive optimizer validation. KompileRemoteBuildEmbedded PR 29 merged on 2026-08-05 with all 20 named scenarios. Triage classified 10 test-side issues and found and fixed one engine defect: explicit prediction outcomes were durable in individual decisions but missing from the aggregate prediction-accuracy query. Scenarios that cannot observe the live pull-dispatch plane assert plan-level properties and record that gap.
- Server-side build watching. kotlin-build-ci PR 211 merged on 2026-08-05. It adds 60-second lost-dispatch detection with bounded self-rerequest, subsuming
MissingCheckRunReconciler, stall/trajectory detection with bounded redrive, merge-queue ref watching, and durable remediation records. - Reliability and hardening. KompileRemoteBuildServiceServer PR 26 merged the
0.1.17server hosting Embedded0.0.19; KompileRemoteBuildEmbedded PR 24 merged atomic journal compaction and projection-failure replay; and kotlin-build-ci PR 206 made saturated runners defer work durably and admit it in FIFO order instead of producing false-FAILED checks, with eight concurrency defects fixed from fail-first evidence. - Still open around the shared-transport incident. ContainerNursery PR 567 is
OPEN: its repository checks are green, but the hosted check has been repeatedly evicted by the incident; it adds the opt-in application-readiness gate. UrlProtocol PR 490 and PR 491 areOPEN: their repository checks are green while the hosted checks currently fail, and they carry the transport fixes for preserving completed responses and isolating failure observation per thread. BuildTestEmbedded PR 504 isOPENand its strengthened telemetry proof is blocked on the same shared-transport incident. The incident is recorded in PlanRepository PR 5114, alsoOPEN.
Four things change:
- kotlin-build-ci can drive builds through
url://kompile-remote-build/when opted in (via UrlResolver); the directurl://buildtest/path remains the default until the production cutover decision is made. The hosted path makes the CI service the front door for CI builds: orchestrating them, recording their results first-hand, and serving the analytics, while buildtest stays a shared backend other callers (kompile-cli --remote, BuildTestCli) still trigger directly. Its orchestration surface is a real API modeled on the proven BuildTestApi shape. - Flaky-test-aware rerun surfaces let the hosted path select and rerun only the handful of failed / suspected-flaky tests, instead of requiring a full-suite retry. The warm no-recompile path and production activation remain future work.
- Richer dashboard surfaces now include a flake-frequency proxy chart, a confirmed-flake headline metric and daily trend, and build-correlated performance results read through from
url://performancetest/, modeled on PerformanceTestWui. These follow-ups landed in KompileRemoteBuildWui PR 15. - The remote build optimizer now exposes capacity planning, admission, fairness, rerun, and resource-stat surfaces. The service records the decisions needed to evolve toward the 10-minute target, memory-aware packing, fleet fairness, and visible pending admission; all 20 named optimizer scenarios are delivered, upstream peak-telemetry capture is delivered, and the full dynamic test-distribution end state remains future work while the live pull-dispatch plane is not yet active.
Current state
The CI system now has two supported submission paths. The direct path remains the default; the opt-in path routes through the hosted KompileRemoteBuild service. Production has not been switched to that opt-in path and remains pending an explicit user decision.
| Repository | Role now |
|---|---|
| kotlin-build-ci | GitHub App on ContainerNursery. With CI_BUILD_ROUTE=kompile-remote-build, it submits through the hosted adapter; an unset value or direct-buildtest keeps the direct url://buildtest/ path, and invalid values fail loudly. It still tracks phases, sets the kotlin.build (remote) check, handles /rerun·/bypass, downstream-dependency builds, and the build watchdog. |
| KompileRemoteBuildApi | Interfaces + data types for lifecycle, archive/chunked submission, logs, per-test/build-rule results, history, check verdicts, rerun decisions, queue introspection, capacity planning, and resource statistics. |
| KompileRemoteBuildEmbedded | Embedded implementation for first-hand CI event/projection capture, durable FIFO admission, rerun decisions, optimizer/resource projections, and the existing buildtest-wide analytics poller. |
| KompileRemoteBuildServiceServer | Hosts the Embedded behind url://kompile-remote-build/ as the durable planner-gated FIFO CI engine, with queue, verdict, rerun, optimizer, and resource-stat RPC surfaces. Published as kompile.remotebuild.server:kompile-remote-build-server:0.1.17, hosting Embedded 0.0.19. |
| KompileRemoteBuildCli / KompileRemoteBuildWui | CLI + server-rendered dashboard over url://kompile-remote-build/; the WUI has optimizer/capacity views, a flake-frequency proxy chart, a confirmed-flake headline metric and daily trend, and a build-correlated performance read-through from url://performancetest/. |
url://buildtest/ |
The default remote build/test execution environment (droplet pool). Named url://netlab-hosted/... routing is also available through BuildTestEmbedded and NetLabManagerServer — see Pluggable Execution Environments. |
Three consequences now matter:
- The direct path remains available, while the opt-in path removes the CI app's direct backend coupling. Production still uses direct
url://buildtest/until the explicit cutover decision is made. - Rerun and verdict surfaces are delivered, but the warm no-recompile path is not. The hosted path can make bounded, flake-aware rerun decisions and record authoritative outcomes; manual commands remain available, and warm classpath reuse stays a later enhancement.
- Capacity planning and admission surfaces are delivered. The service exposes planner, queue, optimizer, and resource statistics, and the exhaustive optimizer validation is delivered with current pull-dispatch gaps recorded; the full production cutover remains open.
Why it accelerates developers
- Flaky CI is exactly the tax component 1 exists to remove. A green check that needed three manual
/reruns is friction on every PR; for an agent swarm — "unforgiving consumers of friction" that "a flaky build … will stall completely" — it is worse than friction, it is a hard stop. Making the check self-heal on flakes turns a recurring human-in-the-loop interruption into a non-event. - Re-running only the failing tests, not the whole suite, is strictly cheaper — even in the interim (one recompile, then just the failing subset instead of every test). And once compilation is cached in the shared build cache (Kotlin Build) and test execution is a routable service (Pluggable Execution Environments), a rerun becomes pulling an already-built classpath and executing one test — seconds, no droplet rebuild at all.
- One front door is easier to evolve and secure. Folding build-triggering behind
url://kompile-remote-build/makes the GitHub App a thin presentation/integration layer over aurl://service (the standard layered architecture the whole plan leans on), giving a single place to add W3Wallet auth, quota, and observability. - Flake and performance trends in one place turn "is this test getting flakier / slower?" from tribal knowledge into a chart.
- Predictable build latency, and no capacity-luck reds. A 10-minute ceiling the system actively provisions for turns CI latency from "depends on the suite" into a budget every PR can count on; memory-aware packing removes the swap-storm/OOM failure mode that today masquerades as test flakiness; and a build that would today die on an over-quota droplet request instead waits visibly in a queue. For an agent swarm, "pending, position 3, starts in ~4 minutes" is actionable — "FAILED: droplet allocation" is an infrastructure flake that poisons trust in every red check.
Plan / roadmap
These milestones were proposed; the checked items below are now delivered, while the unchecked items remain future work. The remote build optimizer section remains deliberately sharp: its hard invariants and decided policy defaults are requirements, not implementation detail.
Route CI through url://kompile-remote-build/
- [x] Give
url://kompile-remote-build/a build-orchestration surface, modeled on the proven BuildTestApi shape. Delivered 2026-07-30/31 via KompileRemoteBuildApi PR 9 and KompileRemoteBuildServiceServer PR 17. The API now covers archive and chunked submission, lifecycle, logs, per-test/build-rule results, history, check verdicts, rerun decisions, and queue introspection;PENDINGis a first-class admission state. - [x] Make KompileRemoteBuildServiceServer the CI orchestrator over
url://buildtest/. Delivered 2026-07-31 by PR 17. The service owns CI admission and orchestration while buildtest remains a shared service forkompile-cli --remote, BuildTestCli, and BuildTestWui. - [x] Repoint kotlin-build-ci at
url://kompile-remote-build/. The opt-in seam delivered 2026-07-31 in PR 201, preserving directurl://buildtest/as the default. Production has not been repointed; settingCI_BUILD_ROUTE=kompile-remote-buildin production is pending an explicit user decision. - [x] Capture CI builds first-hand — in addition to, not instead of, the buildtest-wide poller. Delivered 2026-07-31 via KompileRemoteBuildEmbedded PR 10 and KompileRemoteBuildServiceServer PR 17. The broad poller remains, while CI submissions and decisions are recorded in the first-hand CI event/projection path.
- [x] Track each test's pass/fail status per branch (
mainand PR branches). Delivered 2026-07-31 through KompileRemoteBuildApi PR 9, KompileRemoteBuildEmbedded PR 10, and KompileRemoteBuildServiceServer PR 17.
Remote build optimizer — capacity planning, dynamic test distribution, and admission control
Today's capacity decisions are heuristics (see Current state): shard count from test count, static per-droplet test assignment, parallelism a constant, no memory signal, no fleet accounting, and allocation failure = failed build. The optimizer replaces guesswork with measurement, splitting the problem in two: capacity planning before the build runs (how many worker droplets, and how much concurrency each can safely offer) from the historical record of what each build rule and test actually costs, and dynamic distribution while it runs — tests are handed out to build servers as they finish their previous work, never carved into up-front shards, so the balance shifts in real time based on when servers actually complete. This component is critically important to get right — it sits in the allocation path of every CI build, and its failure modes (a memory-starved droplet in a swap storm, an exhausted droplet fleet, a build failing on a quota error) are precisely the trust-destroying infrastructure flakes component 1 exists to remove. It is correspondingly the most heavily-tested part of this workstream — and because it will be implemented largely by autonomous agents, its requirements are pinned down below with deliberate precision: the hard invariants are MUSTs an implementation cannot trade away, and the decided policy defaults are the initial contract (each is configurable later, but the behavioral shape is the requirement).
Vocabulary. A build's capacity plan is the pair (worker droplet count, per-droplet memory budget) chosen at admission. A droplet's resident set is the tests executing on it at a given moment. The fleet is every droplet in the provider account — CI-created or not. The admission queue is the ordered set of builds waiting as PENDING for capacity.
Hard invariants — an implementation that violates any of these is wrong, whatever else it does well:
- Never fail on allocation. No build ever reaches a failed state because a droplet could not be allocated — neither proactively (fleet accounting says none is free) nor reactively (a droplet-creation request throws an over-quota / cannot-allocate error). Both paths park the build in the admission queue as
PENDING. - Never overcommit memory. At every moment, on every droplet, the sum of the predicted peak memory of the resident set is ≤ the droplet's memory budget (RAM minus the headroom reserve). Admission is checked per pull against the actual current resident set — never against averages.
- No advance sharding. Which droplet runs a test is decided only at pull time, by an idle slot taking the oldest test in the queue that it can currently admit (memory-wise). There is no precomputed test→droplet assignment anywhere. A head test that does not currently fit is skipped past — it stays at the front for the next droplet with room (or runs alone under the oversized-test rule below), so head-of-line blocking never idles the fleet.
- No idling while work remains. A droplet with admissible queue items and a free slot pulls; it never sits idle while its build's queue holds a test whose predicted memory fits the droplet's remaining budget. When the queue is dry but another droplet still holds in-flight work, an idle droplet adopts unstarted leased work, then speculatively duplicates in-flight stragglers — tail rescue test execution, first terminal completion wins. (Invariants 3 and 4 bind on the pull-based dispatch plane — the work-stealing scheduler ExecutionEnvironments builds, which by decision lands early inside BuildTestEmbedded over the existing SSH transport, transport-agnostic at its core, rebinding to
url://when the BuildTest Runner Agent lands; until that plane lands, today's static sharding is a known interim deviation that receives no further investment, not a competing target.) - Exactly-once results. Every discovered test yields exactly one recorded result per build: an execution may be attempted more than once — re-dispatch after a droplet dies, an in-run flake re-lease, or a deliberate tail-rescue duplicate racing a straggler — but a result is recorded once (first terminal completion is authoritative; superseded attempts are kept as attempt history, never counted), and a test that can no longer be run is finalized FAILED with a descriptive error — never silently dropped, never counted twice.
- Fairness floor, no preemption. Every admitted (running) build holds ≥ 1 droplet until it completes. Scarcity shrinks the plans of newly admitted builds; it never takes droplets away from a running build.
- Fleet cap respected in aggregate. The sum of droplets across all concurrently planned + running CI builds, plus observed non-CI fleet occupancy, never exceeds the account's droplet limit — including under concurrent submission races.
- Every decision is explainable. Each plan, shrink, queue entry, requeue, and quota-retry is recorded on the run (and queryable via the API), with the inputs that produced it.
Decided policy defaults — initial values; each configurable, but the shape is the requirement:
-
The 10-minute target is measured from the first droplet-provision request to the last test result collected. Time spent
PENDINGin the admission queue is excluded from the target but reported separately (queue introspection: position + estimated start). -
Predictions: duration for capacity sizing = p50 of the trailing history; memory for admission = p95 of recorded peaks. Trailing history = the last 100 samples within the last 30 days, per key. No-history fallbacks: duration = the declared
@Timeout, else 30s (matching today's sharding fallback); memory = a deliberately generous 2 GB default. -
Headroom reserve: per-droplet memory budget = droplet RAM − max(1 GB, 20% of RAM), covering OS + runner + compile residue.
-
Diminishing returns: stop adding droplets when the next droplet improves the predicted finish by < 30 seconds (and never plan more droplets than tests).
-
Scarcity threshold: fairness mode engages when free fleet capacity (limit − occupancy) drops below 20% of the limit; new plans are shrunk largest-first, floor of 1 droplet each.
-
Queue order: FIFO by submission time. Because the fairness floor means any build can start on a single droplet, the head starts as soon as one droplet frees — no reordering, no jumping, no starvation.
-
Quota-retry backoff: after a provider over-quota/allocation exception, retry allocation with exponential backoff from 30s doubling to a 5-minute cap, driven by the
Clockabstraction (ManualClock-testable), until capacity frees. -
Oversized tests never deadlock the queue: a test whose predicted peak memory exceeds a droplet's entire budget is dispatched alone to an idle droplet (a resident set of one) — sized up to a larger droplet slug when history warrants — rather than waiting forever for a budget it can never fit. If it then genuinely OOMs, that is the test's own failure, reported with full context; the invariant-2 budget check applies to combining tests, not to refusing a lone oversized one.
-
[x] Track per-rule / per-test compute metrics — including memory — captured upstream in buildtest. Delivered upstream on BuildTestEmbedded
main: the CAPTURE pipeline records VmHWM-based test-JVM peaks and build-rule runner peaks, with(type, name)-keyed rollups including p95 and sample counts plus restart persistence. The strengthened public-path proof is BuildTestEmbedded PR 504, which isOPENand blocked on the shared-transport incident recorded in PlanRepository PR 5114; the proof remains pending infrastructure recovery.url://kompile-remote-build/ingests these into its event-sourced history and serves the derived statistics — percentiles, variance, and trend alongside averages and sample counts, since a mean alone under-provisions bimodal tests (the planner consumes the percentiles fixed in the decided defaults above: p50 duration for sizing, p95 memory for admission). Executions with no history get the conservative fallbacks from the decided defaults (declared@Timeoutelse 30s; 2 GB memory), which biases first runs strongly against overcommit — but is not a guarantee: a first run whose true peak exceeds the default can still overshoot. That is acceptable by design — the memory invariant is defined over predicted peaks — and self-correcting: the failure is reported with full resource context, the measured peak enters the history, and the next run is planned against reality. Build-rule/compile peaks are recorded for a different consumer than test peaks: compilation is often a build's true memory high-water mark, so rule peaks inform droplet size selection (a workspace whose compile history exceeds the standard slug's budget gets a bigger slug), while test peaks drive per-pull test admission. -
[x] Plan droplet count against the 10-minute target. Delivered 2026-07-31 via KompileRemoteBuildEmbedded PR 10 and KompileRemoteBuildServiceServer PR 17. Capacity plans, queue admission, fairness decisions, and resource statistics are part of the embedded/service surface; exhaustive optimizer validation is delivered below via KompileRemoteBuildEmbedded PR 29.
-
[ ] Hand tests out in real time — no advance sharding. A build's discovered tests form a single shared queue; each worker droplet pulls its next test the moment a slot frees, so the distribution shifts continuously based on when servers actually complete their previously-assigned work. A test that runs 10× its predicted duration delays only its own droplet's next pull — the remaining queue drains through the other workers — and a droplet that dies mid-test has its in-flight tests requeued rather than stranding an entire pre-computed shard. At the very end of a run, tail rescue test execution closes the last gap: an idle droplet adopts a busy droplet's unstarted work, and when only in-flight stragglers remain it speculatively duplicates them and races — first completion wins, exactly-once results — so a wedged droplet or droplet-specific flake bounds the tail by the healthiest worker, not the sickest. This is the work-stealing scheduler ExecutionEnvironments builds for the default runner pool; this workstream consumes that dispatch plane and adds the layer above it (capacity sizing, memory admission, fleet fairness, build queueing) rather than building a second scheduler.
-
[x] Pack test parallelism by measured memory, not a constant. Delivered 2026-07-31 via BuildTestEmbedded PR 441 and KompileRemoteBuildServiceServer PR 17. The planner uses predicted peak memory for per-pull admission instead of a fixed parallelism constant.
-
[x] Fleet-limit awareness, degrading to fairness under scarcity. Delivered 2026-07-31 via KompileRemoteBuildEmbedded PR 10 and KompileRemoteBuildServiceServer PR 17. The optimizer accounts for fleet capacity and records fairness-constrained plans.
-
[x] Admission control: exhaustion and quota errors queue the build — they never fail it. Delivered 2026-07-31 via KompileRemoteBuildEmbedded PR 10 and KompileRemoteBuildServiceServer PR 17. Exhausted or refused capacity is represented as
PENDINGwith queue information and retry behavior rather than an allocation-caused red result. -
[x] Close the loop: every run sharpens the model. Delivered 2026-07-31 via KompileRemoteBuildEmbedded PR 10 and KompileRemoteBuildServiceServer PR 17; resource history and predicted-vs-actual data are exposed by the service surface.
-
[x] Test the optimizer exhaustively — a minimum bar, not a ceiling. Delivered 2026-08-05 via KompileRemoteBuildEmbedded PR 29, with all 20 named scenarios passing, included, and merged. Triage classified 10 test-side issues and found and fixed one engine defect: explicit prediction outcomes were durable in individual decisions but were not recorded in the aggregate prediction-accuracy query. The scenarios that cannot observe the live pull-dispatch plane assert plan-level properties and record that gap. Per the testing standards: end-to-end through the public API against a real BuildTestEmbedded backed by FakeDropletService (never real DigitalOcean droplets), deterministic via
ManualClock, with fault injection through real in-memory Api implementations. The suite covers edge cases, error conditions, boundary values, concurrency, and integration paths:lightweightSuiteRunsWidelyParallelOnFewDroplets— hundreds of tests with small recorded memory peaks run at high per-droplet parallelism (on the order of 100-wide) on a small droplet count, and the run's recorded plan shows that packing.heavyTestsCapPerDropletParallelism— tests with multi-GB recorded peaks run at low parallelism; the sum of concurrently-resident predicted peaks never exceeds droplet RAM minus headroom, asserted over every admission decision, not on average.mixedSuitePacksHeavyAndLightWithoutOvercommit— heavy and light tests interleaved: the combined-residency constraint holds at every pull while the light tests still run wide; a heavy test that would overshoot waits for residents to finish rather than being admitted.dropletCountMeetsTenMinuteTargetFromHistory— a suite whose history predicts ~40 droplet-minutes of work plans enough droplets that the predicted finish is under 10 minutes; a small suite plans exactly one.slowTestShiftsDistributionToOtherDroplets— one test injected to run 10× its recorded duration: the shared queue drains through the other droplets in the meantime (their completed-test counts grow while the slow droplet is occupied), and total wall clock is bounded by the slow test, not by a stranded pre-assigned shard.deadDropletsInFlightTestsRequeueToSurvivors— a droplet killed mid-run has its in-flight tests handed back out to the surviving droplets; every test still completes exactly once, none dropped, none run twice.noDropletIdlesWhileQueueNonEmpty— with uneven test durations, no worker sits idle while unstarted tests remain in the build's queue (the real-time-handout invariant, asserted via the dispatch event log rather than wall-clock timing).newTestsWithNoHistoryGetConservativeEstimates— a first-ever test is planned with its declared@Timeout(or the 30s default duration) and the 2 GB memory default — never zero.pendingQueueWaitReportedSeparatelyFromRunDuration— a build that waited in the admission queue reports queue wait and run duration as distinct figures; the 10-minute target accounting covers provision-to-last-result only.fairnessShrinksLargeBuildsWhenFleetRunsLow— with the fleet near its cap, a build that would plan N droplets is shrunk — its predicted finish allowed past 10 minutes — rather than exhausting the remaining capacity; the constrained plan is recorded on the run.noBuildStarvedBelowOneDroplet— under maximum contention, every admitted build still holds ≥1 droplet and completes.exhaustedFleetQueuesBuildAsPending— a submission arriving when zero droplets are free yields a PENDING run reporting queue position and estimated start through the API; no droplet-creation attempt is made and nothing fails.quotaExceptionRequeuesInsteadOfFailing— FakeDropletService injected to throw an over-quota allocation error: the build lands back in the queue as PENDING, partially-allocated droplets are released (none leaked), the requeue is recorded in the run's log, and the check never reds.oversizedTestRunsAloneRatherThanDeadlocking— a test whose predicted peak exceeds a whole droplet's budget is still dispatched (alone, to an idle droplet) and the build completes; the queue never deadlocks on it.queuedBuildsStartWhenCapacityFrees— completing or cancelling a running build promotes the queue head automatically (ManualClock-driven), in submission order.cancelledPendingBuildLeavesQueueConsistent— cancelling a queued build removes it without perturbing other builds' positions or leaking reserved capacity.concurrentSubmissionsNeverOverAllocate— many builds submitted in parallel (bounded-concurrency stress) never plan beyond the fleet limit in aggregate, and the queue + running set accounts for every submission exactly once.adHocBuildtestUsageCountsAgainstFleetView— droplets created by direct buildtest callers (BuildTestCli,kompile-cli --remote) reduce what the optimizer plans; CI never pushes the fleet over its limit because of ad-hoc load it did not create.predictionErrorIsRecordedPerRun— predicted-vs-actual duration and memory deltas are queryable after each run.planIsDeterministicForIdenticalInputs— identical history + suite + fleet state produces an identical capacity plan (droplet count + per-droplet memory budget), with no wall-clock or randomness dependence; only the runtime dispatch order varies with actual completion times.
Build watching — absorb build-watchman into the CI service
build-watchman exists today as a local CLI (bundled into the watch-build skill), born from the 2026-07-05 incidents (CI-trigger fragility handoff). It watches a PR to green across every layer that failed that day: check state, dispatch state (a suite queued with zero runs 60s after GitHub attempted it = a lost CI notification, with the exact rerequest remediation), forward progress against a reference trajectory reconstructed from the repo's own recent successful runs (checked at every 10% mark; resume-polluted or degenerate history is rejected), duration-vs-history, and the merge queue (queue-ref validations get the same analysis; an ejection is an explicit alert with the re-enqueue procedure). Every finding is one OK / CHANGE / PROBLEM line carrying its next action, and the watcher itself can never fail silently.
- [x] Refactor build-watchman's functionality into a separate shared GitHub project (a watching library per the standard layered architecture — the CLI becomes a thin client over it) so the same detection logic is integratable into kotlin-build-ci / the CI service, not only runnable by an operator. Delivered 2026-07-30 via build-watchman PR 9; the published CLI is
buildwatchman:build-watchman:0.0.14. - [x] Integrate it server-side. Delivered 2026-08-05 via kotlin-build-ci PR 211: the CI service detects lost check-suite dispatches within 60 seconds and self-rerequests with bounded remediation, subsuming
MissingCheckRunReconciler; it applies trajectory/stall detection with bounded redrive to orchestrated runs, watches merge-queue refs, and records durable remediation decisions.
Flaky-test-aware reruns — green without a full rebuild
- [x] Selective re-execution of failed tests (per-test, batched). Delivered 2026-07-31 via KompileRemoteBuildApi PR 9, KompileRemoteBuildEmbedded PR 10, and KompileRemoteBuildServiceServer PR 17. The service exposes per-test rerun selection and bounded batched execution; warm, no-recompile reruns remain a later enhancement.
- [ ] End-state: warm, no-recompile per-test reruns. Once Pluggable Execution Environments + the shared build cache land (compile once → publish the classpath to the shared cache → run a single test on a warm runner that pulls it), a rerun becomes "enqueue the failing test(s) to a warm runner" — no droplet provision, no recompile. Better-with, not blocked-on.
- [x] The rerun result is authoritative — reruns gather evidence, they never forgive. Delivered 2026-07-31 via KompileRemoteBuildEmbedded PR 10 and KompileRemoteBuildServiceServer PR 17. A test is greened only when a retry actually passes.
- [x] Which failures are "believed flaky" (and so worth rerunning). Delivered 2026-07-31 through the rerun-decision surface in KompileRemoteBuildEmbedded PR 10 and KompileRemoteBuildServiceServer PR 17. The service records the evidence used for rerun decisions; the production adapter remains opt-in.
- [x] Bounded rerun budget (counted in attempts). Delivered 2026-07-31 via KompileRemoteBuildEmbedded PR 10 and KompileRemoteBuildServiceServer PR 17, including the per-build and per-test attempt caps.
- [x] Green the check when reruns pass — and record why. Delivered 2026-07-31 via KompileRemoteBuildApi PR 9, KompileRemoteBuildEmbedded PR 10, and kotlin-build-ci PR 201; verdicts and retry summaries are part of the service/adapter surface.
- [x] Coexist with the manual
/rerunand/bypasscommands. Delivered 2026-07-31 in kotlin-build-ci PR 201, which preserves the existing direct-path behavior and maps the hosted lifecycle/verdict surface.
Flake tracking & visualization
- [x] Flake-frequency-over-time charts. A proxy chart shipped 2026-07-30 in KompileRemoteBuildWui PR 14. KompileRemoteBuildWui PR 15 adds the confirmed-flake headline metric and daily trend, counting only a failed initial outcome followed by a same-code pass and final
PASSEDstatus.
Performance-test results in the dashboard
- [x] Surface the performance dashboard. Links to the performance dashboard shipped 2026-07-30 in KompileRemoteBuildWui PR 14. KompileRemoteBuildWui PR 15 adds the stateless build-correlated
url://performancetest/read-through, inline summary/histogram, and detail links without duplicating analytics.
Cross-cutting
- [ ] Standard-layered-architecture compliance. The GitHub App and WUI are HTTPS frontends that talk only
url://and hold no durable state on disk (per the architecture rules); all domain logic and state live in the Embedded behindurl://kompile-remote-build/. - [ ] W3Wallet-gated access. Trigger-a-build and the dashboards move from "ContainerNursery routing provides bootstrap access control" (today's stopgap, noted in both service READMEs) to real W3Wallet capabilities — a viewer capability for the dashboards, a trigger capability for builds.
- [x] End-to-end tests per the testing standards. Delivered 2026-08-05 via KompileRemoteBuildServiceServer PR 29: through the public
url://kompile-remote-build/surface, a fail-once-then-pass test records both attempts and a GREEN verdict, while an always-failing test records all three attempts and remains RED. The scenarios use the real BuildTestEmbedded/FakeDropletService fixture and isolated in-memory URL nodes. - [ ] Documentation. Keep the six cross-linked READMEs and a future Documentation Repository project page in sync as these land.
Decisions
Resolved as the design firms up; kept here so the rationale survives.
- buildtest stays a shared service — the CI service calls it, doesn't own it.
url://kompile-remote-build/owns CI semantics (which build a PR needs, the rerun policy, check status, result history) and drives buildtest for CI builds, but buildtest remains independently callable bykompile-cli --remote, BuildTestCli, and BuildTestWui. buildtest's own evolution — theTestRunnerApirunner plane, work-stealing, build-once-run-anywhere — belongs to Pluggable Execution Environments; this workstream consumes that plane rather than absorbing it. So "the GitHub App should stop talking to buildtest directly" removes the app's coupling — it does not make the CI service buildtest's sole owner. - The production route cutover is an explicit pending decision. kotlin-build-ci PR 201 provides and validates the opt-in seam, but production remains on direct
url://buildtest/; do not setCI_BUILD_ROUTE=kompile-remote-buildin production without an explicit user decision. - GitHub I/O stays in the App; the CI service is GitHub-agnostic. The GitHub App owns all GitHub-specific work — webhooks, installation tokens, source-tarball download (and combining the up-to-44 downstream tarballs), check-run rendering, comment commands — and hands
url://kompile-remote-build/a source archive + ref metadata. The CI service never holds GitHub credentials and owns the GitHub-agnostic half: build orchestration, the buildtest upload, and reruns. This preserves today's clean split (theKompileRemoteBuild*side has zero GitHub knowledge) and keeps the service reusable by non-GitHub frontends. - Rerun granularity is per-test, and the feature ships on today's buildtest. Reruns are selected and accounted per test (buildtest's native
--testallow-list /testSpecs) and executed batched — onesubmitBuild(testSpecs=…)droplet per round, ≤3 rounds. The interim cost is one recompile per rerun round (today's buildtest recompiles per run;overrideJarRunIdis self-bootstrap-only, not classpath reuse) — accepted, because skipping the passing-test retest is most of the win. Warm, no-recompile per-test reruns are gated on Pluggable Execution Environments + the shared cache — better-with, not blocked-on. - Rerun policy is a cost control, not a trust gate. The rerun outcome is authoritative — a test is greened only by actually passing a retry — so the policy only decides how much droplet time to spend, never what to forgive. The decided knobs: a failing test is "believed flaky" (eligible to rerun) if it had ≥1 green run in the larger of {last 10 builds, last 24 hours}, counting runs across
main∪ the PR branch, or it carries a transient/infrastructure error signature; total rerun attempts per build are capped atmax(6, ceil(3% × total tests))(3% of total tests, not failing) with ≤3 attempts per individual test. Eligibility is intentionally inclusive (a retried regression simply stays red), so the "green-on-main+ PR-fail = likely regression" signal is a ranking input, not a gate: when the budget binds, genuinely-flaky-on-mainand transient failures retry first and reliably-green-on-main-but-PR-failing tests last. Per-test status is tracked per branch in v1, captured first-hand (today'sBuildRunhas no branch field). Build/compile failures are never rerun. - A flake is same-code non-determinism — and the rerun confirms it. A failing test becomes a flake candidate (eligible for auto-retry) per the believed-flaky rule; the service then retries it up to 3 times on the unchanged tree. If any same-code rerun goes green — even with low probability — the code provably can pass, so that failure and every other red on the same code are confirmed flakes (non-deterministic noise, not a defect). A candidate that never greens across its retries is treated as a real failure. This is the ground truth behind "reruns are authoritative," and it is exactly what the flake-frequency metric counts.
- Performance results are federated, not absorbed.
url://performancetest/is already a sibling analytics service fed by the same pipeline (the BuildTest dispatcher auto-submits@PerformanceTestruns). The CI service exposes a thin read-through to it (no re-ingestion, no duplicate analytics); the WUI renders an inline summary + histogram and deep-links to PerformanceTestWui for detail. Unifying the two analytics services is a non-goal unless a concrete cross-analytics query emerges. - Flake-reruns apply to all three check contexts — including the merge queue, from day one. The policy runs wherever the
kotlin.build (remote)check is produced: primary PR builds, merge-group (merge-queue) builds, and downstream-dependency builds (up to the 44 kompile-cli-tree repos combined under one shared conclusion). A flake is most expensive in the merge queue (it evicts the PR and re-runs the whole group) and most numerous downstream. Merge-queue reruns are enabled from day one with a tight budget that fits the existingcheck_response_timeout_minutes— the timeout is never raised to accommodate them — accepting some eviction risk on rerun-heavy builds until the warm/no-recompile end-state removes the recompile latency. For downstream builds the budget is computed over the union of all combined repos' tests, with per-testmainstatus keyed by(repo, test). - Auto-rerun coexists with the manual commands, it doesn't replace them.
/rerun//retry//retriggerstays the human escape hatch (a fresh full rebuild + retest, itself subject to auto-rerun);/bypassis unchanged and is distinct from auto-greening — bypass forces green without building, auto-rerun greens only by tests actually passing — and should be needed less once auto-rerun lands. - The optimizer's policy lives in
url://kompile-remote-build/; measurement and enforcement live in buildtest. buildtest — which provisions the droplets and runs the tests — is where per-execution duration/memory telemetry is captured (per the upstream-fix principle: measured where execution happens, never approximated downstream), where the fleet's live occupancy and the account's droplet limit are read, and where a plan's droplet count and per-droplet memory budget arrive as submission parameters (per-test dispatch is the pull queue's, not a parameter). The CI service owns the policy: the 10-minute target, memory-aware packing, fairness under scarcity, and the CI admission queue. One piece lands upstream in buildtest itself: queue-on-quota semantics — a droplet-creation over-quota/allocation error parks the run asPENDINGuntil capacity frees instead of failing it — because provisioning (and therefore the exception) lives there, and every caller (ad-hoc included) deserves the never-fail-on-allocation behavior, not just CI. The two queues are layered, not duplicates: the CI admission queue is the proactive layer (it plans against fleet accounting before asking buildtest for anything, so CI builds should rarely reach buildtest's backstop), while buildtest's queue-on-quota is the reactive layer at the provisioning boundary where the exception physically occurs. When CI and ad-hoc submissions are both waiting on the same freed capacity, buildtest serves them FIFO — CI claims no priority over ad-hoc callers, and the optimizer sees ad-hoc queued demand the same way it sees ad-hoc occupancy: as fleet load it must plan around, never plan over. - 10 minutes is a target, not an SLA — and fairness beats latency when they conflict. The planner provisions to finish under 10 minutes when capacity allows; under scarcity, builds get slower (shrunk allocations, fairly distributed) rather than denied, and under exhaustion they get later (queued as pending) rather than failed. The strict invariants are "no droplet is memory-overcommitted" and "no build fails for allocation reasons" — never "every build is fast."
- Sizing is the optimizer's job; dispatch is dynamic, pull-based, and shared with ExecutionEnvironments. The optimizer decides how many droplets a build gets and the memory budget governing each droplet's concurrency, from historical duration/memory statistics; which test runs where is never decided in advance — droplets pull from the build's shared queue as they finish their previous work, so distribution rebalances in real time around slow tests, fast droplets, and dead workers, and the end-of-run tail is closed by tail rescue (idle droplets adopt unstarted work, then race duplicates of in-flight stragglers, first completion winning). That pull-based dispatch plane is the work-stealing scheduler ExecutionEnvironments builds for the default runner pool; this workstream consumes it rather than building a second scheduler, and deliberately does not invest in a smarter static partitioner in the meantime — today's duration-balanced pre-sharding is retired in favor of dynamic handout, not replaced by a cleverer version of itself. The historical stats thereby regain a planning role — capacity sizing and memory admission — without reintroducing the per-test static assignment ExecutionEnvironments retires; the two workstreams compose rather than compete. The optimizer's other halves (fleet accounting, fairness, admission queueing) do not wait on the scheduler: they are valuable under any dispatch model and land independently — while the dynamic-handout invariants themselves (hard invariants 3–4) bind once that dispatch plane lands.
- Memory estimates are measured, not modeled. Per-test peak memory comes from the test JVM's own high-water mark captured by the runner; predictions use percentiles and variance over that history, with deliberately conservative defaults (declared
@Timeout/ generous memory) for executions with no history. An estimate that proves wrong shows up as a recorded predicted-vs-actual delta, not a silent correction. - CI decisions are event-sourced; the broad analytics keep ingesting all buildtest runs. The in-path service is the system of record for CI decisions (rerun attempts/recoveries, check verdicts, per-test branch/
mainstatus) — data nothing else holds — recorded CQRS-style as an append-only EventLog. But the buildtest-wide poller is retained, not retired: the flake/duration/pass-rate analytics must keep covering every buildtest run, including ad-hockompile-cli --remote/ BuildTestCli runs. So analytics are projections over both the first-hand CI event stream and the polled buildtest population; the okio JSON rollup is a replayable read-model. Event-sourced shape now (local okio), swapping to a hostedurl://event-log/later. The unbuilt database service is not a dependency.
Open questions
- Production CI route cutover — pending explicit user decision. The opt-in seam is merged and the hosted service is published, but production remains direct. The decision is whether and when to set
CI_BUILD_ROUTE=kompile-remote-buildin the production GitHub App.
Related
-
build-watchman — the local PR/build/merge-queue watcher whose functionality this workstream absorbs into the CI service (see "Build watching" above).
-
Kotlin Build — the toolchain underneath; its shared build cache is what makes "rebuild nothing, re-run one test" possible.
-
Pluggable Execution Environments — reframes
url://buildtest/as the default test runner and introduces compile-once / route-a-single-test execution with work-stealing dispatch — the substrate selective reruns ride on. -
PerformanceTest / PerformanceTestWui — the model (and likely the source) for performance-result rendering in the dashboard.
-
UrlResolver — the P2P fabric the GitHub App uses to reach
url://kompile-remote-build/. -
W3Wallet — capability-based access control for triggering builds and viewing dashboards.
-
EventLog / Observables — candidate substrates for first-hand result capture and live-updating dashboards.
-
Implementation repos: kotlin-build-ci, KompileRemoteBuildApi, KompileRemoteBuildEmbedded, KompileRemoteBuildServiceServer, KompileRemoteBuildCli, KompileRemoteBuildWui.
Graduation
The core route seam, hosted orchestration, bounded rerun surfaces, optimizer/capacity surfaces, exhaustive optimizer validation, server-side build-watchman integration, WUI follow-ups, cross-cutting public-URL coverage, and named NetLab execution-environment route are delivered. This workstream graduates once production builds are explicitly switched to url://kompile-remote-build/ and the end-to-end check is verified on the hosted path in production. Warm, no-recompile reruns remain an enhancement owned by Pluggable Execution Environments, not a production-cutover prerequisite.