← Workstreams

Workstream: PerformanceTest

Status: Planned · Component: Maximize developer productivity · Docs: project page

Goal

Make "did this change make it faster or slower?" a routine, trustworthy answer. A developer marks a test as a performance test, and the platform answers — per test, per commit, and as a trend over a branch's history — whether performance improved or regressed, with enough statistical honesty that the answer survives noisy cloud hardware.

The core problem: everything drifts

Absolute performance numbers are unstable over time. A different physical machine behind the same droplet size, a hypervisor or kernel update, a JVM point release, a noisy neighbor — any of these can move absolute timings by large factors, swamping the effect of the code change being measured. A history of absolute numbers is therefore a history of the fleet, not of the code.

The mitigation this workstream is built around: never trust an absolute number across sessions; only trust ratios measured within one session. A co-run executes two or more candidate versions of the code interleaved on the same machine at the same time-window. Whatever the hardware happens to be doing that day affects every version in the co-run roughly equally, so the relative performance between versions is stable even when the absolute numbers are not. All cross-time analysis is then built out of these within-session ratios, never out of raw timings.

Current state

This extends the existing PerformanceTest project — a full Api / Embedded / Runner / ServiceServer / WUI family (PerformanceTestApi, PerformanceTestEmbedded, PerformanceTestRunner, PerformanceTestServerService, PerformanceTestWui) serving url://performancetest/ with a dashboard at perftest.kotlin.build. Today it already covers the single-version story end to end: a test annotated @PerformanceTest(measuredIterations, warmupIterations) is routed by the BuildTest dispatcher to a dedicated droplet, run with warmup and per-iteration timeouts, and reported with summary statistics (min/max/mean/median/p95/p99/stddev), an SVG histogram, and a per-test run-history page.

What does not exist yet is everything comparative: there is no co-run (each run measures one version on its own droplet), no cross-version ratio, no baseline tracking, no regression verdicts, and no trend charts — comparing two runs today means a human eyeballing two absolute numbers measured on different hardware days apart. This workstream adds that comparative layer.

The relative-performance model

Each co-run of versions u and v yields a paired measurement of their log-duration difference — an observation of θ(u) − θ(v), where θ is the latent log-performance index of a version. Estimating per-version indices from a graph of such pairwise observations is a well-studied problem: it is the continuous-outcome generalization of Bradley–Terry / Elo / TrueSkill rating models, and it is exactly the repeat-sales regression used for the Case–Shiller house price index in finance — which estimates a market-wide price trend purely from pairs of sales of the same house, for precisely the same reason: absolute prices drift with the market the way absolute timings drift with hardware.

Because it is genuinely unclear which estimator best reflects reality on our workloads, the analysis layer is pluggable, and several strategies ship from day one, each producing its own trend line so they can be compared side by side on real data:

  1. Chained ratio (baseline) — each version's index is the previous version's index times the median within-session ratio between them. Trivial to understand, accumulates error along the chain; included as the honesty check the fancier models must beat.
  2. Graph least-squares (primary) — a weighted least-squares fit of all θ values over the entire co-run graph at once (the repeat-sales estimator; equivalently a mixed-effects model with a random effect per session). Every historical co-run constrains the fit, indices are re-estimated as new evidence arrives, and each index carries a confidence interval derived from the residuals.
  3. Online Gaussian rating (streaming) — a TrueSkill-style incremental update where each version holds a Gaussian belief (μ, σ) refined by each co-run. Cheap, order-sensitive, never revises history; included to see how much the global fit actually buys.
  4. Bayesian mixed-effects (reference) — the full generative model: per-session random effects, heavy-tailed noise so a single pathological iteration cannot drag an index, posterior credible intervals on every estimate. The most statistically defensible of the four and the most expensive to compute; it serves as the reference the cheaper estimators are judged against.

Indices are relative by construction — each connected component of the co-run graph is anchored at an arbitrary reference version (index 1.0), and everything else is expressed as a speed-up/slow-down factor against that anchor.

Plan / roadmap

  • [ ] Co-run execution. Extend the submit contract so one submission carries k ≥ 2 versions of the same test. All versions of a co-run execute on one droplet, each version in its own freshly-launched JVM (so JIT profiles never cross-contaminate), never concurrently (so they never contend), interleaved as alternating iteration blocks in randomized order across multiple rounds — so slow drift within the session (thermal, neighbor load) also averages out of the ratios. The existing single-version run remains as the degenerate k = 1 case.
  • [ ] Cold-start mode. @PerformanceTest grows an execution mode: the default warm mode keeps today's behavior (iterations timed inside one long-lived JVM), while COLD_START launches a fresh process for every iteration and times it from process start to completion — measuring the class of performance the warm runner structurally cannot: JVM boot, service bind time, first-request latency (the ContainerNursery lazy-start path is the canonical example of why this class matters). Co-runs interleave cold-start iterations across versions exactly as warm mode interleaves blocks.
  • [ ] Version identity and baseline re-runs. A measured version is identified by the git commit (plus repo, branch, and testSpec) that produces its classpath bundle — the submit contract grows these as first-class fields alongside today's free-form label. The service deliberately does not retain bundles: when an old version is needed for a fresh co-run, it is rebuilt from source at its commit, which the Kotlin Build shared build cache makes cheap for any commit that has been built before. When old code has rotted and no longer builds or runs, its index remains estimable anyway: the graph least-squares and Bayesian models connect it to the present through the historical co-runs it already participated in. Fresh baseline re-runs are therefore a bonus (they add edges and shrink confidence intervals), never a requirement.
  • [ ] Analysis strategies. The four estimators above are domain-agnostic math over pairwise ratio observations, so they live in a standalone reusable library (community.kotlin.statistics.pairedcomparison) — one estimator interface, four implementations, testable in isolation against synthetic data. PerformanceTestEmbedded consumes it, feeding stored co-run results in and surfacing the outputs (per-version indices, per-pair verdicts, confidence intervals) through new Api read methods.
  • [ ] Raw observations are the permanent record. The service retains the raw per-iteration observations of every run and co-run indefinitely; version indices, verdicts, and charts are derived data, recomputable at any time. This is what keeps the estimator layer honestly pluggable — a new strategy is backtested by replaying the full history, without re-measuring anything — and it keeps the capabilities that depend on replay open (A/A self-co-runs to calibrate estimator false-positive rates, automatic bisection of a regression window). Storage cleanup may compact or archive, but never discards raw observations.
  • [ ] Adaptive stopping. Co-runs are sized sequentially rather than with a fixed iteration count: start with a few rounds, stop early once the ratio is decisive in either direction (clearly unchanged or clearly moved), and extend only while the verdict is genuinely ambiguous — using a sequential test (SPRT-style) designed for repeated looks, so stopping early does not bias the verdicts. Most changes do not move performance; this concentrates droplet spend on the few co-runs where more evidence actually changes the answer.
  • [ ] Profiling as a first-class lens. Not just how long, but where the time went: runs can execute with a profiler attached (async-profiler / JFR), giving every performance test an up-to-date flamegraph in the WUI as a routine artifact — and a co-run renders the differential flamegraph between its versions. A confidently-regressed verdict automatically triggers a profiled re-run, so a red verdict arrives with its explanation attached. Profiled blocks are always additional and excluded from the timing observations, so profiler overhead never contaminates the estimators' ratios.
  • [ ] Trend and impact charts. Two new WUI views, rendered — like the existing histogram — as server-side SVG with no client-side JavaScript. The branch view plots a test's performance index over a branch's commit history (one line per analysis strategy, with confidence bands) so improvement and regression trends are visible at a glance. The commit view answers "what did this commit do?": for a given commit, every performance test's speed-up/slow-down factor against its parent, with confidence intervals.
  • [ ] Anti-creep anchor runs. Per-PR verdicts are blind to the boiling-frog failure mode: a long series of individually-insignificant regressions that add up to a slow system. On a schedule, the service co-runs main's head against a fixed older anchor (e.g. the last release tag), rebuilt from its commit like any baseline. That measures the cumulative drift directly, and the long edges it adds to the co-run graph stiffen the least-squares fit across the whole history.
  • [ ] CI integration: part of the existing test plan, not a bolt-on. Performance tests run where they already run — inside the sharded test execution the kompile-cli build performs on every PR, with @PerformanceTest routing a test to url://performancetest/ exactly the way Pluggable Execution Environments routes any test to its declared runner. The comparative layer extends that same dispatch: on a PR, the routed execution becomes a co-run of the PR head against its merge-base with main, built like any other shard (the merge-base build is normally a shared-build-cache hit). Regression verdicts are always informational: a performance test fails its shard only on execution errors (crash, timeout), never for being slower — a confident regression surfaces in the check output, the dashboards, and the charts, where a human decides what it is worth. Verdicts therefore surface through the normal check results and the CI dashboard's read-through federation, per the Kompile Remote Build plan: url://performancetest/ stays the single system of record for performance analytics, and KompileRemoteBuildWui deep-links into the new comparison views rather than re-ingesting them.

Why it accelerates developers

  • Performance regressions get caught when they are one commit big, not six months later when "the system feels slow" and fifty commits are suspects.
  • It closes the last unverified dimension of a change. Correctness is gated by Kompile Remote Build; this makes speed a first-class, per-PR answer with the same zero-ceremony contract — one annotation, no bespoke harness.
  • It makes cheap hardware trustworthy. Because only within-session ratios matter, performance testing can run on ordinary shared droplets rather than dedicated bare-metal runners — the statistics absorb what the infrastructure cannot guarantee.

Graduation

This workstream graduates when co-runs execute end to end, at least two analysis strategies chart the same real test's history side by side in the WUI, and one repository's PRs receive candidate-vs-merge-base verdicts routinely. At that point the comparative layer folds into the existing PerformanceTest project page as shipped behavior.