Explain and fix the three testrunner timeout tests that passed only on a retry in CI run 33e21cc1
Written 2026-10-08 21:30 UTC. RE-VERIFY before acting: every state below is a write-time snapshot. Re-check the PR with gh pr view 54 --repo CodexCoder21Organization/community.kotlin.kompile.testrunner.jvm --json state,mergedAt, the evidence branches with git ls-remote, and the buildtest runs at the linked URLs.
Mission summary
https://github.com/CodexCoder21Organization/community.kotlin.kompile.testrunner.jvm/pull/54 (merged 2026-10-08 21:25 UTC) makes the testrunner's per-test timeout stop the whole owned process tree and run its diagnostics budgets on an injected clock. Its remote CI run https://buildtest.kotlin.build/run?id=33e21cc1 passed overall, but three of the process-tree timeout tests failed on their first attempt and passed on attempt 2 (and attempt 3). The PR landed with this retry class recorded as a known limitation because two investigation lanes could not reproduce it locally or on the real remote runner. This handoff is the follow-up: find the mechanism that makes a first attempt fail, pin it with a reproducer, and fix it in the product or the test harness.
What was found (chain)
- The three retry-only tests and the companion test examined are among:
testTimeoutStopsParentStoppedDuringCleanup, testTimeoutStopsObservedChildAfterParentExit, testTimeoutStopsNestedJvmCreatedByShutdownHook, testTimeoutStopsChildCreatedByDescendantShutdownHook. The exact three, their first-attempt console outputs (9,121, 63,685 and 62,529 characters) and the attempt-by-attempt comparison are in https://github.com/CodexCoder21Organization/community.kotlin.kompile.testrunner.jvm/blob/wip/F54G-evidence-20261008/evidence/F54G/result.md.
- The failing first attempts show a missing nested-hook PID publication at timeout completion and unfinished inner scenarios in two of the tests; the outputs do not reveal which phase was interrupted or blocked. None of the three uses the clock-driven diagnostics driver. Attempt 2 passed for every target; the three failures did not share one droplet. Refuted: no per-attempt output endpoint exists in the WUI; the complete journal is the source.
- Local reproduction (lane F54F): 0/10 failures with the amplifiers tried (evidence https://github.com/CodexCoder21Organization/community.kotlin.kompile.testrunner.jvm/tree/wip/F54F-evidence-20261008).
- Instrumented remote reproduction (lane F54G): three remote runs of the four-test bundle with timestamped compiler, launcher, hook, clock, acknowledgement and completion prints, all passing first attempt 4/4 (runs https://buildtest.kotlin.build/run?id=d93a25a5, https://buildtest.kotlin.build/run?id=66cdd957, https://buildtest.kotlin.build/run?id=c6f926ac). An equal slot count on the same droplet class does not recreate the original companion workload, so these passes do not refute the historical failures. All instrumentation was reverted; the merged PR contains none of it.
- Independent reviews of the PR (lanes R54d, R54e, R54f; evidence branches
wip/R54*-evidence-20261008) found no product defect that explains the retries; every assertion, budget and count in the PR was kept.
Relevant PRs / refs
| Repo |
Branch |
Remote head SHA |
PR |
What is on it |
State |
| CodexCoder21Organization/community.kotlin.kompile.testrunner.jvm |
fix/lnc1-timeout-process-tree |
d7ee13199be5224169bfb891ee5871fbbcd6a3da |
https://github.com/CodexCoder21Organization/community.kotlin.kompile.testrunner.jvm/pull/54 |
the timeout process-tree fix |
merged to main |
| CodexCoder21Organization/community.kotlin.kompile.testrunner.jvm |
wip/F54G-evidence-20261008 |
7e051a398e89cb110abed4a8b46266225e0041af |
no PR |
attempt-output comparison, instrumented remote runs, findings |
evidence only |
| CodexCoder21Organization/community.kotlin.kompile.testrunner.jvm |
wip/F54F-evidence-20261008 |
36034e7abf28ec0eaa2521345928e9a3f7c7f862 |
no PR |
local amplification attempts (0/10) |
evidence only |
Nothing is deployed or published from this effort; the testrunner is consumed as a Maven artifact and no new version was published for this PR (verified: no publish step ran in any lane).
Next steps
- Start from the first-attempt outputs in the F54G result file, not from a fresh local loop: the missing nested-hook PID publication at timeout completion is the one concrete anomaly; find the code path that publishes that PID and enumerate what can prevent it on a loaded two-vCPU runner (a hook that has not registered yet when the timeout fires; a child JVM that is still starting; an interrupt consumed before publication).
- Force that condition deterministically through the public API (a child that delays its hook registration, a timeout that fires during child startup) rather than by machine load; the reproducer must fail on the merged main near 100% and pass after the fix.
- Fix at the cause (product or test harness), keep every existing assertion, budget and count, open a PR with the why above, and run the four tests 10/10 locally plus one remote run.
- If the forced condition cannot be found from the outputs, add a durable per-attempt diagnostic (the attempt's own timestamps and PIDs in its failure text) so the next CI occurrence names the phase, and land that first.
Reusable knowledge
- buildtest's results API paginates (
/api/test-results?id=<run>&page=<n>&pageSize=100); read every page or retried tests are missed.
- Remote test output for passing attempts carries no console events; only superseded failed attempts keep output.