← Priority list
Blocked

Restore the buildtest WUI projection feed: backend getProjectionChanges blocks on startup reconciliation

UrlProtocol PR 589 (transport fix) is MERGED, but on fresh 2026-09-24 processes getProjectionChanges still times out: the backend parks in awaitProjectionChangeStartupReconciliation (BuildTestEmbeddedService.kt:8267) past the 30 s handler budget, so the WUI stays projectionFeedResetRequired=true. Blocked on the L5 projection startup/durable-checkpoint handoff; then deploy + re-verify.

Waiting on the active startup-compaction dependency; read-only health checks still show projectionFeedResetRequired=true, and completion then requires an owner-authorized production redeploy and feed verification.

Handoff document

Markdown

Restore the buildtest WUI projection feed (backend blocks getProjectionChanges on startup reconciliation)

RE-VERIFY: re-scoped 2026-09-24T04:50Z by sweep3-w4-20260924 (handoff-board sweep). Re-check https://buildtest.kotlin.build/api/health, /health/dashboard, and the backend log before acting.

Mission

Restore the live projection feed on https://buildtest.kotlin.build/ so /api/health reports projectionFeedConnected=true AND projectionFeedResetRequired=false, and the health-projection-feed-stale notice is hidden on /health/dashboard and on a /run?id=... page.

What changed since the 2026-09-13 snapshot

  • The transport fix this handoff was waiting on, https://github.com/CodexCoder21Organization/UrlProtocol/pull/589 ("Publish TCP connection close before handler interruption"), is MERGED (2026-09-13T22:20:14Z). That item is done.
  • Production has since been redeployed by other sessions: WUI route https:buildtest.kotlin.build:443 runs buildtest-wui-faf14af1-20260924.jar (restarted 04:33Z and 04:43Z on 2026-09-24); backend route url:buildtest: runs buildtest-server-hotfix-embedded-0.0.68290276-20260924.jar.

Observed on 2026-09-24 04:42-04:50Z (after PR 589 merged, on fresh processes)

  • /api/health: projectionFeedResetRequired=true on every sample; projectionFeedConnected flaps true/false; projectionUpdatedAt advances; unattributedRuns 791 -> 800. /health/dashboard renders id="health-projection-feed-stale" without hidden.
  • WUI log (last 1500 lines): 80x RPC request 'getProjectionChanges' failed: INTERNAL_ERROR - Service handler timed out after 30 seconds without producing a result, plus the same handler timeout for getTestAttemptResultsPaginated (99), getTestResultsPaginated (22), getBuildRun (9).
  • Backend log: the handler is not losing its reply in transport. It is parked here until the 30 s handler budget interrupts it:
    Error handling 'getProjectionChanges' (periodic full trace): java.lang.InterruptedException
      at AbstractQueuedSynchronizer$ConditionObject.await
      at buildtest.embedded.BuildTestEmbeddedService.awaitProjectionChangeStartupReconciliation(BuildTestEmbeddedService.kt:8267)
      at buildtest.embedded.BuildTestEmbeddedService.getProjectionChanges(BuildTestEmbeddedService.kt:7782)
    
    i.e. the projection change journal has no complete committed state after the backend restart, so every getProjectionChanges waits for startup reconciliation, which does not finish within the 30 s handler budget.

Mechanism (inferred from the above, consistent with BuildTestEmbedded main a58fe128)

awaitProjectionChangeStartupReconciliation (BuildTestEmbeddedService.kt ~8360) waits on projectionChangeMaintenanceChanged until projectionChangeStartupReconciliationComplete whenever journal.hasCompleteCommittedState is false. Startup reconciliation (synchronous full journal replay + eager run reads) takes longer than the RPC handler budget on production data, so the WUI never gets a reset baseline and stays resetRequired=true. This is the same defect the L5 projection handoff measures in tests ("feed began 18648 ms and waited for startup reconciliation"; "same-directory restart ... still read all 53,606,885 bytes") and plans to fix by persisting serving checkpoints across restart.

Next steps

  1. Blocked on url://handoff/handoffs/hf-2026-09-15-finish-l5-projection-startup-and-compaction-under-the-original-stress-deadlines (durable serving checkpoint resume / bounded startup work). Do not add a timeout, retry, or cap here - the fix is making startup serve from a durable checkpoint.
  2. After that fix is merged AND deployed to url://buildtest/ (production deploy requires explicit owner approval), re-check: one getProjectionChanges completes; /api/health shows projectionFeedResetRequired=false and projectionFeedConnected=true across several samples; the stale-feed notice is hidden on /health/dashboard and a /run?id= page.
  3. unattributedRuns is now "runs that could not be assigned to lifecycle states because legacy phase timestamps were unavailable" (HealthDashboardServlet.kt) - a permanent legacy count, not a backfill-progress metric; do not wait for it to reach zero.
  4. awaitingTeardownDroplets=-1 (still present) is a separate projection-arithmetic defect; track separately.

2026-09-24 dependency update

The dependency on url://handoff/handoffs/hf-2026-09-15-finish-l5-projection-startup-and-compaction-under-the-original-stress-deadlines (rank 56) was removed: that handoff was completed as superseded on 2026-09-24 because its lane's branch was abandoned. What this feed waits on is unchanged: startup reconciliation after a restart must finish (or serve from a durable checkpoint) fast enough that getProjectionChanges answers inside the 30 s handler budget. The startup-compaction fix is now owned by url://handoff/handoffs/hf-2026-08-24-fix-startup-compaction-delta-heap-exhaustion-then-revalidate-the-extreme-suite (rank 86), which is the new dependency. Read "Next steps" item 1 as pointing at that handoff; steps 2-4 are unchanged. (sweep3-finisher-20260924)

No status reports yet.

Add dependency

Complete this handoff

Moves it out of every priority list and into ArchiveArea.