← Handoffs

Repository · handoffs

Handoff: Finish feed reset diagnostics and open the pull request

View on GitHub ↗

id: hf-2026-10-05-finish-feed-reset-diagnostics-and-open-the-pull-request url: url://handoff/handoffs/hf-2026-10-05-finish-feed-reset-diagnostics-and-open-the-pull-request title: Finish feed reset diagnostics and open the pull request summary: Run the named reset-log tests against the pushed implementation, fix any compile/assertion issues, then open the ready BuildTestEmbedded PR; branch fix/feed-reset-diagnostics at 07547693565d47f669eea2947a27b2c93f8ef678 contains the unverified implementation, and the six fail-first cases passed their red check against old source while remote runner attempts failed before tests. created: 2026-10-05T02:32:32.540Z completed: null blocked-reason: The 60-minute task window expired before any test ran against the changed source or a PR could be opened. dependencies:

Handoff: Finish feed reset diagnostics and open the pull request

Written 2026-10-05 02:31 UTC.

RE-VERIFY: This is a write-time snapshot. Before acting, run git fetch origin, check git ls-remote origin refs/heads/fix/feed-reset-diagnostics, and inspect the pull-request list for that head. Confirm the live BuildTest run state from its page/API before re-driving any remote tests.

Mission summary

The user asked for a small BuildTestEmbedded change so coordinator logs explain each projection-feed reset: authority changes need the old/new IDs, trigger, and full throwable chain; reset answers need the requesting cursor, category, and returned cursors. The public site was stale for more than 17 hours after resets on 2026-10-04, and investigations could not identify why authority changed or which reader was reset. The user's request also called for real on-disk public-API tests, targeted and full-suite verification, a ready PR with the specified first paragraph, then stopping without merge, enqueue, publish, or deploy.

What was found and done

  1. Read br/common.md, the local resettrigger findings, the linked diagnostic specification and timeline, BuildTestEmbedded README, and the linked Testing Architecture guidance. The existing reset callback was called from finally before an advertised snapshot page had completed; the recovery path printed a stack but did not record old/new authority IDs; unavailable-journal rebuild changed authority without an authority event. Ordinary compaction changes baseline generation and is not an authority replacement.
  2. Added structured reset diagnostics with unknown-authority, position-no-longer-retained, snapshot-replaced, and other, plus incoming and returned cursors. Reset-answer notifications are emitted after successful page completion and outside the journal lock. Added authority-replacement events for startup reopen failure, unavailable-journal rebuild, and advertised-reset verification recovery, with a stable JSON cause/cause/suppressed representation. No reset decision, timeout, limit, or retention value was intentionally changed. These source changes have not yet been compiled or tested in their final form.
  3. Added six public-service test scenarios using temporary on-disk journals: the four reset categories (unknown authority and other are both exercised in one test), startup replacement, unavailable-journal rebuild, and advertised-snapshot recovery. To prove fail-first behavior, ran all six tests locally against the pre-change source at main ca96938232d7048c8fd7db57615388bd689b7d14 while the implementation files were stashed. The result was TESTS FAILED (0/6 tests completed successfully, 6 failed); captured assertion output showed the old reset line missing the category and returned cursors. This is expected red evidence, not a green verification of the implementation.
  4. Two remote baseline attempts did not execute tests. Run 1a0c3154 failed during worker provisioning after digitalocean-droplets getDropletCreationStatus lost its RPC stream; its build log contains java.io.EOFException: Stream closed while reading message data (read 0 of 4 bytes). Run 41e22cf3 remained BUILDING until the local client hit its 30-minute limit; build-watchman then reported it CANCELED because the remote workspace closed. Neither run is a code-test result.
  5. The implementation and tests are committed and pushed. The worktree was clean at write time, and the local head matched the remote branch head. No PR exists and no code was deployed or published. The 60-minute timebox expired before running the tests against the changed source, reviewing the final diff, or opening the requested PR.

Relevant PRs / refs

Item Write-time state Link
fix/feed-reset-diagnostics, head 07547693565d47f669eea2947a27b2c93f8ef678 Pushed; current work; tests against this implementation remain unrun Branch, commit
Pull request None No PR URL exists yet
Deployment/publication None performed Not applicable

Next steps

  1. Re-verify current state first: fetch main, confirm the remote branch/head, ensure there is no existing PR for this head, and confirm no tracked test process remains.
  2. Run the six named tests against the pushed implementation. Use --local if BuildTest remote provisioning remains unhealthy; the help output confirms --test may be repeated. Fix any compile or assertion failures and rerun only the affected named tests.
  3. Review the full source/test diff, especially the three authority-replacement sites and that all logging callbacks run outside journal/persistence locks. Check git diff --check and remove ignored build artifacts before any commit.
  4. Resolve the suite-policy conflict explicitly in the PR/findings. br/common.md:49-50 says not to run the full local suite and names CI as the whole-suite gate; the user task asked for one full local suite. The previous run followed the lane rule. No whole-suite run has occurred.
  5. After the targeted tests pass, fetch/rebase on current origin/main, run the required named checks from that exact head, commit/push any fixes, and open a ready PR. Lead the PR body with the user's supplied outage rationale verbatim. Update the body using gh api -X PATCH repos/CodexCoder21Organization/BuildTestEmbedded/pulls/<id> -F body=@/abs/file.md, re-read it, and stop with no merge, enqueue, publish, or deploy.

Reusable / operational knowledge

  • scripts/test.bash --help supports multiple target names with repeated --test flags or a colon-separated list.
  • The failed remote runs above are infrastructure failures and must not be counted as assertions passing or failing against either source version. Do not keep re-running remote jobs blindly; diagnose the runner path first.
  • out/resetlog-findings.md in the session scratchpad contains the contemporaneous findings and process notes; it is outside the repository, so rely on this handoff and the pushed branch for durable state.
  • The report-challenge skill's CLI automatically creates and merges a PlanRepository PR; it was not invoked because the user explicitly prohibited merges and enqueueing for this task.