← Workstreams

Probably we should rename https://github.com/CodexCoder21Organization/kotlin-build-ci to githubci.kotlin.build and it should be the github app for integrating with Github's CI including the WUI for managing your app installation/configuration. It should help users onboard, and should provide settings like who is allowed to use /bypass and should generally allow a multi-tenant users to setup with the CI.

These are the instructions that were provided to me by codex for configurating my githubci app. The instructions were perfect, they were accurate and precise, the URLs included the correct IDs as must have been discovered by querying the real github app/configuration or something. If we are ever configuring a different repository or configuring a different github instance or if the permissions are ever wrong, we want to be able to generate instructions like this:

  ## Exact correction steps

  1. Open the App’s Permissions & events settings
     (https://github.com/organizations/CodexCoder21Organization/settings/apps/kotlin-build-ci-test/permissions).

  2. Under Repository permissions, change Issues from Read-only to Read and write.
  3. Under Organization permissions, confirm Members is Read-only. The App registration already requests this, but the
     installation has not approved it.

  4. Under Subscribe to events, enable Pull request. Leave Check run, Check suite, Issue comment, and Push enabled.
  5. Click Save changes.
  6. Open installation 97022712’s configuration page
     (https://github.com/organizations/CodexCoder21Organization/settings/installations/97022712).

  7. Accept the pending permission request using Review request / Accept new permissions. The approved request must include:
      - Issues: read and write
      - Members: read-only

  8. Refresh the live audit endpoint
     (https://githubci.kotlin.build/health/configuration?repository=CodexCoder21Organization/kotlin-build-ci). Success is HTTP
     200 with "status":"healthy".

  The relevant PRs remain open:

  - https://github.com/CodexCoder21Organization/kotlin-build-ci/pull/190
  - https://github.com/CodexCoder21Organization/KotlinBuildHealthCheck/pull/10

Any multi-tenant user settings and github configuration settings should follow https://github.com/CodexCoder21Organization/DocumentationRepository/blob/main/architecture/STANDARD_LAYERED_ARCHITECTURE.md the only weird thing here is that we are combining the WUI with the github API endpoints because we want a single URL for both and because both are HTTPS anyway so it really just makes sense.

Backend service: url://githubci/

The HTTPS app at https://githubci.kotlin.build (the GitHub App webhook receiver, the WUI, and the GitHub-facing API endpoints combined on one URL, per above) should be a thin frontend in the sense of the standard layered architecture: it talks only url:// and holds no durable state on disk. All domain state and logic live behind a backend url://githubci/ service, with the standard project family around it (GithubCiApi / GithubCiEmbedded / GithubCiServiceServer / GithubCiCli; the WUI half is the existing app frontend itself).

Responsibilities of url://githubci/:

  • Incoming-request log. An append-only record of every request the frontend receives — webhook deliveries above all (delivery id, event type, repository/installation, received-at, ACK latency, processing outcome), but also /rerun·/bypass comment commands and WUI/API actions. This is our own first-hand ledger of what GitHub actually delivered, complementing GitHub's GET /app/hook/deliveries view — exactly the evidence the recurring lost-delivery / lost-check-suite-dispatch incident class has had to reconstruct by hand each time.
  • Multi-tenant settings, including user settings. Tenants (installations / orgs / repositories) and their configuration for the multi-tenant CI service: who is allowed to use /bypass, onboarding state, app installation/configuration audit expectations (the health/configuration checks above), and per-user preferences.
  • Periodic lost-build reconciliation. A recurring sweep (see the dedicated section below) that finds builds GitHub is expecting the CI service to handle but the service is not actually tracking or progressing, and repairs them.

Tenancy shape follows Single-Tenant and Multi-Tenant APIs. This is a multi-tenant first-party service, so it gets both data-plane tiers under the uniform naming rule (FooManager manages many Foos): a GithubCiTenantManager holding the cross-tenant, lifecycle-global operations — creating a tenant when an app installation onboards, listing/addressing tenants, ownership and permissions — which vends a single-tenant GithubCiTenant handle carrying the tenant-scoped operations (read/write that tenant's settings, query that tenant's request log). Per the "when to publish Foo separately" criteria, the single-tenant interface starts out internal to the Embedded and only the manager is published, until multiple backends or detached-handle consumers earn its extraction. Per-user settings authorization should eventually ride W3Wallet capabilities rather than bespoke ACLs.

Periodic lost-build reconciliation sweep

GitHub's view of a build and the CI service's view drift apart in practice, and today nothing notices until a human finds a PR wedged. The recurring causes:

  • A webhook delivery never arrived (transient GitHub or network failure), so the service never learned a check_suite was requested. GitHub shows the check as expected/queued forever; the service has no record at all.
  • The service's state was cleared or disrupted — maintenance, a redeploy, a run deletion — so a build GitHub believes is in_progress no longer exists on our side. GitHub shows a yellow "building" check that will never conclude.
  • A build was lost or canceled mid-flight (executor death, queue eviction) without the check run being concluded, leaving the same stuck-in_progress shape.

The backend url://githubci/ service should run a reconciliation sweep on a recurring schedule — roughly every 6 hours (configurable; per-tenant override is fine) — that finds these orphans and repairs them.

Discovery must NOT enumerate every repository individually. The sweep needs a bounded number of API calls regardless of how many repositories are installed. Two complementary bulk sources, in preference order:

  1. The GitHub App webhook delivery log as a catch-up journal. GET /app/hook/deliveries is GitHub's own record of every delivery it attempted for our App — one paginated, cross-repository feed. Diff it against our incoming-request log (above): any delivery GitHub attempted in the window that our ledger never recorded as successfully processed is a lost dispatch. Failed or missed deliveries can then be replayed with POST /app/hook/deliveries/{delivery_id}/attempts, which re-enters the normal webhook path rather than needing a bespoke repair path. This is exactly the evidence the recurring lost-delivery incident class has had to reconstruct by hand each time; the sweep automates that reconstruction.
  2. A bulk cross-repository search for stuck check states. The delivery log catches deliveries GitHub attempted; it cannot catch state lost after successful delivery (the maintenance/state-wipe case). For that, issue one paginated GraphQL search across the whole tenant — e.g. search(type: ISSUE, query: "org:<org> is:pr is:open") selecting each PR's head-commit statusCheckRollup / check runs — and filter to check runs owned by our App that are still QUEUED / IN_PROGRESS (or a check suite still merely requested) older than a staleness threshold comfortably above any legitimate build duration. One search query per tenant per sweep, not one query per repository. (is:pr is:open status:pending can pre-narrow the search where the rollup qualifier is precise enough.)

Only if neither bulk source is available for some tenant may the sweep fall back to per-repository enumeration, and that fallback should be loudly logged as the degraded path.

Reconciliation, once candidates are found: cross-check each candidate head SHA against the service's tracked builds and act on the disagreement, in whichever direction it points:

  • GitHub expects a build the service has no live record of → re-trigger it (replay the delivery, or re-request the check suite) so it flows through the normal build path.
  • GitHub shows in_progress but the service's record is terminal (completed/canceled/lost) → conclude the check run with the real outcome the service knows, or re-request the suite if no trustworthy outcome exists. Note the known trap that stale durable records can look live (the phantom-record class seen after buildtest run deletions, where kotlin-build-ci retained durable records for runs that no longer existed and rerequests silently reattached to them) — the sweep must trust verified live state, not the mere existence of a record.

The sweep must be idempotent and conservative: an age threshold keeps it from racing legitimately fresh builds, a repair already in flight is never issued twice, and every repair action is written to the incoming-request log with a reason (reconciliation: lost delivery <id> / reconciliation: stale in_progress since <t>) so a rule-repaired build can never be mistaken for a normally-dispatched one. All GitHub reads and writes the sweep performs — the deliveries feed, the search query, redeliveries, check-run conclusions — go through GithubProxy like every other outbound call (per the section below), which also centrally meters the sweep's rate-limit spend so a large tenant's sweep cannot starve interactive traffic.

CI build-request routing

Where a CI build request coming from GitHub (a push / check_suite:requested for a configured repository) gets routed should be a user-decidable setting, not hard-wired into the app. Routing rules are part of the multi-tenant settings held behind url://githubci/ and are edited through the tenant surface and the WUI, taking effect without a redeploy.

Rule resolution is by specificity — the most specific matching rule wins:

  1. Default — applies to everything configured with githubci.
  2. Organization rule — overrides the default for a particular organization.
  3. Project rule — overrides both for a particular repository/project.

Each rule chooses one of three dispositions:

  • Auto-pass — conclude the check green immediately; no real CI is invoked. (For repositories that opt out of CI — archives, docs-only projects.)
  • Auto-fail — conclude the check red immediately; no real CI is invoked. (A freeze/lockdown lever.)
  • Invoke a CI provider — dispatch the build to a particular service, e.g. url://buildtest/ directly or the hosted url://kompile-remote-build/ orchestration.

An auto-concluded check must say so in its output — "auto-passed by routing rule (no CI run)" — so a rule-produced conclusion can never be mistaken for a real build result (per Never Hide a Failure).

A primary everyday use of routing is keeping non-kompile repositories from submitting kompile builds at all. Not every repository the App is installed on is actually a kompile-ci repository: some are pure documentation, and some build with a different build system entirely. Today those repositories either get pointless build submissions (which can only waste build capacity or fail confusingly) or need ad-hoc opt-outs; with routing, a project rule (or an organization default plus per-project overrides) simply gives them a disposition that never dispatches to a kompile CI provider — typically Auto-pass, with its explicit "no CI run" annotation. Concrete examples at the time of this writing: Observable builds with Gradle rather than kompile, while DocumentationRepository and PlanRepository (and related repositories of the same kind) are documentation repositories with no code to build — all of these should be routed away from kompile CI by rule rather than special-cased in the app. (A repository on another build system could alternatively point an "Invoke a CI provider" rule at a provider for that build system, once one exists behind the common provider interface below.)

The CI provider is swappable behind a common provider interface. Rules name a provider rather than baking one in, so the underlying CI backend can be replaced per project, per organization, or globally. This generalizes today's deploy-wide CI_BUILD_ROUTE environment seam (kotlin-build-ci PR 201, whose production cutover is still a pending decision in KompileRemoteBuild) into per-tenant routing settings — a cutover becomes a settings edit, not an environment change and restart.

The point of the swappability is graceful migration. A routing rule should therefore support the parallel-run cutover shape directly: one authoritative provider whose check gates, plus optional informational shadow providers that run on the same requests and render their results without gating anything. Migrating a project (or the whole fleet) to a new CI backend then follows the documented sequence — run both in parallel, watch until they agree, flip which provider is authoritative, retire the old one — and every flip is a small reversible settings change with the previous provider still running.

Division of labor with KompileRemoteBuild: that workstream's url://kompile-remote-build/ stays GitHub-agnostic and owns build orchestration, reruns, and CI-decision events; url://githubci/ owns the GitHub-side state — the delivery/request ledger and the tenant/user settings. The App frontend is a thin composition over both, which preserves the recorded decision there that "GitHub I/O stays in the App" while removing the App's durable on-disk state.

Managing CI lanes from the WUI

Beyond routing, the WUI should let users manage a repository's CI lanes directly — where a "lane" is one named check that runs against a PR (a GitHub Actions workflow like a repository's own build workflow, or an App-owned check such as kotlin.build (remote)). Today changing any of this means hand-editing YAML in the repository and clicking through GitHub's branch-protection settings; it should instead be a self-service surface in the githubci WUI, backed by the tenant settings behind url://githubci/. Concretely, a user should be able to:

  • Enable and disable lanes. For GitHub Actions workflows this maps to GitHub's own workflow enable/disable endpoints; for App-owned check lanes it maps to the routing rules above (an Auto-pass rule is "this lane is off for this repository", with its explicit "no CI run" annotation). The WUI presents both as one uniform lane list per repository, whatever the mechanism underneath.
  • Add, remove, rename, and edit the CI config files. The workflow files under .github/workflows/ (and any other CI config files a lane is driven by) are viewable and editable in the WUI, with changes written back to the repository through the GitHub contents/commits API. Because these are ordinary files in the repository, WUI edits land as ordinary commits — on a branch with a PR when the target branch is protected — so the repository's own review rules and CI apply to CI-config changes exactly as to any other change. The WUI never bypasses branch protection to edit a config file.
  • Choose which lanes are required, and related settings. GitHub models this as required status checks in branch protection rules / repository rulesets; the WUI exposes it as a per-lane "required" toggle plus the closely related settings (strict up-to-date-with-base requirement, merge-queue participation, which branches the rule applies to). Marking a lane required or un-required is an outward-facing settings change, so it is loudly surfaced and recorded, never a silent side effect of another action.

Every one of these operations goes through the configured GitHub API proxy (GithubProxy, per the section below) — reading and writing workflow files, toggling workflows, and editing branch protection / rulesets are all outbound GitHub calls, so the same rules apply: the credential lives in the proxy, the proxy's chokepoint scopes what the WUI-driven surface may touch, and rate spend is metered centrally. Each mutation is also recorded in the incoming-request log with who requested it, so lane changes have the same audit trail as /bypass usage. Authorization for who may edit lanes and requiredness is part of the multi-tenant settings (eventually W3Wallet capabilities, like the rest of the per-user authorization above) — being able to un-require a gating check is at least as sensitive as /bypass and should be gated accordingly.

Outbound GitHub calls go through the GithubProxy

Whenever the githubci service needs to call out to GitHub — creating and concluding check runs, reading GET /app/hook/deliveries, running the health/configuration audits, posting PR comments — it should make those calls through the GitHub proxy api service (GithubProxy, url://githubproxy/), per the standing third-party service proxies rule: the GitHub credential is contained in the proxy rather than held ambiently by githubci, the proxy's chokepoint scopes what githubci may do, and GitHub's per-token rate budget is metered centrally across all consumers.

Direct local GitHub API calls — githubci authenticating with a configurable API key of its own — remain supported, but only as an emergency mode that must be explicitly enabled in the configuration, to facilitate graceful migration and fallback. The direct path keeps a working system available while the proxy route becomes (and remains) the default — covering both the migration period and a proxy outage without a redeploy — but engaging it is a deliberate, visible operator action: a configuration flip that is loudly surfaced while active and reverted once the emergency is over. It is never an automatic or inferred fallback — a proxy-path failure must surface to the operator rather than silently rerouting to the direct path, exactly the "unavailable must be verified, never inferred" discipline, since a silent reroute would hide a real defect in the proxy path (per Never Hide a Failure).