← Workstreams

Workstream: AiCliConversationSearch

Status: Planned · Component: Enable agentic swarms · Consoles: AiCliSupervisorManager · Engine: AiCliHostSupervisor

This is an intention and requirements document, not an implementation plan. It describes making every conversation that has ever run on the supervisor fleet visible and searchable from the fleet console at aiclisupervisor.wasmserver.com/conversations — including the conversations that exist only as harness-native session files (the .codex / .claude jsonl transcripts) with no live thread and no agent association. It is the near-term increment toward the history read models that the AiCliHostSupervisor redesign (§9) already envisions: "history can be looked up by query" via a rendered transcript and a SearchableBucket recall index.

North star

An operator (or, later, an agent using a recall tool) can open the fleet console's Conversations page and see the fleet's complete conversational memory: every conversation, on every supervisor, whether it is live right now, paused, or finished months ago — and can search that memory by content ("which conversation debugged the gossip storm?") rather than only by scrolling a list of titles. Nothing that a harness CLI wrote to disk is invisible to the console.

The problem today

The Conversations page renders only live conversations, from the manager's cached fleet snapshot of currently-running PTY sessions. But the machines the daemons manage carry a much larger record: every codex instance keeps its full conversation history as jsonl session files inside its .codex directory (and claude keeps the same under .claude), including conversations that predate the manager, conversations started manually outside any agent, and conversations whose PTY has long exited. Today that record is:

  • Invisible — a fleet with hundreds of on-disk conversations shows an empty Conversations page the moment nothing is live.
  • Unsearchable — the page's search box is a client-side filter over the rows already rendered; there is no way to find a conversation by what was said in it.
  • Only partially owned — the supervisor already reads these same session files for spend history (AiCliHostSupervisor spend-history decisions), but exposes nothing of their content.

Goals and requirements

  1. Completeness — every on-disk conversation appears. The unit of inventory is the harness-native session file, not the manager's thread record. A conversation with no agent, no live PTY, and no operator-given name is still listed, with whatever identity can be derived from its file (harness, session id, start/last-activity times, working directory, model, a title derived from its first user message). "No agent associated" is a first-class, honestly-labeled state — never a reason to drop a row.
  2. Fleet-wide, through the manager. The console renders the fleet's conversations from the durable manager service, aggregated across every registered supervisor — consistent with the standing principle that the manager aggregates; the console renders (AiCliSupervisorManager). A page load never fans out to every supervisor.
  3. Content search. The search box on /conversations graduates from a client-side row filter to a real fleet-wide search over conversation content (transcript text) as well as titles and metadata, with sensible filters (supervisor, harness, agent, date range) able to arrive incrementally. SearchableBucket — the org's Lucene-backed document store — is the intended search backend, fulfilling the recall-index role the AiCliHostSupervisor redesign assigns it; the workstream validates that fit before committing to it.
  4. Readable transcripts. A listed historical conversation can be opened and read — a rendered, human-readable transcript view (who said what, when, with tool activity legible) — distinct from the live PTY terminal, which remains the surface for live sessions. Reading history requires no running session.
  5. Read-only over the native record. The harness CLI's session store stays the canonical record, per the redesign's continuity principles (§9). This workstream only ever reads session files: it never rewrites, moves, or deletes them, and never disturbs a live session to index it. There is no platform-side transcript reconstruction or summarization tier.
  6. Bounded resources at every layer. Session stores hold thousands of files, individual transcripts run to megabytes, and this fleet has already been through outages caused by unbounded list marshaling. Listing is paginated and incremental; transcript content moves in bounded chunks; indexing is incremental (a grown file re-indexes cheaply, an unchanged file re-indexes not at all); memory use on daemon, manager, and search backend stays flat as history accumulates.
  7. Idempotent, freshness-bounded ingest. Conversations keep growing after they first appear; re-ingesting a conversation updates its single search document rather than duplicating it. New or newly-grown conversations become visible and searchable within minutes, not only on restart — and the staleness bound is stated, not implied.
  8. Graceful mixed-version fleet. Supervisor daemons upgrade one machine at a time. A daemon that predates the new capability degrades honestly: the fleet page still renders, that supervisor's historical record is reported as unavailable with the upgrade instruction (the same pattern the Agents page uses for daemons that predate agent management), and nothing errors or blanks the page.
  9. Association where derivable, honesty where not. When a historical conversation can be tied to a known agent or to a live thread (agent home directories, workspace paths, session ids), the row links to that agent/thread. When it cannot, the row stands alone. Derivation never invents an association, and the same conversation never appears twice because association failed.
  10. Live and historical are one list. The Conversations page stays a single surface: live conversations (clicking into the PTY terminal) and historical ones (clicking into the transcript view) share the list, visually distinguished by status, sorted and searched together. Operators should not need to know which storage tier a conversation lives in to find it.

Non-goals

  • Not the durable event-sourced transcript home. The AiCliHostSupervisor redesign makes the supervisor durably own transcript history in its event-sourced store; that remains the destination. This workstream reads the harness-native files that exist today, so the fleet's memory is usable now — it must not build a second durable transcript store on the manager that would compete with the supervisor's ownership. (The manager may durably hold its aggregated index — metadata and search documents — per its existing cached-snapshot role; it does not become the home of transcript truth.)
  • No context seeding or resume reconstruction. Search and reading serve humans (and later, recall tools). Rebuilding an AI CLI's working context stays the CLI's own job, per §9.
  • No lifecycle management of the native stores. Retention, cleanup, archival, and workspace reclamation of .codex/.claude directories are out of scope (the fleet already states retained-workspace counts without reclaiming them; that honesty pattern continues).
  • No new authentication/authorization model. The conversations surface inherits the console's existing operator-facing access model; per-conversation ACLs are not this workstream.

The shape of the system

The layering follows the standing architecture — each layer only widens an existing responsibility:

  Operator ──HTTPS──► AiCliSupervisorManagerWui  /conversations
                         │  (renders manager state; search box queries manager)
                         ▼  url://
              AiCliSupervisorManager (durable)
                - aggregated conversation inventory across the fleet   (widened: live-only → live + historical)
                - search index over titles + transcript content        (new; SearchableBucket-backed)
                         │  url://  (poll/ingest, bounded + incremental)
                         ▼
              AiCliHostSupervisorDaemon (per machine)
                - inventories its harness session stores (.codex, .claude)   (new API surface)
                - serves conversation metadata and transcript content, paginated/chunked
                - already scans these same files for spend history — one shared, incremental view of the stores

Requirements on each subsystem

  • AiCliHostSupervisorDaemon — grows a read-only historical-conversations capability: inventory the machine's harness session stores (all of them it manages — agent homes and any default/unassociated store), serve per-conversation metadata and bounded transcript content, incrementally (building on the incremental session-file index the spend-history work already established). Mixed-version tolerance is part of the contract: the capability's absence must be detectable.
  • AiCliSupervisorManager — widens its aggregation role: durably maintain the fleet-wide conversation inventory (live + historical, merged and deduplicated), keep it fresh within the stated bound, and answer content-search queries; owns the relationship with the search backend.
  • SearchableBucket — serves as the search index if validated (document-per-conversation, idempotent upsert under a stable id, bucket-scoped isolation, query semantics fit for "find the conversation where X was discussed"). If validation shows it unfit for transcript-scale indexing, the finding — and the chosen alternative — is recorded here before implementation proceeds.
  • AiCliSupervisorManagerWui — the Conversations page grows the unified live+historical list, server-backed search, and the read-only transcript view; empty states stay honest ("no conversations" only when the fleet truly has none; per-supervisor unavailability labeled inline).

Acceptance criteria

  • A fleet whose supervisors carry harness session files shows them all on /conversations — including conversations with no agent and no live session — with a page that stays responsive at thousands of rows.
  • Searching a distinctive phrase that occurs only inside one conversation's transcript finds that conversation, fleet-wide, from the page's search box.
  • Opening a historical conversation renders its transcript readably without any running session.
  • A supervisor running a pre-upgrade daemon is labeled unavailable-for-history on the page while everything else renders normally.
  • Ingest is demonstrably idempotent (a re-ingested conversation appears once) and demonstrably bounded (indexing a large store does not degrade daemon or manager memory), with tests exercising these properties end-to-end.

Relationship to other workstreams

  • AiCliHostSupervisor — the engine this reads from today and the eventual durable owner of transcript history; its §9 read-model vision (rendered transcript + SearchableBucket recall index) is what this workstream begins delivering, from the native files that already exist. The incremental session-file indexing built for spend history is shared machinery, not duplicated.
  • AiCliSupervisorManager — the console and durable aggregation layer this widens; its "fleet home" surface is where the result lands.
  • SearchableBucket — the candidate search backend; this workstream is its first transcript-scale consumer and will surface any capacity/contract gaps.
  • AiCliSpend family — the existing session-file parsers (spend/usage) prove these files are tractable; conversation inventory generalizes the same source of truth from "how much did it cost" to "what was said".

Graduation

This workstream graduates when the acceptance criteria above hold on the production fleet — the Conversations page is the fleet's complete, searchable memory — and the design decisions made along the way (daemon API shape, ingest cadence, search-backend verdict) are recorded in the relevant project documentation.