← Workstreams

Workstream: AiCliHostSupervisor

Status: Planned · Component: Enable agentic swarms

This document is an aspirational vision and design-scope document, not a milestone plan. It describes the end-state we are building toward and the architecture decisions that define it. Decomposition into concrete tasks and milestones happens later, in separate workstream notes. AiCliHostSupervisor is the durable home of the agent; its user-facing surfaces are two audience-specific consoles that sit over a fleet of supervisors — MyLittleAgi for end users and AiCliSupervisorManager for fleet admins.

North star

Make running an agent a first-class, supervised, capability-scoped activity — so that a user can define a persistent agent once and then have conversations with it on demand, where each conversation is a live AI-CLI session running safely on someone else's machine, wired up to exactly the tools, data, credentials, and peer agents that agent has been granted, and nothing else.

The defining split is the durable agent host vs. audience-specific consoles over it — the supervisor is where an agent lives, and the two frontends are just different lenses onto a fleet of supervisors:

  • AiCliHostSupervisor is the durable home of the agent. It is a daemon on a host machine that both supervises live sessions and durably owns everything that makes an agent a persistent entity: its definition (persona/system prompt, CLI choice, granted capabilities, allowed peers), the permissions/capabilities it holds, its runtime configuration (the concrete config files each CLI reads — .claude, .codex, MCP-server wiring, the mounted-tool set), its conversation catalog (every conversation it hosts, active and paused/archived, with resume state), its transcript history, and the inter-agent mailbox among the agents it hosts. An agent homes on exactly one supervisor; that supervisor is the source of truth for what the agent is, what it may do, and everything it has said. This state is durable and event-sourced (CQRS over the Event Log) and survives a restart. This is a deliberate evolution: the supervisor started as "stateless muscle," became stateful for config + catalog, and now owns the whole agent — it is both brain and muscle for the agents on its host, and is a pet, not cattle.
  • MyLittleAgi and AiCliSupervisorManager are consoles over supervisors, differentiated by audience — not stores of the agent. Both are frontends backed by a fleet of supervisors; neither owns an agent's definition, history, or mailbox (those live on the supervisor that homes the agent). MyLittleAgi is the end-user surface — a person just creates and talks to their agents and neither knows nor cares which supervisor runs them; the backing supervisors surface only deep in a settings page. AiCliSupervisorManager is the fleet-admin surface — the same supervisors underneath, but supervisor configuration, health, and placement are front and center. Each console holds only thin, audience-facing durable state of its own (which supervisors back it, human-friendly names/preferences, a cached aggregate), the same shape AiCliSupervisorManager already uses.

An agent is a persistent entity homed on one supervisor — its supervisor-owned definition (history, granted tools, granted data, granted credentials, an allow-list of peers it may talk to), runtime configuration, transcript history, and mailbox. A conversation is a PTY of that agent on that same supervisor: launched from the stored definition, kept in the supervisor's durable catalog across its whole active → paused/archived lifecycle, and paused or terminated on the user's request. These are two of the four canonical AiCli terms — supervisor (the privileged daemon on a machine managing many agents/models), model (a pretrained AI callable as a function; the choice that powers an agent), agent (a package of model + files/knowledge + permissions), and conversation (a thread of messages to/from an agent) — defined once on the AiCli project page. This document specifies the engine that makes that true.

Guiding principles

  • Privileged supervisor outside; guest and PTY inside. The supervisor is a privileged daemon that supervises from outside the container boundary: it requires Docker, launches guest containers, and never runs an AI CLI as its own child process. The guest (AiCliDockerGuestDaemon) lives inside the conversation's container: it directly owns the PTY, and it — not the supervisor host — is where the AI CLI binaries (claude, codex, gemini, opencode, …) are installed. The supervisor drives the guest over a url:// side channel (§2).
  • The supervisor is the agent's durable home, and restart-survivable. The supervisor durably owns the agents homed on it in full: their definitions, permissions, runtime configuration, conversation catalog (active and paused/archived), transcript history, and mailbox — event-sourced and surviving a restart. Restart-survivability means the supervisor's own durable store comes back intact; nothing about the agent depends on a separate brain. Because an agent's entire durable state lives on its home supervisor, that supervisor is a pet, not cattle: losing its durable store strands its agents until restored, and an agent is bound to the host that homes it (the trade-off, and the deferred cross-host migration/federation, are faced directly in §8). The consoles over the fleet (MyLittleAgi, AiCliSupervisorManager) hold no agent state — they read and drive the supervisors.
  • An agent is a W3Wallet principal, not an API key holder. An agent never holds raw secrets for its tools, data, and delegated credentials. Those are W3Wallet capabilities it has been delegated; the secret stays in the owner's daemon and the agent receives only results. Authority is possession of a capability, scoped and revocable, never ambient "it's running on our box." (The one pragmatic exception, until W3Wallet capability provisioning matures, is the agent's own model credential — see §5.)
  • Least authority by construction. A conversation is launched holding exactly the capabilities its agent definition grants — no host credentials, no network reach, no peer access beyond its allow-list. A tool the agent wasn't granted is not merely hidden; it is uninvokable because the agent holds no capability for it.
  • Tools are functions, mounted as CLI executables. An agent's tools are LambdaServer functions it is permitted to invoke, presented to the AI CLI as command-line executables mounted into the conversation's container (a Docker filesystem mount) — not via MCP, at least for now. The set of mounted tools is the grant: mounting an executable gives the agent an ability, omitting it removes it.
  • Inter-agent messaging is CQRS, not RPC. Agents communicate through an asynchronous mailbox: sending is a command (append an event), reading is a query (an Observable projection of the recipient's inbox). No agent ever blocks on another agent being alive. An offline agent simply has unread mail waiting.
  • Conversations are isolated. Each runs in its own container, driven through the in-container guest daemon, so a misbehaving agent cannot reach the host, the supervisor, or its siblings except through capabilities it holds.
  • One public API serves consoles and agents alike. The supervisor's url:// surface is the single observation/control plane: AiCliSupervisorManager, MyLittleAgi, and agents themselves — notably monitor agents (§12) — consume the same operations (attributed transcript reads and subscriptions, atomic message send, start/continue, catalog). The sufficiency bar for the API is that a console and an agent can each build on it with nothing private.
  • Initiation is console policy, not an engine restriction. The engine supports human-opened and trigger-woken conversations alike (§8); it must, because it powers SpawningPool's autonomy as well as MyLittleAgi's deliberately human-initiated-only experience. Each console states its own policy; the Manager, as the direct admin surface, exposes everything the engine supports.
  • Continuity is the AI CLI's own machinery. Session resume and context compaction are things every supported AI CLI already does natively; the platform's job is to never lose the CLI's own state (a durable home mount) and to relaunch with the CLI's native resume invocation — not to rebuild resume or summarization as platform components (§9).

Current state — what exists today vs. what is aspirational

The AiCli family is built and deployed end-to-end as described on its Documentation Repository project page, and the fleet console runs in production at aiclisupervisor.wasmserver.com — see its operational page in ServiceAtlas. Naming what already exists keeps this plan honest about the remaining work.

Already implemented (the foundation to build on):

  • PTY emulation and remote control. aiCliPtyApi defines the AiCliPty contract (start, sendString/sendEnter/sendCtrl, renderScreen, resize, isAlive, close, AiCliExecutionStatus BUSY/IDLE); AiCliPtyEmbedded implements it over JediTerm + pty4j, and harness-specific implementations exist for each supported CLI (ClaudePty, CodexPty, GeminiPty, OpenCodePty, AntigravityPty), shipped as the aiCliPty artifact.
  • A multi-session supervisor daemon. AiCliHostSupervisorDaemonApi / Embedded / ServiceServer manage multiple concurrent PTY sessions by session id (create, list, control, shutdown, health/status) and expose them over a single url://...aiclisupervisor/ endpoint via SJVM client bytecode, in both ContainerNursery lazy-start and standalone P2P modes. The daemon is deliberately PTY-implementation-agnostic and not runnable alone: the stock service-server refuses to create sessions until composed with a PTY factory through its public runServiceServer seam.
  • A host-local launcher — the architecture this workstream inverts. AiCliHostSupervisorLauncher is the daemon's runnable main entrypoint. Today it detects AI CLIs (agy, claude, codex, gemini, opencode) on the supervisor host's own PATH and composes the daemon with PTY implementations that run them as direct child processes of the supervisor JVM. That is exactly backwards for a privileged supervisor — the CLIs belong inside guest containers, not on the supervisor's host — and §2 replaces this composition while keeping the launcher as the entrypoint.
  • A PTY-proxy seed for the guest. AiCliDockerGuestDaemonEmbedded is today a transparent PTY proxy between a parent terminal and one child process: it spawns the child under a real PTY, forwards bytes bidirectionally (with window-resize support), and its tests already drive the real claude CLI end-to-end against a fake Anthropic server. It is the seed of the in-container guest daemon — but it is not yet a daemon: it has no container packaging, no lifecycle beyond proxying one child, and no control channel by which a supervisor could drive it from outside.
  • Usage/quota reporting at every scope. AiCliPty.quotaStatus() reports the backing CLI's spend/quota metrics (each a max quantity + current quantity — tokens, requests, or a percentage — with window and reset info; e.g. Codex's 5-hour and weekly limits parsed live from /status), and the daemon's getUsageJson aggregates those metrics at supervisor / model / agent / conversation scope for the fleet console to render. (Planned extension, 2026-08-01:) beyond live quota, the supervisor also computes historical token consumption and its dollar cost, adjusted per API pricing: it scans the past conversations in the harness CLIs' local session directories and computes per-conversation token usage and spend as ClaudeSpendApi defines it — that API's scope expands from Claude-only to cover codex sessions too. Per-model pricing is fetched dynamically rather than hard-coded, so price adjustments and newly released models are picked up without a code change. The fleet console renders this as the per-supervisor "Token Consumption" chart (AiCliSupervisorManager).
    • Codex session parsing (decided 2026-08-01). Claude sessions come from ~/.claude/projects/*/*.jsonl as today; codex sessions come from ~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl — one file per conversation, keyed by the embedded session id — reading each turn's timestamped token_count event joined with the prevailing turn_context model. Codex counts are normalized into the same four disjoint token buckets the Claude parser produces: uncached input = input_tokens − cached_input_tokens, cache-read = cached_input_tokens, cache-write = cache_write_input_tokens, and output = output_tokens unchanged (reasoning_output_tokens is a subset of output and is priced at the output rate, not separately). The parser tolerates schema drift across CLI versions (each file records its cli_version): unrecognized events are skipped, and unparseable files are surfaced as a count rather than silently dropped.
    • Spend library shape (decided 2026-08-01). The library generalizes rather than growing under the Claude name: ClaudeSpendApi / ClaudeSpendEmbedded / ClaudeSpendCli are renamed AiCliSpendApi / AiCliSpendEmbedded / AiCliSpendCli (GitHub redirects preserve existing links), with coordinates aicli.spend.api:aicli-spend-api and package aicli.spend.api — a one-time cost paid now, while the family has no consumers outside itself. The harness-neutral core is exactly what the API already defines (the four disjoint token buckets, ConversationCost with per-model breakdown, the service methods); harness specifics live behind a SessionParser interface in the Embedded layer with claude and codex implementations, each owning its session-root discovery and file format, keeping the Api layer logic-free per the standard layered architecture. The static ModelFamily/ModelPricing tables stop being the primary pricing mechanism and become the bundled static fallback of the pricing chain below, gaining OpenAI entries. (A parallel CodexSpend* family was rejected: it would duplicate the cost model and force the supervisor to stitch two APIs together for one chart.)
    • Pricing source (decided 2026-08-01). The spend library exposes a ModelPricingSource interface (per-model dollar rates for the four token buckets), and the real implementation fetches OpenRouter's public models endpoint — which carries per-token dollar rates including cache-read/cache-write buckets and long-context overrides for both Anthropic and OpenAI models — normalizes the raw model ids recorded in session files to OpenRouter slugs, and caches the table durably on the supervisor with a ~24h refresh plus an eager refresh whenever an unknown model id is encountered. The fallback chain is live fetch → last-good cached table → a bundled static table, so pricing degrades to slightly-stale rather than failing offline. A bucket with no listed price counts as zero (e.g. OpenAI cache writes are free); a model unknown even after refresh still has its tokens counted but its spend flagged unpriced rather than silently charged at $0.
    • Spend-history endpoint (decided 2026-08-01). The spend series is exposed through a new supervisor RPC, not an extension of getUsageJson — that method's consumers expect live quota, not a 30-day retroactive series. getSpendHistoryJson(sinceDay) returns daily UTC buckets — per model, the four token-bucket counts and the dollar cost — plus the unpriced-model and unparseable-file counts, so data-quality problems surface in the payload rather than vanishing. The supervisor keeps an incremental index over the append-only session files (tracking file offsets/mtimes): one full scan ever, cheap calls thereafter, even with thousands of accumulated sessions. The Manager polls this at a slow cadence and durably stores the merged series (AiCliSupervisorManager).
  • Frontends and the fleet manager. AiCliGui (Compose Desktop/TUI) drives sessions through the supervisor, and the single-supervisor diagnostic web console has been elevated into AiCliSupervisorManager (graduated 2026-07-05) — a durable url:// manager service plus stateless fleet console giving an operator one pane over every supervisor and every live conversation. These are operator/diagnostic surfaces over raw supervisors, distinct from the end-user agent products MyLittleAgi and SpawningPool.

So today the supervisor can already launch, render, and remote-control raw PTY sessions running an AI CLI, over url:// — but it runs those CLIs on its own host, with no container isolation, and holds nothing durable. What it cannot yet do is run conversations inside supervised guest containers, or be the durable home of a full agent — a defined, capability-scoped, history- and mailbox-carrying entity that persists on the supervisor across conversations.

Aspirational / not yet built (the substance of this workstream):

  • The supervisor–guest inversion. The supervisor does not yet launch guest containers at all, and the guest is not yet a daemon: there is no per-harness guest image, no url:// side channel, no remote AiCliPty proxying, and no launch handshake. This is the structural change everything else in this document stands on (§2).
  • Launch-from-stored-definition. The supervisor takes raw "run this command" requests, not "run agent X's next conversation." There is no notion of a stored agent definition it launches from, and no workspace/home continuity across conversations.
  • Capability-scoped sessions. Sessions today run with whatever ambient authority the host gives them. There is no per-conversation W3Wallet agent principal, no delegated tool/data/credential capabilities, and no enforcement that a conversation can only do what its agent was granted.
  • Tools as mounted CLI executables. There is no mechanism to mount a granted set of LambdaServer-backed CLI tools (plus built-ins like SearchableBucket and the mailbox) into a conversation's container, scoped per conversation, each performing a capability-gated invocation.
  • The inter-agent mailbox. Agents cannot message each other. There is no message event stream, no inbox projection, and no permission model for "may talk to all / a subset of peers."
  • Monitor agents. No agent can watch another: there is no attributed transcript-event subscription, no atomic message-send operation, no monitor attachments or trigger-matching watch loop, no trigger-wake, and no circuit breaker — the whole §12 surface is unbuilt.
  • The supervisor as the agent's full durable home. Today the supervisor holds only in-memory session state keyed by session id. It does not durably store agent definitions, permissions, runtime configuration (.claude/.codex, MCP servers, mounted tools), a conversation catalog (active + paused/archived), transcript history, or the mailbox — the whole notion of a persistent, capability-scoped agent living on the supervisor is unbuilt. Making the supervisor the durable owner of all of this — with the two frontends as thin consoles over it — is the change this workstream now centers on.
  • Placement. There is no scheduler that picks a supervisor host for a new conversation; the Manager's fleet registry is the natural substrate but nothing reads it for placement yet.
  • Per-agent network isolation. A container today gets whatever network the host gives it; the default-deny egress policy of §10 is unenforced (deliberately deferred).

The design

1. The agent / conversation / supervisor model

   ┌──────────────── Consoles over the fleet — no agent state of their own ──────────────────────┐
   │  MyLittleAgi (end-user: just talk to your agents; supervisors hidden except deep in settings) │
   │  AiCliSupervisorManager (fleet-admin: supervisor config, health, placement front and center)  │
   └───────┬───────────────────────────────────────────────────────────────────▲───────────────┘
           │  (1) create agent · start/continue a conversation                  │  (5) read for display:
           │      for agent X — a command to X's home supervisor                │      transcript · catalog ·
           │                                                                    │      usage · mailbox views
           ▼                                                                    │      (over the same url://)
   ┌────────────── AiCliHostSupervisor — the agent's durable home (per host) ───┴───────────────┐
   │  DURABLE (event-sourced): agent definitions + permissions · runtime config (.claude/.codex,  │
   │  MCP, tools) · conversation catalog (active + paused/archived) · transcript history · mailbox │
   │                                                                                              │
   │   ┌──────────── conversation = container + PTY ────────────┐   ┌──────────── … ───────────┐ │
   │   │  AiCliDockerGuestDaemon  →  AI CLI (Claude Code/Codex)  │   │  another agent's session  │ │
   │   │        │ runs as shell commands                        │   └───────────────────────────┘ │
   │   │        ▼                                                │                                  │
   │   │  mounted CLI tools (granted set only, via docker mount)│   (2) supervisor+guest negotiate  │
   │   │        │ invoke AS the agent's negotiated W3Wallet id  │       a W3Wallet identity for the  │
   │   │        ├─► LambdaServer functions (capability-gated)   │       agent; auto-revoked on exit  │
   │   │        ├─► SearchableBucket (recall)                   │                                   │
   │   │        ├─► SimpleFileSystem (workspace)                │   (3) the supervisor delegates    │
   │   │        └─► Mailbox: send=command, inbox=projection ────┼──►    the granted caps to that id   │
   │   └────────────────────────────────────────────────────────┘       (owner's wallet is root)   │
   └──────────────────────────────────────────────────────────────────────────────────────────────┘
                            (4) paused/terminated on user request — the agent and all its state stay on its home host
  • Agent — a durable entity owned and stored by its home supervisor in full: which AI CLI to run, its persona/system prompt, the set of capabilities it holds (tools, data buckets, credentials, peer-messaging rights), its runtime configuration (the CLI's config files, MCP wiring, mounted-tool set), its transcript history, and its mailbox. Persistent across conversations; it exists whether or not any conversation is currently running. Each agent homes on exactly one supervisor.
  • Conversation — a PTY session of that agent on its home supervisor. It is launched from the stored definition and stays live until the user pauses or terminates it (§8); the supervisor keeps it in its durable conversation catalog through the whole active → paused/archived lifecycle, so the host that ran it remembers it (and its resume state) afterward. The transcript is committed to the supervisor's own durable history.
  • Supervisor — a stateful, privileged daemon on a host that is the durable home of the agents assigned to it and launches, multiplexes, and supervises their conversations — each in its own guest container — exposing everything (live PTYs, catalog, history, mailbox) over one url:// endpoint. It durably owns each hosted agent's definition, permissions, runtime config, conversation catalog, history, and mailbox; restart-survivable via that durable store; an agent is bound to the supervisor that homes it.

2. The supervisor–guest split: privileged outside, PTY inside

The engine's structural rule: the supervisor is privileged and lives outside the container; the guest owns the PTY and lives inside it. The supervisor requires Docker. It never runs an AI CLI as its own child process, and no AI CLI needs to be installed on a supervisor host — the CLIs live only inside guest images.

  • The guest daemon owns the terminal. AiCliDockerGuestDaemon runs as the container's entrypoint. It spawns the AI CLI under a real PTY and embeds the harness-specific AiCliPty implementations, so all terminal emulation and harness intelligence — screen rendering, BUSY/IDLE detection, quota parsing, launch-flag/config construction — is colocated with the binary it drives. It keeps its transparent stdio-proxy behavior: docker run -it on a guest image behaves exactly like the wrapped CLI, and the container's logs show the CLI's verbatim output, so every conversation container is debuggable with nothing but Docker.
  • The side channel is url://. In addition to stdio transparency, the guest exposes two control surfaces to the supervisor over a UrlResolver connection: the full AiCliPty contract (start, send, renderScreen, resize, liveness/status, quota) as the primary drive surface, and a subscribable raw PTY byte stream for remote mirroring and debugging. The supervisor-side counterpart is a remote AiCliPty proxy that forwards the contract over this channel.
  • The guest dials the supervisor. At container launch the supervisor injects a bootstrap environment: its own dialable address including its peer identity, the conversation id, and a single-use launch token. The guest mints an ephemeral keypair (its transport identity), makes an identity-pinned dial to the supervisor — outbound-only, so it works on any Docker network with no inbound ports into the container — and presents the conversation id and token over the authenticated connection. Because the connection is bidirectional and long-lived, the supervisor then drives the guest's control surfaces back over the same connection. (This dial-out-then-drive-back shape is exactly the direct-connection pattern the UrlResolver plan prescribes for non-public peers; the env-var bootstrap mirrors NetLab's injected bootstrap convention.)
  • Per-harness guest images. There is one guest image per supported harness — aicli-guest-claude, aicli-guest-codex, aicli-guest-gemini, … — each bundling a pinned version of that CLI plus the guest daemon, with the image definitions owned by the guest repository. The supervisor's harness catalog is the set of guest images it can provide (built on demand from the shipped definitions, or pulled at a pinned tag), replacing today's host-PATH lookup. The agent definition's CLI choice selects the image.
  • Repository roles. The daemon repos (Api/Embedded/ServiceServer) stay PTY-implementation-agnostic and not runnable alone. AiCliHostSupervisorLauncher remains the supervisor's main entrypoint and composition point: CLI flags, the configured per-harness guest images, and the docker-guest PTY factory that launches containers and hands the daemon remote AiCliPty proxies. The composition seam stays fully general — composing any AiCliPty implementation (in-process local PTYs, fakes) remains supported and is the intended shape for non-Docker unit tests — but the docker-guest factory is the production composition, and the launcher's current host-PATH composition is retired with it.
  • The durable home is supervisor-owned agent runtime configuration. The container's home directory is a durable per-agent SimpleFileSystem mount that the supervisor owns and maintains on its host — this is the concrete embodiment of the supervisor's stateful agent runtime configuration. It continuously carries the CLI's own config files (.claude, .codex, MCP-server definitions, the mounted-tool wiring), its credentials, and its native session store (§5, §9). The container itself stays disposable; the supervisor-held home is what makes pause, crash, and resume equivalent — and, because it lives durably on this supervisor, it is part of what binds the agent to this host (§8). The supervisor materializes this configuration from the agent's own stored definition when the agent is first created on it, and is its durable owner throughout.

3. Launch-from-definition: the rehydration bundle

The single most important new operation is "start (or continue) a conversation for agent X." A console — MyLittleAgi or AiCliSupervisorManager — issues this as a command to agent X's home supervisor, the one that already durably stores X's definition, permissions, runtime configuration, history, and mailbox. The supervisor then assembles a self-contained rehydration bundle for the guest container from its own stored agent state — nothing durable is handed in from the console, which knows only which agent to run, not what the agent is. The bundle the supervisor prepares for the container contains:

  1. Agent identity & runtime — which AI CLI to launch (selecting the per-harness guest image) and how, read from the stored definition. The agent's W3Wallet identity is negotiated between supervisor and guest at launch (not a stored secret); the supervisor knows which agent this is, so it can delegate the agent's granted capabilities to that negotiated identity (§5).
  2. Resume reference — when continuing a paused conversation, the CLI's native session reference read straight from the supervisor's own durable conversation catalog (the console names the conversation; the supervisor already holds its resume state). The relaunched CLI restores its own context from its own session store on the durable home mount (§9). A fresh conversation omits this.
  3. Tool manifest — the set of granted LambdaServer-backed CLI tools and built-in tools to mount into the container (§4), each paired with the capability handle that authorizes it — all read from the stored definition.
  4. Granted capability set — the set of W3Wallet capabilities (tools, data, credentials, peer-messaging) the supervisor will delegate to the negotiated agent identity (on behalf of the owner's wallet) once the supervisor↔guest handshake completes (§5). These resolve to references/proxies, never raw secrets.
  5. Workspace handle — the SimpleFileSystem mounts for the agent: its working files and its durable home (§2).
  6. Inbox cursor — where in the agent's own inter-agent message stream it last read (§6), so its inbox projection materializes correctly.

As the conversation runs, the supervisor commits durable history to its own event-sourced store continuously: transcript events (what was said/done), any outbound mailbox messages, and workspace deltas. Because the supervisor commits these as they arrive, a crash mid-conversation loses at most the in-flight, uncommitted tail — never committed history. The consoles subscribe to this state for display (rendered transcript, live status, usage) but hold none of it: everything durable about the agent — definition, config, catalog, history, mailbox — lives in the one place, the home supervisor.

4. Tools: granted LambdaServer functions, mounted as CLI executables

An agent's tools are functions stored in LambdaServer that the agent is permitted to execute — plus a small set of well-known built-ins. (We are not using MCP, at least for now.) Instead, each granted tool is presented to the AI CLI as a command-line executable mounted into the conversation's container via a Docker filesystem mount — e.g. onto a tools directory on the agent's PATH. The AI CLI invokes a tool the same way it runs any shell command, and the tool does the work:

  • The set of mounted tools is the grant. At launch the supervisor mounts in exactly the executables for this agent's granted toolset — no more. A tool the agent wasn't granted is not merely hidden, it is not present on the filesystem, so it is uninvokable. Adding a tool is mounting its executable; revoking it is not mounting it — this is how a user "gives an agent a new ability."
  • Each tool invokes as the agent's W3Wallet identity. When the AI CLI runs a mounted tool, the tool reaches out over url:// to do its job — invoke the granted LambdaServer function, run a SearchableBucket query, perform a SimpleFileSystem operation, or post a mailbox message — exercising the W3Wallet capabilities the supervisor delegated (on behalf of the owner's wallet) to this agent's negotiated identity (§5). A tool that needs a secret API key invokes a Proxy/reference capability whose secret runs inside the owner's daemon, returning only the result. The agent can run the tool; it never sees the key, and it can only invoke capabilities the agent was actually granted.
  • Built-in tools every agent gets (subject to grant), mounted the same way: SearchableBucket for searchable context/recall, SimpleFileSystem for a working directory, and the mailbox send/read tools (§6). SearchableBucket is the natural substrate for "the agent's context associated with its tasks": the user (or the agent itself) stashes relevant documents into the agent's buckets and the agent recalls them by query rather than by key.
  • MCP is a possible future surface, not a current dependency. The mounted-CLI-tool model needs nothing beyond a filesystem mount and url:// reachability; if an MCP surface is wanted later, it can be layered on without changing the grant model (the mounted set still defines what is available).

5. Identity, credentials, and authority — the agent as a W3Wallet principal

Each conversation runs as a W3Wallet agent principal — precisely the "agent (native process, instance-scoped, auto-revoked on exit)" principal kind that the W3Wallet workstream calls out — and is the worked example of W3Wallet's general capability provisioning into hosted workloads.

The supervisor and the guest negotiate the agent's W3Wallet identity at launch, and the handshake rides the side channel of §2 — no separate protocol is invented:

  1. The guest mints an ephemeral keypair for this conversation (the private part never leaves the guest container — it lives only in the guest daemon's memory, written to no mount and no image layer) and makes the identity-pinned dial to the supervisor. The transport's mutual cryptographic handshake is the proof of possession, on both sides.
  2. Over that authenticated connection the guest presents the conversation id and the single-use launch token from its bootstrap environment; the supervisor validates and spends the token, binding this peer identity to this conversation.
  3. The supervisor delegates the agent's granted capabilities to that negotiated identity, on behalf of the owner's W3Wallet. Because the supervisor durably holds the agent's definition and grants, it is the natural broker: the owner's wallet is the root of authority, and the supervisor requests delegation of exactly the capabilities the definition grants. The grant is instance-scoped and its leases are short — per the W3Wallet model, lifecycle is enforced by renewal, not memory: the supervisor renews only while the container lives, so a reaped conversation's authority dies within one lease TTL even if explicit revocation is missed.
  4. The guest now holds an identity whose granted capabilities it can exercise; mounted tools invoke url:// targets as that identity, so each tool automatically gets exactly the W3Wallet capabilities the supervisor delegated to the agent — no more, and nothing a separate per-tool secret could leak.

Whether the W3Wallet principal is the guest's transport keypair or a second keypair bound over the authenticated channel is deliberately deferred to the W3Wallet workstream, whose hosted-workload provisioning design owns that choice; this engine works identically either way.

Because the unit of authority is the agent's negotiated identity plus the capability set granted to it, there is no finer-grained per-tool credential to protect or bypass: the agent is meant to be able to invoke any capability it was granted, and tools are simply how it does so. The implications:

  • No raw secrets in the container for tools, data, and delegated credentials. These are delegated capabilities the agent holds. A credential the agent "has access to" is a Proxy or reference capability: the secret lives in the owner's daemon, the agent invokes the capability, the daemon runs the secret and returns the result. The agent never sees the key, and the owner sees every use and can revoke it.
  • The model credential is the pragmatic exception, held by the agent itself. The agent's own AI-CLI account is configured via the CLI's native OAuth login, inside the container, and the CLI connects directly to its LLM provider — exactly as the CLI does on a developer's machine. The credential and the CLI's session store live on the durable per-agent home mount (§2), so login happens once per agent (drivable through the conversation's own PTY — a login flow is just terminal interaction) and survives every pause and relaunch. Delivering the model key as a W3Wallet Proxy capability instead — so the agent drives the model without ever holding its credential — remains the aspirational refinement, adopted when W3Wallet's capability provisioning matures.
  • Delegation by reference + attenuation. The supervisor (acting for the owner's wallet, from the agent definition it holds) delegates exactly what the definition grants and no more — to the negotiated identity — and may attenuate at the grant (which functions, pinned parameters, capped quota, shortened expiry). Authority strictly diminishes — an agent can never amplify what it was given, and can only re-delegate (e.g. to a peer agent in the same roster) a narrower slice.
  • Instance identity, auto-revoked on exit. The principal is scoped to the specific running conversation; when the container exits (or is reaped), its authority is automatically revoked — by lease non-renewal at the latest. A terminated conversation cannot act.
  • Quota as a spend-rate allocation. An agent's budget is a W3Wallet accrual-rate token bucket (the existing spend-key / spend-limit / spend-registration primitive) — a conversation is allocated resources at a rate, e.g. $3/hour, enforced at the owner's daemon on each capability invocation. A runaway agent drains its capped, rate-limited bucket rather than the host or the owner's wallet. The bucket travels with the delegation (§3/§7), so a delegated conversation inherits a budget the owner controls.

This is the same model W3Wallet already describes for agents; this workstream is one of its first real consumers, so the two plans must stay in lock-step.

Budget exhaustion mid-conversation — wait, then ask. Because the budget is an accrual-rate bucket, running low is usually not a failure, it is a pause:

  1. Wait for accrual (preferred). When the bucket can't cover the next step but is still accruing, the conversation simply pauses and waits until enough credits have accrued to continue — the live session stays warm (§9) so it resumes seamlessly. A slow agent throttles to its allocated rate rather than erroring out.
  2. Ask the user via a W3Wallet funds dialog (when accrual won't suffice). When there are not — and will not soon be — enough credits to make meaningful progress (the accrual rate is zero or too slow for the work), the service asks W3Wallet to surface a funds dialog for this service, using W3Wallet's lazy/dynamic permission-request UX. The user decides how to continue: top up a specific amount, raise the spend-rate allocation, or stop. The service may suggest a rate (e.g. "$3/hour") or request a specific amount to finish the task. Approving tops up / re-rates the bucket and the paused session continues; declining leaves the conversation paused (durable state already flushed, resumable later) or ends it.

The enforcement wiring this depends on (connecting W3Wallet's spend-bucket primitive to capability invocation, the service-initiated funds dialog, and the tuning of pause thresholds and suggested rates) is W3Wallet's design surface — deferred to its workstream.

6. Inter-agent messaging: async mailbox over CQRS

Agents talk to each other through an asynchronous mailbox, modeled directly on CQRS: messages are written like commands and read like observable projections. There is no synchronous agent-to-agent RPC. Messaging is scoped to an owner's own agents that share a home supervisor — an agent can only message another agent the same owner owns and (for now) one homed on the same supervisor, so the mailbox event stream and inbox projections live in that supervisor's own durable store, right alongside the agents themselves. There is no cross-user (cross-wallet) messaging. Messaging across supervisors — when an owner's roster is spread over several hosts — would require federating these per-supervisor mailboxes and is a deferred decision (below); until then, agents that must talk are homed together.

  • Sending is a command. When an agent sends a message to a peer, the mailbox send tool appends a message_sent event to the inter-agent message Event Log — {from, to, conversationContext, body, timestamp}. The send is authorized by a "may message peer X" capability (§7); it never reaches into the recipient's process. It is fire-and-forget from the sender's perspective.
  • Reading is a query (a projection). Each agent's inbox is an Observable projection over the message event stream, filtered to messages addressed to it, materialized from the agent's last-read inbox cursor. When a conversation runs, its inbox projection is surfaced to the agent — as injected context and/or a pollable mailbox read tool (§4) — and advances reactively as new message_sent events arrive. Reading is non-destructive; the event log remains the source of truth and the inbox can be rebuilt by replay.
  • Liveness is decoupled. Because sending only appends an event, the recipient need not be alive. An offline agent simply accumulates unread mail in its projection; the next time the user opens a conversation with that agent (conversations are human-initiated — §8), it sees the backlog. This is the whole point of choosing a mailbox over RPC: no agent ever blocks on another agent's availability, and no message is ever lost because a recipient happened not to be running.
  • Delivery is into the next conversation, not a push. A message is addressed to an agent; the home supervisor holds it in that agent's inbox projection (in its own durable store) until a conversation with the agent runs, then surfaces the backlog (and any mail that arrives mid-conversation) to the live session. Because sender and recipient share a home supervisor, the whole send→project→deliver loop is local to that supervisor; a console merely displays the inbox the supervisor projects. (The mailbox itself never wakes anyone: mail waits for the next conversation, however that conversation is initiated — §8. Monitor agents (§12) are the deliberate contrast — they push into a live session via the supervisor's public API and do not ride the mailbox at all, which is also why cross-supervisor monitoring needs no mailbox federation. Mail-driven autonomous waking as a product is SpawningPool's surface.)

The finer delivery semantics — strict ordering of the inbox projection, and whether backlog re-display after an un-advanced read cursor is acceptable (at-least-once) or must be exactly-once — are this workstream's design surface now that the mailbox lives on the supervisor; they are noted as a deferred detail rather than owned by a console.

7. Permission to talk: all peers vs. a subset

"An agent may be granted to talk with all agents or only with a subset of agents" — meaning all or a subset of the owner's own agents homed on the same supervisor (§6) — maps cleanly onto W3Wallet capabilities held in the agent definition:

  • A subset — the agent holds a distinct "may message peer X" capability for each permitted peer (or a single capability attenuated to an explicit allow-list). The mailbox send tool requires the matching capability, so a message_sent to a non-granted peer is uninvokable.
  • All of the owner's agents — the agent holds a broader "may message any agent in this owner's roster" capability, optionally attenuatable to a scope (e.g. only agents tagged into a shared project).
  • Direction is configured per pair. The owner decides each direction independently — granting agent A the right to message agent B does not imply B may message A — so messaging graphs among the owner's agents can be one-way.
  • Revocation is intrinsic. Because the right to message is a held capability, the owner revokes it the W3Wallet way and the next send fails — no change to the recipient required.

8. Lifecycle, concurrency, and placement

  • The supervisor durably owns the conversation catalog. Every conversation the supervisor hosts is an entry in its durable conversation catalog, and the entry persists across the conversation's whole lifecycle: active (live PTY), paused (no live PTY, resumable from its recorded native session reference and durable home), and archived (terminated, retained for the record). The catalog and each agent's runtime config (§2) are the supervisor's durable state; they survive a supervisor restart. This is the concrete meaning of "the supervisor is stateful."
  • A conversation stays live until the user says otherwise. There is no idle timer and no automatic reaping: a running conversation keeps its container and PTY warm — the AI CLI retaining its full native, in-process context — until the user pauses or terminates it from a managing surface. Warmth is cheap (an idle CLI spends no tokens) and continuity beats reconstruction. This is not the keep-warm anti-pattern: it is the supervisor's first-class session policy for an active conversation the user has deliberately left open. Pausing moves the catalog entry from active to paused; the entry itself never leaves the supervisor.
  • Anything else that ends warmth is just an unplanned pause. Host loss, resource pressure, or supervisor maintenance may still force a session down — but because the CLI's session store lives on the supervisor-owned durable home continuously (§9) and the catalog entry is itself durable, a forced shutdown (even an ungraceful crash) leaves exactly the same resumable state as a deliberate pause: on restart the supervisor rediscovers the conversation as paused in its own catalog. Durability never depends on warmth.
  • Pause/resume custody is the supervisor's; AiCliSupervisorManager orchestrates and aggregates. Because the supervisor now durably holds the conversation catalog and resume state, pause/resume custody lives with the supervisor, not the Manager. On pause the supervisor simply flips its own catalog entry to paused (recording harness + native session reference + durable-home reference) — it does not flush that record elsewhere and forget. The Manager's role becomes relaying the user's pause/terminate/resume commands to the owning supervisor and aggregating the supervisors' durable catalogs into one fleet view; resuming is asking the owning supervisor to relaunch from its own catalog entry. (This walks back the Manager workstream's earlier "session-state custody = the Manager holds resume records" extension — custody moves down into the now-stateful supervisor; the two documents are updated together.)
  • One live conversation per agent. An agent is a single stateful entity (one history, one workspace, one runtime config), so at most one conversation per agent is live at a time; the supervisor serializes them. A second open against a busy agent continues the existing live session rather than forking a second PTY — two concurrent sessions would race on the agent's shared history and workspace. (Distinct agents run fully concurrently, multiplexed across supervisors.)
  • Initiation is console policy; the engine supports both human-open and trigger-wake. A conversation starts when a console opens or continues one for a human, or when an authorized trigger wakes the agent — the monitor wake of §12 today, SpawningPool's schedule- and mail-driven autonomy on the same mechanism as it lands. The engine must support autonomous waking because it powers SpawningPool; "conversations are human-initiated" is MyLittleAgi's product policy, stated in its own doc, not an engine rule — and AiCliSupervisorManager, as the direct admin surface over supervisors, exposes everything the engine supports. The mailbox alone still wakes no one: an agent with unread mail carries that backlog until its next conversation, however initiated (§6).
  • Placement is chosen when an agent is created; thereafter the agent is pinned to its home. Supervisors register themselves on the url:// fabric and report health/capacity via their existing status API; the AiCliSupervisorManager fleet registry is the natural substrate for discovering which supervisors exist and have headroom. Placement therefore decides which supervisor an agent is created on; a failed first launch can retry on a different host. But once an agent exists, its entire durable state — definition, config, catalog, history, mailbox — lives on that supervisor, so every conversation and resume for the agent happens on the same supervisor. Moving an agent to a different supervisor requires migrating its durable state, a deferred decision (see below); until then homing is fixed, and losing a supervisor's durable store strands its agents until that store is restored. The specific placement signals (load, workspace locality, cost) are deferred to the console/scheduler side.

9. Session continuity and resume

"Persistent entity with history" must feel continuous to the user — and the AI CLI's own machinery is what provides it. Every supported CLI already owns a native session store, a native resume invocation, and native context compaction/summarization; the platform deliberately builds none of those as components. Its whole job is two-fold: never lose the CLI's own state, and relaunch with the CLI's own resume.

  1. Warm session (default). The conversation's PTY/container stays alive until the user pauses or terminates it (§8). While warm, "continuing the conversation" is just more input to the same live CLI — zero rehydration, full native context. When the CLI's context window fills, the CLI compacts it itself, natively — that is its concern, not the supervisor's.
  2. Pause and resume (the CLI's own mechanism). The CLI's session store lives on the supervisor-owned durable home continuously — resume is a property of where the store lives, not of a clean shutdown, so deliberate pause, forced shutdown, and outright crash all leave the same resumable state. On pause, the supervisor records the resume state (harness + native session reference + home reference) in its own durable conversation catalog and flips the entry to paused (§8) — the AiCliSupervisorManager is notified so its aggregated fleet view stays fresh, but the durable custody is the supervisor's. On continue, the owning supervisor starts the same per-harness guest image against the same home and invokes the CLI's native resume with that session reference — the CLI restores its own context. A fresh W3Wallet instance principal (§5) is negotiated for the resumed run and re-granted the same capabilities from the durable definition.
  3. If a CLI cannot resume, that is surfaced honestly — the next conversation starts fresh (the CLI's limitation, not papered over by a platform reconstruction layer), while the durable history remains fully readable through the supervisor's read models below (surfaced by the consoles). There is deliberately no platform-run transcript-replay or summarization tier.

History read models are for humans and search, not for context seeding. The CLI's native session store is the canonical record of the conversation. As history is committed to the supervisor's own store (§3), the supervisor derives a rendered transcript and a SearchableBucket recall index — read models the consoles render and the agent's own recall tool queries — so history can be looked up by query. These views serve display and recall; rebuilding an AI CLI's working context stays the CLI's own job.

10. Isolation and network egress

The capability model (§5) bounds an agent's authority — what privileged actions it can take. It does not, on its own, bound blast radius: an agent runs an arbitrary AI CLI that executes code, so without a network policy it could exfiltrate its workspace/context to any host or probe the wider network. Egress policy is the second half of containment.

  • What a conversation container must reach: the supervisor (the guest's outbound side-channel dial — §2); the agent's LLM provider endpoints directly (OAuth and inference — §5); and the url:// targets of its granted, mounted tools. Everything else is what default-deny denies.
  • Default-deny egress, per agent — the policy end-state. Beyond the reachability set above, a conversation's container gets no outbound network by default; everything external flows through a granted capability or mounted tool, where it is authorized and auditable. The owner can opt a specific agent into broader egress as part of its definition (a web-research or coding agent that must git clone), ideally as an explicit allow-list rather than wide-open.
  • Enforcement is deliberately deferred. Initially, conversation containers run on Docker's ordinary bridge network with the policy above unenforced — the priority is landing the supervisor–guest inversion and launch-from-definition first. When enforcement is picked up, the candidates on the table are: an internal-only Docker network whose sole exit is a supervisor-run egress proxy enforcing a per-agent domain allow-list (the leading candidate, since provider endpoints are CDN-backed domains that IP-level rules track poorly); per-conversation kernel-level filtering; or a NetLab-style controlled network. NetLab is in any case the natural harness for testing whatever enforcement ships.
  • Isolation is layered with the capability boundary, not a replacement for it. Even an egress-enabled agent still holds no raw secrets for its tools and data and can still only invoke the capabilities it was granted; egress policy limits where its own code can reach, while capabilities limit what authority it can wield.

11. Testing: docker-required, modeled on NetLab

The supervisor requires Docker, so its test suite is a docker-required suite modeled on NetLab's — the platform's established pattern for exactly this shape of system:

  • Docker's absence is a loud failure, never a skip. A dedicated gate test asserts the Docker daemon is reachable and fails (not skips) otherwise, with script- and CI-level fail-fast checks in front of the suite — the environments that run these tests are expected to have Docker, so its absence is a real failure to surface.
  • NetLab's container-test conventions apply throughout: one scenario per test, unique per-run resource names so concurrent runs never collide, idempotent teardown in finally so cleanup failure never masks the real assertion, a label-based orphan reaper as the backstop for tests that die without cleanup, guest images built on demand from their shipped definitions (no manual setup on ephemeral CI builders), and CI running on a raw host runner where the real Docker daemon is available.
  • Two tiers of test payload. A fake-harness guest image — the real guest daemon wrapping a scripted stand-in CLI — covers everything the supervisor owns, fast and deterministically: container launch, the bootstrap/dial-back handshake and launch-token binding, driving the remote AiCliPty surface, the raw byte-stream subscription, pause/resume records, reconnect after a supervisor restart, and teardown. A real-CLI tier then proves the full stack per harness: the genuine per-harness guest image driven against a test-hosted fake provider server (the pattern the guest repository already proves today, running the real claude against a fake Anthropic server via the CLI's base-URL override), asserting a prompt→response round trip through supervisor → container → guest → real CLI → fake provider.
  • Non-Docker unit tests stay cheap by composing fakes or in-process PTYs at the daemon's general PTY-factory seam (§2) — the seam's generality exists precisely so most logic is testable without a container.

12. Monitor agents: keyword-triggered observation and timely advice

Any agent can be monitored by one or more other agents. A monitor agent watches a target agent's conversations, wakes when a keyword trigger fires, analyzes the conversation — typically with a different model than the target's — and injects an attributed advice message into the target's live conversation to help get it back on track. The goal is progressive timely disclosure: the monitor remembers what it has already disclosed and reveals the next piece of advice at the moment it helps, rather than dumping everything up front or repeating itself. Creating and attaching such monitors is deliberately trivial from AiCliSupervisorManager.

  • A monitor is an ordinary agent, not a new entity kind. It has a definition (persona/system prompt, CLI/model choice, capabilities), homes on a supervisor, and appears in catalogs like any agent. What makes it a monitor is purely additive: its definition carries one or more monitor attachments (target + trigger spec), and it holds the capabilities to observe and inject into its targets. Analyzing with a different model than the target's is the natural configuration (the monitor's definition picks its own model), not an enforced rule. Monitors may monitor monitors — a monitor conversation is an ordinary conversation, and every rule in this section applies to it unchanged.
  • Monitors are clients of the supervisor's public API — the same one the consoles use. A monitor may home on the same supervisor as its target or a different one; either way its machinery consumes exactly the operations AiCliSupervisorManager consumes — the attributed transcript subscription and ranged read, the conversation catalog, start/continue, and the atomic message send — over url:// to the target's home supervisor, with same-supervisor attachment just the degenerate case of the same path. One API surface serves both consumers, and its sufficiency bar is that neither the fleet console nor a monitor needs anything private (guiding principles). Notably, monitors do not ride the mailbox (§6) — injection is a push into a live session where mail is delivery into a future one — so cross-supervisor monitoring works without mailbox federation.
  • The attachment lives in the monitor's definition, on the monitor's home supervisor. Attaching a monitor changes nothing in the target's definition; the target's supervisor merely validates the presented capabilities on each observe/inject call. Attachments are agent-scoped by default (covering every conversation the target starts) or conversation-scoped for one-off shepherding; topology is many-to-many and unrestricted beyond what capabilities express. Attachments are same-owner only: they require the owner's wallet to delegate observe/inject authority over the target, and cross-wallet delegation does not exist anywhere in the platform yet (cross-owner/admin monitoring is deferred — see the deferred decisions).
  • Each attachment delegates two capabilities: may-observe and (optionally) may-inject. They are split because a watch-only monitor — one that analyzes and alerts without ever steering — is a legitimate shape, and because reading a session and writing into it are very different powers to revoke independently. Both follow the standard model of §5: delegated by the owner's wallet via the supervisor at attach time, attenuated to the target agent (or single conversation), lease-renewed. One nuance: while the monitor is dormant, its home supervisor exercises may-observe on the monitor's behalf to run trigger matching — the one place a capability is held at the supervisor level per attachment rather than minted fresh per conversation launch.
  • Trigger matching runs on the monitor's home supervisor, over clean attributed transcript text — never raw PTY bytes. The watch loop holds a subscription to the target conversation's attributed transcript-event stream — a required API surface: harness-parsed text deltas, each carrying a speaker (human, target agent, tool, or injected-by-monitor-M) and timestamp. Raw PTY bytes are unusable for matching: ANSI escapes and screen redraws split keywords mid-token and re-emit the same text, producing both misses and duplicate fires; each clean delta is instead evaluated exactly once, as it arrives, keeping wake latency low. A trigger is an OR-combined set of keywords, each in one of two modes: substring (fires on an occurrence anywhere, even inside a longer token) or word (fires only on a delimited occurrence — the boundary is a text edge or any non-alphanumeric character, i.e. regex \b semantics). Matching is case-insensitive by default, with a per-trigger case-sensitive flag. All speakers are scanned by default — the target agent's output, the human's input, and tool output all count; a per-trigger speaker filter (e.g. ignore injected-by-monitor speech) is the authoring tool for noise and loop immunity. Richer conditions (AND, regex, proximity) are deliberately out of scope: the escape hatch is waking more often and letting the monitor's model cheaply dismiss false alarms.
  • Waking is the existing continue-conversation operation, against a persistent monitor conversation per target conversation. The first trigger from a given target conversation starts a monitor conversation; every later trigger continues that same conversation (warm or native resume — §9 unchanged) with a structured trigger message: which target conversation, the matched keyword(s), the matched excerpts with speaker attribution, and a transcript cursor. That persistence is what makes disclosure progressive — the monitor remembers what it already advised; a fresh conversation per trigger would produce an amnesiac monitor repeating the same advice on every hit. A trigger is a prompt, not a mandate: deciding "false alarm, stay silent" is a valid outcome. The §8 one-live-conversation-per-agent rule holds — a monitor watching several targets activates serially and simultaneous triggers queue; truly concurrent monitoring means several monitor agents, which is fine because they are trivial to create. When the target conversation is archived, its monitor conversation is archived too; an agent-scoped attachment spawns a fresh monitor conversation when the target's next conversation starts.
  • Reading during analysis is a mounted built-in tool. The woken monitor pulls as much surrounding transcript as its analysis needs through a conversation-read-style built-in mounted into its container per §4, invoking as its negotiated identity under may-observe — a ranged, cursored read of the same attributed transcript surface. Attribution matters to the analysis itself: the monitor can see who said the triggering thing and shape its response accordingly.
  • Injection is an atomic message send — harness-aware in timing, attributed in-band. The supervisor API's sendMessage operation delivers a whole message as one serialized unit: concurrent senders (the human at a console terminal, several monitors) queue message-by-message and never interleave keystrokes; raw keystroke access remains only for the interactive terminal surface where a human genuinely is typing. Delivery timing is the target guest's job, not the monitor's: the harness-specific AiCliPty, which owns BUSY/IDLE detection, delivers immediately where its CLI supports mid-turn steering or queued input, and otherwise holds the message until IDLE. The injected text is platform-prefixed with its provenance (e.g. [Advice from monitor "name"]: …) — the target's model sees only PTY text and must know this is monitor advice, and the prefix being platform-applied means a monitor cannot impersonate the user — and the transcript event records speaker = injected-by-monitor-M. An injection consumes a turn of the target's budget (it processes the message like any input); the analysis consumes the monitor's — both already bounded by the spend-rate buckets of §5.
  • Loop safety: structural damping plus a circuit breaker. First-line damping: text injected by monitor M is not evaluated against M's own keyword triggers, and hits arriving while M is already awake (or has a wake pending) for that target coalesce into one wake with batched excerpts — a fifty-hit stack trace is one analysis, not fifty. These are damping, not guarantees: a monitor doing whole-conversation analysis re-reads its own injections regardless, and while a well-designed monitor should have no loop, bugs happen. The guarantee is the circuit breaker: per attachment, the supervisor counts wake→inject cycles since the last human message in the target conversation — the loop signature is agents ping-ponging with no human in the loop — and trips at a configurable threshold (platform default: 5 consecutive human-less cycles), with a rate backstop (N wakes within M minutes) for degenerate fast loops. Tripping is an emergency stop: the attachment stops evaluating, waking, and injecting; both conversations are left paused and intact as forensic evidence (nothing is unwound — the transcript showing the ping-pong is the finding); the attachment carries a durable tripped state that surfaces loudly in the Manager, and the owner is notified. Re-arming is an explicit human action — the breaker never auto-resets, because an auto-reset merely turns a loop into a slow loop. The spend bucket remains the outermost economic backstop; the breaker exists because throttling-to-accrual grinds silently, whereas a breaker stops the loop in seconds and tells someone.
  • Creation is a one-form console flow. Making minimal monitors trivial is a first-class requirement: the AiCliSupervisorManager attach-monitor flow (detailed in its doc) creates the monitor agent from a built-in advisor persona template, takes the keywords + match mode, target scope, watch-only vs. advising choice, and breaker thresholds (defaults pre-filled), defaults the monitor's home to the target's supervisor (any healthy supervisor is allowed), and relays create-agent → delegate-capabilities → attach to the supervisors in one submit — the console durably stores none of it. A user creating a simple keyword monitor types keywords and picks a model; the template's system prompt does the rest ("read the transcript, weigh who said what, inject one concise piece of advice or explicitly decide silence is right, never repeat advice already given").

Subsystem deltas — what each dependency needs (to make this shovel-ready)

This vision is mostly integration of existing primitives, not green-field. The concrete deltas:

Subsystem Delta required
AiCliHostSupervisorDaemon* (Api/Embedded/ServiceServer) Stay PTY-implementation-agnostic (the runServiceServer composition seam is the contract). Become the durable home of the agent: durably own, event-sourced and surviving restart, each homed agent's definition + permissions, runtime configuration (.claude/.codex, MCP, mounted tools — §2), conversation catalog (active + paused/archived, with resume state — §8), transcript history + read models (§9), and the inter-agent mailbox among its homed agents (§6). Expose all of it over url:// for consoles to read and drive. Add create-agent and start/continue-conversation-for-agent-X operations that assemble the rehydration bundle from stored state (§3); mount the granted tool executables and the durable home into the container; broker the W3Wallet identity negotiation (§5) and delegate the agent's grants on behalf of the owner's wallet, renewing leases only while the container lives; serialize to one live conversation per agent (§8); hold pause/resume custody in its own catalog. Monitor-agent support (§12): an attributed transcript-event API (subscription + ranged cursored read; speaker = human / agent / tool / injected-by-monitor), an atomic sendMessage operation (whole messages serialized — no keystroke interleaving), durable monitor attachments carried in the monitor's definition, the watch loop (keyword matching over subscribed targets, exercising may-observe while the monitor is dormant), trigger-wake via continue-conversation with a structured trigger message, and the per-attachment circuit breaker with durable tripped state.
AiCliHostSupervisorLauncher (entrypoint) Pivot the composition from host-PATH PTYs to the docker-guest factory (§2): configured per-harness guest images, container lifecycle (launch with bootstrap env, ensure-image-on-demand, teardown), and remote AiCliPty proxies over the side channel. Home of the docker-required, NetLab-modeled test suite (§11). Today's host-PATH harness detection and per-harness config assembly migrate into the guest images; the seam it composes stays general so unit tests can still wire local/fake PTYs.
AiCliDockerGuestDaemon (guest) Grow from transparent PTY proxy into the in-container daemon: embed the harness AiCliPty implementations; add the url:// side channel (mint the ephemeral identity, dial the supervisor from the bootstrap env, present conversation id + launch token, serve the AiCliPty contract and the raw byte stream); participate in the W3Wallet identity negotiation (§5); own the per-harness guest image definitions (pinned CLI + daemon, per harness). Keep the transparent stdio-proxy behavior — it is the debugging story.
AiCliSupervisorManager The fleet-admin console over supervisors — a frontend, not a store of agent state (§8). Relay create/start/continue/pause/terminate/resume commands to the owning supervisor and aggregate the supervisors' own durable state (agents, catalogs incl. paused/archived, usage) into one fleet view, with supervisor configuration/health/placement front and center. Keeps only thin operator metadata (aliases, thread names) + a cached aggregate. Durable session/agent state lives on the supervisors. (Revises that workstream's earlier "the Manager holds resume records" extension.) Adds the attach-monitor flow and monitor visibility — watched-by badges, attributed injections in the thread terminal, loud tripped-breaker surfacing (§12 and its own doc).
MyLittleAgi (all layers) Becomes the end-user console over a fleet of supervisors — the consumer-friendly sibling of the Manager, not a durable brain. A user creates and talks to their agents without knowing which supervisor runs them (backing supervisors surface only deep in settings). It holds no agent state — definitions, history, mailbox, config, and catalog all live on the home supervisor; MyLittleAgi issues create/converse commands and renders the read models the supervisor exposes. Keeps only thin user-facing durable state (backing-supervisor list, names/preferences, a cached aggregate) plus its consumer surface — the LittleAGI native app, voice/hotword, littleagi CLI, tray, and file-manager entry points. Home of the "conversations are human-initiated" console policy (§8) — an engine that supports trigger-wake, restricted by product choice; whether it exposes monitor creation is its own UX decision. Covered in detail in its own workstream.
LambdaServer Capability-gated invocation (already on its roadmap) + a way to enumerate "the functions this agent may call" and a thin CLI wrapper so a registered function can be mounted as a command-line tool the AI CLI runs.
W3Wallet The agent / hosted-workload principal — established by the supervisor↔guest negotiation (§5) and granted its authority by the supervisor's delegation on behalf of the owner's wallet; delegation-by-reference for tools/data/credentials/peer-messaging; quota-as-attenuation for agent budgets; renewal-based instance leases. Owns the deferred decisions this engine depends on: whether the principal reuses the guest's transport keypair, the spend-bucket-to-invocation wiring, and the service-initiated funds dialog. Adds the may-observe / may-inject monitor capabilities (§12), and eventually the cross-wallet delegation that cross-owner monitoring waits on.
SearchableBucket Per-agent buckets for task context/recall, reached via a built-in mounted CLI tool, capability-scoped to the agent; also the home of the derived history recall index (§9). (Largely exists; needs the capability scoping + a CLI wrapper.)
SimpleFileSystem Per-agent / per-conversation workspace mounts, capability-scoped, surfaced as a built-in tool — and the supervisor-owned durable per-agent home mount that carries the AI CLI's own config files, credentials, and session store, making every pause/crash resumable and embodying the supervisor's stateful agent runtime config (§2, §9). (Largely exists; needs the per-agent scoping and durable supervisor ownership.)
Event Log + Observables The durable substrate for definitions, history, and the mailbox; inbox and roster projections are reactive Observables. (Exists; this is adoption.)
ContainerNursery Hosts the consoles (MyLittleAgi, the Manager) and can host supervisors themselves. Conversation containers are not ContainerNursery-managed: the supervisor launches and owns them directly via Docker, because their lifecycle is user-driven (§8) rather than lazy-start/idle-reap.

Scope

In scope for this vision: the supervisor-as-durable-agent-home / consoles-over-the-fleet split — the supervisor durably owns the whole agent (definition, permissions, runtime configuration, conversation catalog active + paused/archived, transcript history + read models, and the inter-agent mailbox among its homed agents), event-sourced and surviving restart, with each agent homed on exactly one supervisor; and MyLittleAgi (end users) and AiCliSupervisorManager (fleet admins) as audience-specific consoles over the fleet that hold no agent state; the supervisor–guest inversion — a privileged, Docker-requiring supervisor driving per-harness guest containers over a url:// side channel (guest dials out; supervisor drives back; the guest keeps its transparent stdio proxy and additionally serves the AiCliPty contract plus a raw byte stream); launch of a conversation for a stored agent via a rehydration bundle the supervisor assembles from its own durable state; capability-scoped conversations with the agent as a W3Wallet instance principal whose identity is negotiated supervisor↔guest over the side channel and granted its capabilities by the supervisor on behalf of the owner's wallet; tools as capability-gated, LambdaServer-backed CLI executables mounted into the container that invoke as that negotiated identity (the mounted set being the grant) plus SearchableBucket / SimpleFileSystem / mailbox built-ins; in-container native OAuth for the agent's own model credential, persisted on the supervisor-owned durable home; the asynchronous inter-agent mailbox (send = command, inbox = projection), scoped to an owner's agents homed on the same supervisor, with all-vs-subset peer permissions as capabilities; user-driven lifecycle — conversations stay warm until paused/terminated, with pause/resume custody in the supervisor's own durable catalog (a console relaying the command) and continuity supplied by each CLI's native resume and compaction over the durable home; default-deny per-agent egress as the policy end-state (enforcement deferred); a docker-required, NetLab-modeled test suite; monitor agents — any agent watchable by one or more monitor agents, same- or cross-supervisor, that wake on keyword triggers matched over an attributed transcript stream, analyze with their own model, and inject platform-attributed advice via an atomic message send, guarded by trip-open circuit breakers (§12); and a one-live-conversation-per-agent lifecycle — with initiation (human-open vs. trigger-wake) a per-console policy (§8) — with an agent placed on a supervisor at creation and pinned to that home thereafter.

Explicit non-goals: running AI CLIs as the supervisor's own host processes in production (the launcher's current composition — retired by §2, though the general PTY-factory seam remains for tests); raw secrets inside an agent container for tools, data, and delegated credentials (the agent's own model credential is the named exception — §5); platform-built resume, transcript replay, or summarization (continuity is the CLI's native machinery — §9); a console (MyLittleAgi or the Manager) durably owning any agent state — definition, permissions, config, catalog, history, and mailbox all live on the home supervisor; the consoles hold only thin audience-facing metadata + a cached aggregate; migrating an agent's durable state to a different supervisor, or cross-supervisor mailbox federation (both deferred — an agent is pinned to its home and messaging is same-supervisor for now, see below); idle-timer reaping of live conversations; synchronous agent-to-agent RPC; autonomous-execution product surfaces — the engine itself supports trigger-woken conversations (§8, §12), but schedule-/mail-driven autonomy as a product (waking agents on mail or a timer with no live session, large autonomous farms, and their loop/economic controls) is SpawningPool's surface, and "human-initiated only" is MyLittleAgi's console policy, stated in its doc; cross-owner monitor attachments (same-owner only until W3Wallet's cross-wallet delegation exists — §12); cross-user / cross-wallet messaging (the mailbox is single-user only — no cross-wallet directory or delegation); MCP-based tool surfacing (tools are mounted CLI executables for now); concurrent conversations for a single agent; ambient (non-capability) authority for any privileged action. This document also intentionally defers task and milestone breakdown.

Note — this is the second, larger step of an ownership migration. This document once framed the supervisor as stateless muscle over a durable MyLittleAgi brain. A first revision made the supervisor stateful for runtime config + conversation catalog. This revision completes the move: the supervisor is the agent's full durable home — definition, permissions, config, catalog, history, and mailbox all live on it — and MyLittleAgi and AiCliSupervisorManager converge into sibling consoles over the supervisor fleet, differing only by audience (end users vs. fleet admins). The companion MyLittleAgi and AiCliSupervisorManager docs are updated in lock-step.

Deliberately deferred decisions

Per this repository's convention, no open question is left dangling: everything below is a decision explicitly chosen to remain open here, each with the workstream that owns resolving it.

  • Whether the agent's W3Wallet principal is the guest's transport keypair or a separate bound keypair → W3Wallet (hosted-workload capability provisioning). The engine's handshake (§5) works identically either way.
  • The egress-enforcement mechanism → this workstream, revisited after the supervisor–guest inversion and launch-from-definition land; candidate mechanisms are recorded in §10. Until then containers run on the ordinary bridge network.
  • Mailbox delivery semantics (ordering; at-least-once vs exactly-once backlog re-display) → this workstream, now that the mailbox lives on the supervisor (§6).
  • Placement signals (load, workspace locality, cost) for where a new agent is created → the scheduler design on the console side (the AiCliSupervisorManager / MyLittleAgi frontends), atop the Manager's fleet registry.
  • Migrating an agent's durable state across hosts (moving an agent's whole durable state — definition, config, catalog, history, mailbox — to a different supervisor, so it can be re-homed, and to recover the agents of a lost supervisor) → this workstream, revisited after the durable-agent-home core lands. Until then, an agent is pinned to the supervisor that homes it (§8).
  • Cross-supervisor mailbox federation (letting an owner's agents that are homed on different supervisors message each other — §6) → this workstream, revisited after the single-supervisor mailbox lands. Until then, agents that must talk are homed together.
  • Budget mechanics (spend-bucket wiring to capability invocation, the service-initiated funds dialog, pause low-water marks, suggested-rate computation) → W3Wallet.
  • Cross-owner monitoring (a fleet admin's or another user's monitor observing an agent whose owner's wallet is not the monitor owner's — §12) → W3Wallet, alongside the cross-wallet delegation the mailbox's no-cross-user rule also waits on. Until then, monitor attachments are same-owner only.
  • Circuit-breaker default thresholds (the consecutive human-less wake→inject count and the rate backstop of §12) → this workstream, tuned at implementation against real monitor traffic; the mechanism and trip semantics are decided, only the numbers are open.

Relationship to other workstreams

  • MyLittleAgi — the end-user console over a fleet of supervisors, and the consumer-friendly sibling of the Manager. It holds no agent state (that lives on the home supervisor); it lets a person create and converse with their agents without caring which supervisor runs them, and renders the read models the supervisor exposes. The two docs are companions and must stay coherent. Its doc is updated in lock-step with this one.
  • AiCliSupervisorManager — the fleet-admin console over the same supervisor fleet: a frontend, not a store of agent state. It relays create/start/continue/pause/terminate/resume to the owning supervisor and aggregates the supervisors' own durable state (agents, catalogs, usage) into one fleet view, with supervisor configuration/health/placement front and center; its fleet registry is the placement substrate for where new agents are created. It is MyLittleAgi's admin-facing sibling — same supervisors underneath, different audience. Its doc is amended in lock-step with this one.
  • SpawningPool — the same engine, a different surface. MyLittleAgi is about spinning up individual agents and having human-driven conversations with them; SpawningPool is about large farms of mostly-autonomous agents. The boundary is console policy over a common engine capability: the engine supports both human-opened and trigger-woken conversations (§8, §12) — it must, because it powers SpawningPool — while MyLittleAgi restricts itself to human-initiated conversations as product policy, and SpawningPool embraces autonomous waking (mail- and schedule-driven execution and the farm-scale loop/economic controls it requires). They share AiCliHostSupervisor as the runner; the split is who drives, not what runs — and the Manager, as the raw admin surface, exposes all of it.
  • W3Wallet — supplies agent principals, delegated capabilities (tools/data/credentials/peer-messaging), and quota; this is one of its first real consumers, and it owns several of the deferred decisions above.
  • UrlResolver — the fabric under both the supervisor's public endpoint and the supervisor↔guest side channel; the guest's dial-out follows its direct-connection model.
  • NetLab — the model (and, where topologies are needed, the harness) for the docker-required test suite; also the natural test bed for whatever egress enforcement ships later.
  • LambdaServer — supplies the agent's tools as capability-gated functions.
  • Event Log / Observables — supply durable history and the reactive mailbox/roster projections (the CQRS substrate).
  • SearchableBucket / SimpleFileSystem — supply an agent's recallable context, working files, and the supervisor-owned durable home mount that carries its runtime configuration.

Graduation

AiCliHostSupervisor exists today as part of the AiCli project, but only as a raw PTY-session supervisor running CLIs on its own host, holding session state only in memory. This workstream graduates from "vision" when the design above is decomposed into concrete tasks and four things are real end-to-end: the supervisor–guest inversion — the supervisor launches a per-harness guest container and drives the real CLI inside it over the side channel, with the docker-required test suite green; the durable-agent-home core — the supervisor durably owns each homed agent in full (definition, permissions, runtime configuration, conversation catalog spanning active + paused/archived, transcript history, and mailbox), all surviving a supervisor restart, with the consoles holding no agent state; and the create-and-converse path — an agent created through a console (MyLittleAgi or the Manager) runs a capability-scoped conversation on its home supervisor, invokes a granted LambdaServer tool mounted as a CLI command, recalls context from SearchableBucket, exchanges a mailbox message with a permitted peer homed on the same supervisor, and is paused and later resumed from the supervisor's own durable catalog (a console relaying the command) using the CLI's native resume; and the monitor path — a monitor agent attached through the Manager's one-form flow wakes on a keyword trigger in a watched conversation (homed on a different supervisor as readily as the same one), reads the attributed transcript, injects a platform-attributed advice message via the atomic send, and a deliberately forced monitor loop trips the circuit breaker into its loud, human-reset emergency stop (§12). At that point it is no longer tracked as a workstream here. Milestones and sequencing are deliberately out of scope here and will be captured separately.