← Workstreams

Workstream: Hard Drive Manager

Status: Planned · Component: Maximize developer productivity

Goal

Give anyone with data scattered across many storage surfaces — external hard drives that sit offline in a closet, remote object stores like S3, NAS shares, in-use production machines — one durable catalog that knows what lives where and whether it is still intact. A user registers each surface as a named drive by handing the manager an implementation of the SimpleFileSystem contract, along with a human-chosen name and free-text description/notes; the manager records when the drive was seen and captures checkpoints — content-indexed snapshots of every path with its checksum and metadata (size, modified time). From then on the catalog answers, without plugging anything in: what is on that drive? where do copies of this file live? which files exist on only one surface? what changed since the last checkpoint? which drives haven't been seen in months? — and, when a drive is attached, have any of its bytes rotted since last time?

Concretely, the catalog is a data artifact, not a service concept: a file-metadata index — per-file records of path, size, dates, permissions where the surface exposes them, and content hash (plus, in v3, escrowed encryption keys) — grouped by drive, together with the drive records, policies, and identity records, with the journal as its history. Everything else in this plan — manager, daemon, CLI, WUI — is machinery around that artifact: storing it, syncing it, querying it, policing it. The hosted manager stores many catalogs the way a filesystem stores many files; holding a catalog's capability is what it means to control one, and drive names are unique within the catalog that records them — no global namespace, no accounts.

Drives that live on a user's own always-on servers — a NAS, a homelab box, a production machine — register through a self-hosted daemon paired to the manager, so their sightings and scans happen automatically instead of waiting for a human to attach anything.

Expectations are explicit, not implied: a replication policy — a default for all files, overridable for particular files — states what "safely stored" means (for example: at least three replicas, at least one offsite, at least one always-online, and a replica whose drive hasn't been seen in a month is at risk), and the catalog continuously reports what falls short of its policy and suggests what would fix it.

The manager stores index data only, never file contents — it is a catalog and integrity auditor over storage the user already has, not a backup destination. Contents stay where they are; storing bytes is what BlobStorage and SimpleFileSystem are for, per the storage decision guide.

Current state

Greenfield — nothing exists yet. The seams it builds on do:

  • The drive-facing contract is SimpleFileSystemApi. A drive is a single filesystem, so the natural seam is the neutral single-tenant SimpleFileSystem interface the SimpleFileSystem plan is extracting; until that lands, an adapter over today's SimpleFileSystemServiceApi plus a filesystem name serves the same role. Anything that can implement the contract can be a drive: a local disk, an S3 bucket, a hosted url://simple-filesystem/ filesystem, a remote machine.
  • The contract exposes per-entry size and modified time without touching content (listRecursively, metadata) but no checksum, so computing a hash means reading the file's bytes through the adapter. That constraint shapes the hashing policy below.

The delivery shape is the standard layered architecture — HardDriveManagerApi / Embedded / ServiceServer (url://hard-drive-manager/) / CLI / WUI — plus a self-hosted HardDriveManagerDaemon for drives that live on the user's own servers, enrolled per Manager–Daemon Pairing.

Why it accelerates developers

  • "Which drive is that on?" becomes a query. Today the answer is plugging in drives one at a time and browsing; the catalog answers by drive name, path, or content hash — including for surfaces that are offline or remote.
  • Bit rot becomes detectable instead of discovered at restore time. Cold backup drives fail silently. A verify checkpoint re-hashes content and flags every file whose bytes no longer match the recorded hash — for most backups, the first integrity audit they have ever had.
  • Redundancy is auditable — and policed. Because entries are content-addressed, "which files exist on fewer than two surfaces" is a report, not a guess; replication policies turn that report into a standing contract, with the catalog flagging every file that drifts below its declared expectations — the difference between believing data is backed up and knowing it is.
  • Agent-ready. Reachable over url://, an agent can answer "where do copies of X live" or "what changed on drive D since May" without a human mounting anything.

Plan / roadmap

The roadmap is staged: v1 delivers the catalog end to end, v2 teaches it to see inside archives and disk images, v3 makes drives self-identifying, health-aware, and protected — encrypting drives through existing vault tools and hardening the catalog itself against both disclosure and loss.

v1 — the catalog: drives, checkpoints, policies

  • [ ] Api + Embedded — drive registry and checkpoints. HardDriveManagerApi defines the contracts: register a drive (unique human-chosen name, free-text description/notes, and the declared attributes replication policies constrain over — site/location and availability class, so an S3 bucket registers as offsite and always-online while a closet USB drive is onsite and usually-offline), record sightings, take a checkpoint against a supplied SimpleFileSystem implementation, list drives and checkpoints, diff two checkpoints. A sighting is cheaper than a checkpoint but still meaningful — sighting a reachable drive runs a metadata sweep: the entries on the drive are re-listed and their sizes and modified dates compared against the index, confirming the recorded contents are still present (hashes carried forward on a full match) without transferring any file data — cheap even when the drive is just a network mount (SMB/CIFS). A sweep that finds drift is the cue to take a real checkpoint; every checkpoint records a sighting too, so "last seen" always means "seen, with contents metadata-confirmed." The Embedded owns the tree walk and index build, Clock-injectable so tests drive time with ManualClock. Per-drive and catalog-wide exclusion rules (trash folders, caches, build artifacts) keep noise out of checkpoints from the start. Catalog persistence goes through a pluggable catalog-backend interface in the Api — the catalog being a data artifact, where it lives is a backend choice, and custom backends are a supported extension point. The default backend keeps the catalog's home replica on the user's own machine, synchronized/replicated to one or more of the managed drives themselves (the loop-free reserved-area placement machinery v3 completes); the hosted manager below is simply another backend, and all replicas converge through the multi-writer event store below.
  • [ ] Index structure (decided): a content-addressed hash DAG with sharded directory nodes. Each file entry records (name, content hash, size, modified time, permissions/ownership where the surface exposes them — the base SimpleFileSystem contract deliberately omits permissions today, so daemon-fronted drives capture them from the local filesystem); each directory is a node whose own hash covers its children, so an unchanged subtree is a single shared hash and a checkpoint with no changes stores only a new checkpoint record pointing at the existing root hash. The pathological case for a plain merkle tree — one directory with a huge number of children, whose node would re-serialize entirely on any single-child change — is handled by sharding: a directory node above a fanout threshold splits into B-tree-style pages keyed by name, so a million-entry directory rewrites the O(log n) pages on the update path, not one giant node. Nodes are immutable and keyed by their own hash, so the node store deduplicates across checkpoints and across drives automatically.
  • [ ] Catalog journal — checkpointing and self-integrity (decided shape). Every catalog mutation — a drive registered, a sighting, a checkpoint root adopted, a policy edit, an identity record (below) — is an entry in an append-only, hash-chained journal: the catalog's write-side source of truth in the CQRS sense and a natural fit for the Event Log. The journal is periodically checkpointed into a compact snapshot so recovery and replay stay bounded — restore is the latest snapshot plus the journal suffix — and integrity checking turns the auditor on itself: every DAG node must re-hash to its key, the journal chain must verify end to end, and snapshots must reconcile against the chain, run as a scheduled self-check raising the same alerts as drive verify. A catalog that polices bit rot on drives must not silently rot itself. Concurrency and write freshness (decided): the write side is a multi-writer event store. A single writable primary would only postpone the problem — a daemon that checkpoints while disconnected is already a concurrent writer, and buffer-and-deliver-later is synchronization with the conflicts hidden. So the journal is an event store that all of a catalog's writers — CLI, daemons, the hosted manager — append to concurrently: each writer appends domain events (checkpoint adopted, sighting, registry edit, identity record, policy change) to its own hash-chained sequence, and synchronization is a union of immutable events — idempotent and commutative, so there is no head to contend for and no conflict at the store level. The heavy work of a checkpoint is contention-free anyway (immutable content-addressed nodes; two writers producing the same node is dedup, not conflict). A deterministic order over the merged set — per-writer sequences with a writer-id tiebreak, the EventSourcingStreams/HierarchicalClock shape — makes every replica converge to the same state: where domain semantics collide (two concurrent checkpoints of one drive, two edits of one policy), the order picks the survivor by reducer rule and the event set preserves the loser, so conflict resolution is deterministic and never a user prompt. Events carry both when the drive was scanned and which writer observed it; a disconnected writer's events are durable in its own log the moment they happen and simply sync later. Tamper evidence survives per writer — each sequence self-verifies, and the merged journal is a DAG of chained segments — while snapshots are taken at a frontier (a vector of per-writer positions), which is also the read-consistency token: query answers state the frontier they reflect, read-your-own-writes means "my writer's position ≥ N," and derived indexes replay to a requested frontier or say they cannot. This is the Event Log/EventSourcingStreams model applied to the catalog — implemented standalone if the shared primitives aren't ready when v1 lands, converging on them when they are.
  • [ ] Hashing policy (decided): full hash first, metadata fast-path after, explicit verify. Content hashes are SHA-256 (uppercase hex, matching Blobstore). The first checkpoint of a drive reads and hashes every file. Subsequent checkpoints carry the stored hash forward when size and modified time are unchanged, reading only new and changed files — so routine re-checkpoints of a multi-terabyte drive run at metadata speed. An explicit verify checkpoint re-hashes everything and reports files whose content no longer matches the recorded hash. When the SimpleFileSystem contract grows a contentHash (planned there as the CAS ETag), backends that supply it let the manager skip content reads entirely.
  • [ ] File-identity merge/redirect primitive (decided shape). Content addressing answers "same bytes?", but users also need "same file?" — an edited document, a re-encoded photo, or a repacked archive changes hash while remaining the same logical thing. A small append-only layer of identity records sits above content hashes: merge declares two identities to be one (their replicas, history, and policy obligations pool), and redirect declares one superseded by another (obligations transfer to the target; the superseded content's copies stay queryable as prior-version replicas instead of counting toward the current policy). Queries, forensics, and policy evaluation resolve through the identity layer; the records live in the journal, so every merge is auditable and reversible by counter-record — and the duplicate-content report below is exactly where merge suggestions come from.
  • [ ] Adapters. Ship the adapters that make registering a real surface a one-liner: local disk (Okio-backed, so tests use FakeFileSystem), the hosted url://simple-filesystem/ service, and S3. Any user-supplied implementation of the contract works identically.
  • [ ] ServiceServer — the hosted catalog at url://hard-drive-manager/. Scanning must run where the drive is — an offline USB disk is reachable only from the machine it is plugged into — so the Embedded walks and hashes locally and pushes the finished checkpoint to the catalog; because nodes are content-addressed, a push transfers only the nodes the catalog does not already hold. The ServiceServer is stateless, persisting through the same catalog-backend interface — its production backend a scalable structured store (CockroachDB-class: easy horizontal scaling, fast access to individual records), per the off-container discipline the SimpleFileSystem hosted-instance milestone sets — and, a catalog of everything a person owns being sensitive, access is W3Wallet-gated via the standard secure-my-service path.
  • [ ] Self-hosted drive daemon — servers that host drives enroll via Manager–Daemon Pairing. A drive on a user's always-on server is served by a long-running HardDriveManagerDaemon on that host: it wraps the host's disks as SimpleFileSystem implementations, runs the Embedded's walk-and-hash locally, and pushes checkpoints to the catalog. Because the daemon stays connected, its drives' sightings are automatic — daemon liveness is the sighting, with the metadata sweep running locally — and the manager can run scheduled checkpoints and verifies remotely, which is what makes a replication policy's always-online replica class continuously verified rather than merely asserted. Sitting next to the disk, the daemon confirms presence at whatever strength the host affords — the metadata sweep, the underlying filesystem's own checksumming and scrub verdicts where it has them (ZFS/btrfs-style), or an outright local re-read that never crosses the network — the daemon decides, and asserts the level it achieved; assertions strong enough to prove bytes intact satisfy verify-class policy constraints without pulling a byte through the API. Enrollment follows the pairing handshake in both directions: the WUI mints a coursier one-liner carrying a fresh time-bounded registration key while dropping the controlling W3Wallet capability into the operator's wallet (Direction A), or a running daemon emits a claim token the operator pastes into the WUI's add-by-URL flow (Direction B). Per that doc's invariants, discoverability is never authority, the host gets the default-No consent gate, the terminal state is a revocable capability in a wallet — never an identity — and un-pairing is a capability revocation.
  • [ ] Cross-drive queries. The queries the catalog exists to answer, over the content-addressed index: where-is by content hash or path fragment; files present on fewer than N surfaces; what changed on a drive between two checkpoints; drives not sighted in N days. Answers are honest about observation staleness (decided): the catalog is a record of observations, so a drive's data is always "as of" its last checkpoint or sweep — reality can drift the moment a scan ends, and no design removes that. Rather than pretending liveness, every answer carries the observation times it rests on, the WUI displays them, and the policy freshness machinery is what bounds them — observations aging past their window go at-risk and get nagged, so staleness exists but is never silent.
  • [ ] Index-powered analytics. The reports the index makes nearly free, surfaced in the CLI and WUI: duplicate-content reports (the same hash at many paths on one drive is reclaimable space), largest-files and per-directory usage breakdowns, and search by name, path fragment, or content hash across every drive at once — including drives that are offline. The checkpoint history doubles as forensics: browse any drive as of any past checkpoint, and ask when a file (by hash) last existed anywhere and in which checkpoint it disappeared.
  • [ ] Pluggable derived indexes over the change feed (decided shape). The journal doubles as a subscribable change feed — every checkpoint diff, sighting, identity record, and policy transition is an event — and derived indexes are pluggable projections over it: rebuildable read models that are never authoritative (the DAG and journal stay the source of truth; a stale or corrupted derived index is dropped and replayed, and replay must converge to the same answers). The analytics reports above are themselves the first derived indexes riding this seam, and new ones plug in without touching the core — a media timeline over photo dates, a full-text index of names. Scan-time extractor plugins may enrich entries while a file's bytes are already in hand for hashing — the one moment content is read anyway — so projections can consume facts richer than names and sizes without a second pass over any drive. Change notification to external consumers follows Observables.
  • [ ] Replication policies (decided shape). A policy is a set of constraints a file's replicas must jointly satisfy — a minimum replica count (e.g. at least three), at least one replica offsite, at least one replica always-online, and a freshness window: a replica whose drive has not been sighted within the window (e.g. one month) is at risk and stops counting toward the other constraints. One default policy applies to every file unless a more specific per-file (or per-path-subtree) policy overrides it. Targeting (decided): selection by path, obligation on identity. A policy rule is a selection pattern over (drive, path-subtree) — the most specific rule wins per location, the default policy is the root fallback, and excluded paths carry no obligations at all. The obligation binds to the logical file identity currently at each selected location (its content hash, resolved through the merge/redirect layer): that identity must have the required replicas anywhere in the catalog — any drive, any path, loose or archived. Paths are the durable carrier of intent, so an edit re-binds automatically: the next checkpoint finds a new hash at the selected path and it inherits the policy, while the superseded content drops out of selection (remaining visible as a prior-version replica where a redirect records the succession). When several selected locations give one identity conflicting policies, the effective policy is the union of constraints — strictest per axis — so double selection never weakens anything. Evaluation rides the content-addressed index: a file's replicas are the distinct drives whose latest checkpoint contains its content hash, qualified by the drives' declared attributes and sighting recency — yielding, per selected file, satisfied, under-replicated, or at-risk, with the specific unmet constraints, always reported against the selected path the user recognizes rather than a bare hash.
  • [ ] CLI + WUI. CLI subcommands for register, sighting, checkpoint, verify, diff, where-is, and the staleness/redundancy/policy-compliance reports. WUI to browse drives and their notes, drill into a checkpoint's tree, view diffs and the redundancy report, and manage policies: set the default policy, override it for particular files or subtrees, and a suggestions view that lists every item currently failing its policy together with what would fix it — "add an offsite replica of these 212 files", "drive closet-wd-4tb hasn't been sighted in 25 days; its replicas go at-risk in 5". The WUI also owns the daemon enrollment surfaces from the pairing milestone: mint the one-liner for a new server, add a running daemon by URL and claim token, and un-pair by revoking its capability.
  • [ ] End-to-end tests following the testing standards: a real ServiceServer in-process, FakeFileSystem-backed drives, ManualClock — proving the shape-critical behaviors: a no-change checkpoint stores only a root record; a huge flat directory updates O(log n) pages on a single-child change; verify detects a single flipped byte; the metadata fast-path never misses a file whose size or modified time moved; the catalog self-check flags a flipped byte in its own node store and a broken journal chain; a derived index dropped and replayed from the journal converges to the same state; merge and redirect records resolve correctly in queries and policy evaluation (and a counter-record undoes them); policy evaluation flags an under-replicated file with the exact unmet constraint, and a replica crossing the freshness window (driven by ManualClock) flips to at-risk; a paired daemon's drives record sightings automatically while it is connected, and an unredeemed registration key expires after its window.

v2 — see inside archives and disk images

Backup drives are full of containers — zip and tar archives, disk images — whose contents a path walk cannot see. v2 teaches checkpointing to descend into them, so a file that sits loose on one drive and inside a tarball on another counts as two replicas of the same content.

  • [ ] Archive-aware indexing (decided shape). A container's entry keeps its own content hash and gains an inner tree, indexed into the same DAG exactly like a directory; inner files are hashed on their decompressed content, so where-is, redundancy, and policy evaluation unify loose and archived copies of the same bytes. The container hash is the fast-path boundary: an unchanged container carries its entire inner tree forward for free, and any change re-reads the whole container — whole-file reads are all the SimpleFileSystem contract offers, and the byte-range reads planned there are what would later let a zip central directory be read without streaming the archive. Formats are tiered: zip and tar (with gzip/bzip2/xz) come first via commons-compress; disk images — partition tables plus read-only filesystem parsing — are a deliberately later tier; encrypted or unrecognized containers are recorded as opaque with the reason, never silently skipped. Nesting is indexed to a bounded depth under entry-count and expansion-ratio caps, so a malformed or malicious container (a zip bomb) fails that entry, not the checkpoint. Policy evaluation counts archive-backed replicas but marks them as such: a compressed container is a single fault domain — one corrupt block in a .tar.gz can take out every member after it — so verify treats container integrity as covering all inner entries.

v3 — drive identity, health & protection

Drives should identify themselves and report their own condition, so the catalog notices problems before a restore does — and both the drives and the catalog itself gain an encryption story. The catalog is at least as sensitive as the data it maps: file names are sensitive, content hashes are sensitive (anyone holding a candidate file could probe whether you have it), drive locations and notes are sensitive, and once vault keys are escrowed (below) the catalog becomes the most sensitive artifact in the system — for disclosure and for loss. The resolution is that "never lost" and "never disclosed" only conflict for plaintext: replicate ciphertext promiscuously, and guard one small root secret with the specialized durability small secrets allow.

  • [ ] Drive fingerprinting. Registration writes a small ID marker into a reserved area at the drive root — the manager's writes to a drive are confined to that area, and the checkpoint walk always excludes it from the drive's own index (the exclusion that makes index placement loop-free, below) — so a re-attached drive is recognized regardless of mount point or machine: sightings record automatically on attach (the daemon and CLI watch for known fingerprints), and a renamed mount can never create a duplicate drive record. Surfaces that must stay pristine or are read-only can register marker-free, identified by name alone.
  • [ ] Health telemetry. Each sighting records what the surface can report: total and free capacity everywhere, and SMART attributes where a daemon fronts a physical disk. Failing-health trends (reallocated or pending sectors) mark the drive's replicas at-risk before the drive dies — feeding the same policy machinery as the freshness window — and capacity trends feed the suggestions view ("closet-wd-4tb is 92% full; future replicas won't fit").
  • [ ] Verify budgets — progressive scrub. A full re-hash of a multi-terabyte drive in one sitting is unrealistic, so verify runs on a budget (say, 100 GB per attachment), rotating through the drive across sessions with per-file last-verified timestamps, and policies gain a matching constraint — "every replica's bytes verified within N months" — turning bit-rot detection from an occasional heroic pass into routine hygiene.
  • [ ] Catalog confidentiality & durability (decided shape). The catalog's at-rest format migrates to client-side encryption — mechanical, because nodes are content-addressed: re-encode and re-push (v1–v2 run the W3Wallet-gated plaintext format until then). Durability comes from replicating ciphertext — but not whole-catalog-everywhere: each drive's reserved area carries that drive's own index by default (a small drive never inherits a big drive's index) plus the tiny encrypted registry, so a drive found in a drawer years later still identifies itself and its contents to whoever holds the root secret. Where else a drive's index is replicated — the hosted store, other drives — is the user's choice per drive, expressed as an index-placement policy evaluated by the same machinery as any file: an index replica riding an offline carrier is a snapshot-as-of, the freshness window applies, and a drive whose index or escrowed key material falls below its placement policy raises the same at-risk alerts and suggestions as any under-replicated file — the map and the keys are never quietly less protected than the data they unlock. Self-reference is cut off structurally: reserved-area contents are excluded from every drive's walk, so writing an index replica onto a carrier never changes the carrier's own checkpoint — without that exclusion, drive A's index landing on drive B would dirty B's index, whose replica on A would dirty A's, looping forever; with it, placement reaches a fixed point trivially, and the replica looks like part of the carrier without ever entering its index hash. Which carriers hold which index replicas, and how fresh each is, is registry bookkeeping — never part of any drive's content tree. Disclosure shrinks to custody of one root secret: a W3Wallet capability day-to-day, with a mandatory recovery kit minted at catalog creation (printable mnemonic, optionally split k-of-n across locations or people) whose unverified state is a standing policy violation the suggestions view nags about. Content identifiers become keyed tokens (HMAC under a catalog key) rather than raw SHA-256, so a leaked index fragment lets no one probe whether you hold a given file. The disclosure posture is the user's choice, per catalog — one policy engine, pluggable evaluation host: structure-sighted (default) shares only opaque tokens, sizes, and timestamps with the manager, which can then count replicas, evaluate policies, and push at-risk notifications server-side without ever seeing a name, hash, or key; fully blind shares nothing but encrypted nodes, and a designated paired daemon (or the client) evaluates policies and emits notifications instead. Blind can later be upgraded to sighted; sighted cannot retract what the manager has already seen — the switch is a one-way ratchet, stated up front.
  • [ ] Drive & file encryption via existing vault tools — the catalog as key escrow. Users can encrypt individual files or entire drives through established vault tools behind a pluggable seam — LUKS for daemon-fronted physical disks, file-granular tools for cloud and attachable surfaces — never home-grown cryptography. The vault keys are escrowed in the catalog, wrapped under the root key (an envelope hierarchy: one recovery kit recovers every encrypted drive — which is precisely why the recovery-kit obligation above is mandatory, since root-key loss now means losing drives, not just an index). File-granular encryption records both the encrypted plaintext identity and the ciphertext hash, so verify scrubs an encrypted drive with no keys present, and a plaintext copy on one drive still unifies with an encrypted copy on another as replicas of the same content.

Graduation

When Api, Embedded, ServiceServer, and CLI land, the hosted catalog is deployed, and a real multi-drive catalog is exercised end to end — drives registered (at least one via a paired self-hosted daemon), repeat checkpoints riding the fast path, at least one verify pass, cross-drive queries answering, and policy suggestions surfacing real replication gaps — this graduates to a first-class project: a project page in the Documentation Repository (plus ALL_PROJECTS.md) and the deployment recorded in ServiceAtlas. Graduation is judged on v1; the v2 and v3 stages then carry forward in the project's own documentation.