← Workstreams

Workstream: File Relational Filesystem

Status: Planned · Component: Maximize developer productivity

Goal

A filesystem where organization is a graph, not a tree: files are related to each other by links, no file has a canonical path, and the link structure may contain cycles. A traditional filesystem forces every file into exactly one place in one hierarchy; here a path is just a route through named links, so the same file participates in as many organizations as there are ways to reach it, and "what links to this file?" is a first-class query instead of a full scan. It removes "which single hierarchy do I file this under?" from the list of things a service, an agent, or a person has to solve — file it under all of them.

This is a reusable primitive for component 1 and a natural memory substrate for component 2: where SimpleFileSystem serves data whose single path hierarchy carries the meaning, the File Relational Filesystem serves data whose relationships carry the meaning.

It is a service, not a disk format. Every instance is reached over url:// — from a shared hosted store down to a local instance backed by a single physical file on disk (journal, indexes, and a local content-addressed blob layer packed inside, the way a SQLite database is one file), served at a local url:// address. The host operating system's filesystem is only a byte substrate: the graph is stored in files, never as files, so no OS filesystem limitation — forbidden directory cycles, case-insensitive names, path-length caps — ever touches the model, on NTFS or anywhere else. By the same token, the graph layer ships no human surface: browsing, organizing, and crossing the OS boundary belong to applications built on the public Api — first among them the File Relational Filesystem Explorer — and anything that must render the graph as a tree is an application-side projection, never a constraint on the graph.

The model

The semantics below are settled decisions.

Files are nodes; names live on links. A file is a stable, randomly-assigned identity carrying a mutable set of outgoing links, free-form tags and attributes, and optional content. Content bytes are not stored by the graph layer at all: they live in BlobStorage (Blobstore) addressed by content hash, while the graph layer owns identity, links, and metadata. This split is forced, not incidental — links stored inside content-hashed objects could never form a cycle (a cycle of hashes has no fixed point), so supporting cycles requires the link structure to live in a mutable layer outside the content hash.

Links come in two kinds. A named link has a name that is unique among its source file's outgoing links; names are the only labels a path segment can follow, so every path step is deterministic by construction. An anonymous link carries only tags and properties and expresses a relation (cites, derived-from) without joining the path namespace. A multi-valued relation that wants navigation reifies an intermediate container file (paper/cites/knuth76) — directories re-emerge as a pattern for plural relations rather than as a primitive. Both files and links carry tags; incoming links are indexed, making backlinks first-class.

Paths are routes, not identities. Absolute paths resolve from a single distinguished root file. Cycles merely mean some files have infinitely many spellings, so traversal operations are defined over the reachable file set, never the path set. Handles and the working location are trails — the walk actually taken — so .. steps back along the trail even though a file may have many referrers. A file with no named route from the root is still alive and well, addressed by ID, backlink, or query; a display path (shortest named route from the root) exists for human surfaces only and is never treated as identity.

Lifetime is unlink plus tracing GC. Nothing is deleted directly: links are removed, and files unreachable from the root and the pin set are garbage-collected — the generalization of POSIX link counts that survives cycles, where reference counting fails. Pins with a TTL (leases) cover a file between creation and first attachment, protect in-flight batch work, and provide a retention window that makes undelete a lookup rather than a feature.

Every mutation is journaled. Link, tag, and attribute changes append to a journal — the same append-only source-of-truth pattern the Event Log owns — yielding provenance ("who linked this here, and when"), point-in-time reconstruction, undo, and a change feed for replication, all from one primitive. Mutations group into atomic batches: a multi-step reorganization lands as a single journal entry, so observers and point-in-time reads see all of it or none of it, never a half-applied change. Because content in Blobstore is immutable, the journal makes the entire filesystem state reconstructible at any past instant.

Queries are files; views are live. A saved query is itself a file whose results materialize as its outgoing links — the smart folder as a first-class citizen. Virtual links are derived state: never GC references and never journaled. Consumers watch a file, a tag, or a standing query and receive deltas via Observables, with tag, attribute, and backlink indexes keeping queries index-backed rather than scans. Access control attaches to files and to capabilities (W3Wallet), never to paths — there is no canonical path to attach it to.

Merging is forwarding, not destruction. When two files are discovered to name the same real-world thing — the accidental identity split an autonomously-filing agent inevitably creates — they merge by a directional, journaled redirect: the caller picks the survivor, and the loser becomes a permanent forwarding identity. No ID is ever destroyed, because exported subgraphs, federated stores, and agents hold IDs that can never be rewritten; every dereference — path steps, content reads, queries, backlinks — transparently follows redirects (chains compress, re-merging within an alias set is a no-op, so redirect cycles cannot form), and rewriting physical inbound links is opportunistic compaction, never part of the operation. Set-shaped state unions freely (tags, anonymous links); conflicts — clashing named links or attributes, or differing content — are refused unless resolved in the same atomic batch, and differing content usually means the caller wanted a supersedes link between two live files, because merge is strictly for identity splits, never for versioning. A merge requires write capability on both files, carries the survivor's access control forward explicitly (a silent union would be a privilege escalation), records who asserted the sameness and why, and appears as a first-class event on the change feed so watchers re-point and external ID holders can heal their references; un-merging is a journal replay, with writes made after the merge staying with the survivor unless explicitly reassigned. Deciding that two files are the same thing stays with the consumer — the graph layer owns the mechanism, never the judgment.

Why it accelerates developers

  • Every organization at once. Tags plus live queries end the single-taxonomy fight: the same artifact appears in the release view, the dependency view, and the author view with zero duplication and nothing to keep in sync.
  • "What references this?" is finally answerable. First-class backlinks make the question POSIX cannot answer without scanning the world into an indexed query — the foundation for dependency tracking, impact analysis, and safe cleanup.
  • Provenance and undo are built in. The journal answers who organized what, when, and makes any organizational change reversible.
  • Agent-ready relationship memory. An agent's knowledge is relationship-shaped — X depends on Y, A supersedes B, this run produced that artifact — and real relationship webs contain cycles. A graph filesystem over url:// is a candidate substrate for the Agent Memory knowledge-compounding workstream.

Plan / roadmap

Milestones in proposed order; each earns its own detailed design in its implementation repository as it starts — this plan owns the shape, not the internals.

  • [ ] Semantic contract and Api. Capture the model above as the Api layer of a standard layered architecture service: files, the two link kinds, trails, pins/leases, atomic batches, identity merge/redirects, and the journal as the public contract, served over url://.
  • [ ] Core graph engine. The Embedded implementation: identity, link tables with per-file name uniqueness, tags/attributes, trail-based resolution, and the journal — with an in-memory backend for tests, a durable single-file backend for local instances (which pack their content-addressed blob layer into the same file), and content otherwise delegated to Blobstore.
  • [ ] Tracing GC and leases. Reachability collection from the root and pin set, with the retention window that yields undelete.
  • [ ] Identity merge and redirects. The forwarding merge above: permanent redirect entries that every dereference and backlink query follows, conflict refusal outside an atomic batch, merge events on the change feed — and redirects the GC never collects, since a forwarded ID must stay dereferenceable forever.
  • [ ] Query layer. Tag/attribute/backlink indexes, saved and live queries as files, and watch/subscribe via Observables.
  • [ ] Display paths and accounting. Shortest-named-route display paths, plus reachable vs. exclusive size (what a subgraph holds vs. what cutting it would actually free — du generalized to shared structure).
  • [ ] Subgraph export/import. The cp -r/tar generalization: reachability-bounded, cycle-safe serialization with IDs remapped on import and content deduplicated by hash.
  • [ ] Later phases. An optional constraint layer (declared invariants over curated regions of the graph); federation, for which the groundwork is already laid — stable IDs and content-hashed blobs make cross-store links natural, and the journal is the replication substrate.
  • [ ] Adoption examples. At least one real consumer whose data is genuinely relationship-shaped — the Agent Memory knowledge store is the headline candidate, and the External Brain's note-and-link graph is the headline human-facing one.

Graduation

This is greenfield: no repositories exist yet. It remains a workstream until the core engine, GC, and query layer are built and at least one consumer stores real relationship-shaped data through it over url://, at which point it becomes a first-class project with a Documentation Repository project page and an ALL_PROJECTS.md entry.