Repository · workstreams
Workstream: Blobstore
Status: Shipped · Deferred milestone: W3Wallet ownership/quota/billing, W3Wallet-gated reads, and SetupTool key bootstrap · Component: Maximize developer productivity · Docs: project page
Goal
Give every service and agent a way to store and retrieve opaque binary blobs on demand — large, atomic byte payloads addressed by content, reachable over url:// — so that nobody has to stand up object storage, manage an S3 bucket, or hand-roll an upload protocol just to keep a JAR, an image, or a serialized archive around. A developer (or an agent) should be able to say "store these bytes, here is their SHA-256, give them back to me later by that hash" and have deduplication, integrity verification, streaming, and retention handled for them.
This is the BlobStorage mechanism in the approved storage taxonomy made concrete as a hosted service. It is a reusable primitive for component 1, the binary-object counterpart to the path-structured SimpleFileSystem: it removes "where do I put these bytes?" from the list of things you have to solve when building a service. It is also the content layer beneath the planned File Relational Filesystem, which stores file contents as content-addressed blobs here while owning the link graph itself.
Current state
This is the most mature of the storage workstreams: it already exists and is deployed as the Blobstore project (documented in the Documentation Repository as the Blobstore project and listed in ALL_PROJECTS.md). It is a content-addressed binary store served over url://blobstore/, following the standard layered architecture:
| Layer | Repository | What it provides |
|---|---|---|
| Api | BlobstoreApi | The BlobstoreService contract and BlobReference type: putBlob/getBlob (streaming over InputStream), listBlobs, pinBlob/unpinBlob, blobExists. Blobs are identified by their uppercase-hex SHA-256 hash |
| Embedded | BlobstoreEmbedded | Local in-process implementation; blob bytes in AWS S3, metadata (content hashes, pin/ownership state, timestamps) in CockroachDB |
| ServiceServer | BlobstoreServiceServer | Hosts the Embedded behind url://blobstore/; SJVM client bytecode; ContainerNursery lazy-start (URL_BIND_DOMAIN) or standalone P2P; session-based 4MB chunked streaming (putBlobStart/putBlobChunk/putBlobFinish, getBlobStart/getBlobChunk/getBlobClose) |
| Client | BlobstoreClient | Reusable convergent and random-key AES-GCM encryption plus BZip2 compression, with permanent legacy AES/ECB decryption compatibility |
| CLI | BlobstoreCli | Thin command-line consumer of BlobstoreClient for upload, download, list, pin, unpin, and existence checks |
| Setup | BlobstoreSetupTool | Compose Desktop GUI that generates the AES keys and initialization code for configuring a Blobstore client |
| HealthCheck | BlobstoreHealthCheck | ProductionHealth check for the url://blobstore/ service |
| Test support | BlobstoreInMemory | Shared faithful in-memory BlobstoreService, S3, and metadata implementations for end-to-end tests |
What already works: content-addressed storage with automatic deduplication and integrity verification (blobs keyed by SHA-256); stage-verify-promote uploads that never expose unverified bytes at a final key; explicit pinning semantics (a publicKeyHash owner pins a blob to retain it indefinitely); optional pin-gated reads and reference-counted garbage collection, shipped behind production-safe LOG_ONLY / DISABLED defaults; session-based chunked streaming so multi-gigabyte blobs never have to fit in one message or in memory; a reusable client with convergent and random-key authenticated AES-GCM encryption + BZip2 compression; permanent decryption compatibility for legacy AES/ECB identifiers; and a dual-backend design (S3 for bytes, CockroachDB for metadata). It is deployed on two transports (per its ServiceAtlas page): url://blobstore/ on ContainerNursery and an HTTPS API at api.blobstore.wasmserver.com on Cloud Run. Production runs BlobstoreServiceServer 0.0.8 with read gating in LOG_ONLY and garbage collection DISABLED; on 2026-07-17 both registered ProductionHealth checks reported SUCCESS, including read-back and hash verification of the permanently pinned 10 MiB multi-chunk canary.
The deployed service already holds critically important production data. Because addressing is by content hash and blobs are write-once, everything already stored — bytecode JARs, build-cache bytes, serialized archives that other services depend on — must stay byte-for-byte readable indefinitely. That makes preserving read (and write) access to existing blobs the overriding constraint on this workstream: avoiding data loss takes priority over every hardening item below. Concretely, no change may alter the storage contract that already-stored blobs depend on — the SHA-256 content address and its casing, how the storage key is derived from that hash, the pin metadata that records ownership, or the chunk-reassembly path that streams a multi-chunk upload into a single stored object and streams it back on read. A regression in how chunks are reassembled, or in how chunk data and its ownership metadata are persisted, would silently strand existing blobs; the storage-format invariants are therefore treated as a frozen contract, documented on the project page and guarded in production (see the backward-compatibility milestone below).
So this workstream was not "build it from scratch" — it was "harden the deployed service into the canonical, platform-adopted BlobStorage primitive and close the gaps, without ever losing or breaking access to what is already stored."
Why it accelerates developers
- No object-storage boilerplate. A new service stores binary artifacts in Blobstore instead of provisioning an S3 bucket, choosing a region, and writing an upload protocol — and gets deduplication and integrity checking for free because addressing is by content hash.
- The right tool for opaque bytes. Where SimpleFileSystem is for data whose path hierarchy carries meaning (configs, templates, documents), Blobstore is for data consumed whole as an atomic unit — compiled bytecode JARs, images, serialized archives. The storage decision guide draws the line.
- Encrypted client-side. The bytes the platform stores are opaque even to the platform — a property an agent swarm handling sensitive intermediate artifacts can rely on. Convergent mode preserves cross-owner deduplication for shareable artifacts; random-key mode prevents equality leakage for genuinely secret data.
- Agent-ready. Reachable over
url://, an agent can stash a build artifact or serialized payload and recall it later by hash, while operators can deliberately progress from lifecycle observation to pin-gated reads and reference-counted collection — a natural fit for component 2.
Plan / roadmap
Milestones that hardened the deployed service into the canonical binary-blob primitive:
Standing guardrail — established across 2026-07-09 and 2026-07-10, and continuously enforced. Preserving existing data is never a completed checkbox. BlobstoreHealthCheck runs both a fresh multi-chunk round-trip and read-back of a permanently pinned 10 MiB canary (plus an optional known production blob), while the ServiceServer's golden reassembly test guards the same frozen storage contract at build time. Storage-affecting dependencies remain pinned; changing the S3 client, hashing, metadata schema, or chunk path is a deliberate data-migration decision. Every future change remains subordinate to this guardrail.
- [x] Harden
putBlobwith stage-verify-promote. Uploads stream to a unique staging object, verify the declared SHA-256, and promote the exact verified object with an ETag-conditional copy. A mismatched upload can no longer overwrite or delete a valid final blob; interrupted staging objects are recoverable garbage rather than visible content. - [x] Implement the pin-as-lifecycle model: read-gating + reference-counted GC. BlobstoreApi 0.0.2 defines owner-scoped reads. BlobstoreEmbedded 0.0.4 ships
LOG_ONLY/ENFORCEread gating andDISABLED/DRY_RUN/ENABLEDexplicit GC, with production remaining atLOG_ONLY/DISABLEDuntil an operator completes the documented flip checklist. The enabled path uses grace periods, database-clock arbitration, token-fenced condemned markers, timeout-bounded single-attempt S3 deletes, and recovery for staging orphans and stale markers. Re-pinning a collected blob cannot resurrect it; the caller must re-upload. A deliberately documented residual race window remains because the pinned S3 SDK cannot conditionally delete an object; closing it requires an explicit storage-affecting dependency decision. - [x] Promote the client to a reusable library. BlobstoreClient publishes
community.kotlin.blobstore.client:community-kotlin-blobstore-client:0.0.1; convergent encryption and BZip2 moved out of the CLI, and BlobstoreCli is now a thin consumer. - [x] Ship two authenticated encryption modes. All new uploads use the
BSC2:v2 identifier and AES-GCM. Convergent mode uses a payload-bound deterministic nonce and remains dedup-stable; random-key mode uses a fresh key and nonce for secrets. Legacy AES/ECB identifiers remain decryptable forever and are locked by immutable golden vectors. - [x] Adopt Blobstore for ServiceServer client bytecode. BlobstoreServiceServer publishes its own SJVM client bytecode JAR at startup, content-addressed and pinned, best-effort off the serving path. The architecture pattern is documented; framework-level fetch-by-hash remains a proposal, not an implementation.
- [x] Reconcile naming, immutability, and SimpleStorage.
BlobStorageis the mechanism;Blobstoreis the canonical project and service. The storage taxonomy, project page, and BlobstoreApi KDoc state the write-once contract. Blobstore deliberately does not implement SimpleStorage's incompatible mutable/delete blob interfaces; the rationale is recorded on the project page. - [x] Publish faithful test doubles and exercise the chunk protocol end to end. BlobstoreInMemory publishes
community.kotlin.blobstore.inmemory:blobstore-in-memory:0.0.3; BlobstoreServiceServer runs the storage-contract suite over the real chunked session protocol, including golden multi-chunk reassembly, deduplication, mismatch rejection, and lifecycle-GC behavior. - [ ] W3Wallet-gated ownership, quota, billing, and reads; SetupTool key-bootstrap folding. Verify ownership instead of trusting caller-supplied
publicKeyHash, meter each owner's pins independently, gate reads/existence so convergent encryption's confirmation-of-file surface is closed completely, and fold BlobstoreSetupTool key bootstrap into the W3Wallet-backed flow. (deferred by user direction 2026-07-17; read gating ships pre-W3Wallet against caller-suppliedpublicKeyHash)
Open questions
- HTTPS API parity. Two transports are deployed (
url://blobstore/andapi.blobstore.wasmserver.com). Do both expose the full pin/quota/streaming surface, and is the HTTPS path also W3Wallet-gated once ownership becomes a verified capability?
Graduation
Hardening shipped (2026-07-17). The standing data-preservation guardrail, hardened upload path, pin lifecycle, reusable authenticated-encryption client, production adoption, naming reconciliation, and end-to-end contract coverage are delivered. The W3Wallet-dependent ownership/quota/billing, fully gated-read, and SetupTool-bootstrap items were deferred by user direction on 2026-07-17 and remain visible above as future work. The shipped architecture, operator modes, compatibility commitments, and known residual window live on the Blobstore project page; the HTTPS-parity question remains deliberately open.