flockfs

Large files

flockfs keeps metadata, the event log and live text in PostgreSQL, and can keep file bodies in object storage: the local filesystem or any S3-compatible service (AWS S3, Cloudflare R2, MinIO). PostgreSQL storage costs roughly 10–20× more per GB than R2 (see vision.md) and caps files at 25 MiB; a blob store removes the cap (5 GiB by default) and makes large drives cheap.

WhatWhere
Live text (UTF-8, ≤ 2 MiB): current content, CRDT state, live editsPostgreSQL (small, hot)
Other files ≤ the threshold (256 KiB)PostgreSQL
Other files > the threshold (binaries, text > 2 MiB)blob store
Historical versions > the threshold (text included)blob store (moved in the background)
Everything, without a blob store (the default)PostgreSQL, files ≤ 25 MiB

Configuration

VariableDefaultPurpose
FLOCKFS_BLOB_STOREempty (= postgres)fs, s3, or postgres/empty (bodies stay in PostgreSQL).
FLOCKFS_BLOB_DIR<data>/blobsfs: where blobs live. Back it up with the database.
FLOCKFS_S3_ENDPOINTempty (AWS)https://<account>.r2.cloudflarestorage.com, http://minio:9000, … (http:// allows plain HTTP).
FLOCKFS_S3_BUCKET— (required for s3)Bucket name.
FLOCKFS_S3_REGIONus-east-1AWS region; auto for R2.
FLOCKFS_S3_ACCESS_KEY_ID, FLOCKFS_S3_SECRET_ACCESS_KEYfrom AWS_* / instance credentialsCredentials.
FLOCKFS_S3_PREFIXemptyKey prefix, so several deployments can share one bucket. Never share a prefix between deployments (their garbage collectors would not see each other's references).
FLOCKFS_S3_VIRTUAL_HOSTEDoffVirtual-hosted-style URLs (default path-style, which MinIO and R2 accept).
FLOCKFS_BLOB_THRESHOLD262144Bodies above this many bytes go to the blob store.
FLOCKFS_MAX_FILE_BYTES5368709120 (5 GiB)Largest file with a blob store (without one: 25 MiB).
FLOCKFS_BLOB_GRACE_SECS3600Unreferenced blobs younger than this are never collected.
FLOCKFS_SPOOL_DIRfs: <blob dir>/.tmp; s3: <data>/spoolWhere uploads are written while they are hashed. Needs free space for the largest concurrent uploads.

flockfs serve prints which store it uses at startup. Sandboxes (serve --sandboxes) always keep bodies in PostgreSQL.

Why the default stays in PostgreSQL

An unset FLOCKFS_BLOB_STORE keeps every body in PostgreSQL, as before. Switching an existing deployment to fs silently would make its pg_dump backups incomplete (bodies would live in the data volume), and a container run without a volume would lose them on the next recreate. Opting in is one variable; new installs with large files should set fs (backed up with the data volume) or s3.

Docker compose

# Filesystem, in the existing `data` volume
echo FLOCKFS_BLOB_STORE=fs >> .env && docker compose up -d

# MinIO next to flockfs (official MinIO images are no longer published; the profile uses the
# maintained pgsty/minio build, override with MINIO_IMAGE / MC_IMAGE)
cat >> .env <<'EOF'
FLOCKFS_BLOB_STORE=s3
FLOCKFS_S3_ENDPOINT=http://minio:9000
FLOCKFS_S3_BUCKET=flockfs
FLOCKFS_S3_ACCESS_KEY_ID=flockfs
FLOCKFS_S3_SECRET_ACCESS_KEY=<openssl rand -hex 24>
EOF
docker compose --profile minio up -d

# Cloudflare R2
FLOCKFS_BLOB_STORE=s3
FLOCKFS_S3_ENDPOINT=https://<account id>.r2.cloudflarestorage.com
FLOCKFS_S3_REGION=auto
FLOCKFS_S3_BUCKET=flockfs
FLOCKFS_S3_ACCESS_KEY_ID=<R2 token access key>   # Object Read & Write on the bucket
FLOCKFS_S3_SECRET_ACCESS_KEY=<R2 token secret>

# AWS S3: leave FLOCKFS_S3_ENDPOINT empty, set FLOCKFS_S3_REGION (or rely on instance roles)

The credentials need get, put (including multipart), head, list and delete on the bucket.

Layout

Blobs are content-addressed (SHA-256) under a per-drive space, a random name kept in the drive's own blob_space table:

{FLOCKFS_S3_PREFIX}{space}/{hash[0..2]}/{hash}
  • Identical bodies in one drive are stored once (a re-upload, a restore, a copy, every version of an unchanged file).
  • A drive is one key prefix: purging a drive deletes it (resumable through the control table blob_purges); an export reads only that drive.
  • files.blob / versions.blob hold the hash (and content is empty); blobs(hash, size, touched_at) lists what the drive uploaded.

Writes

Uploads never hold the workspace lock or a whole body in memory:

  1. The body streams to a spool file while it is hashed (POST /api/upload, PUT /api/files/{id}/raw), refused past FLOCKFS_MAX_FILE_BYTES (413 file_too_large).
  2. Small bodies and live text (≤ 2 MiB UTF-8) take the ordinary path (CRDT, three-way merges).
  3. Otherwise: an intent row (blobs.touched_at = now) is committed, the object is uploaded (filesystem: hard link or copy + fsync + atomic rename; S3: one PUT, or a multipart upload of 8 MiB parts, 4 in flight), skipped if an object of that size already exists.
  4. A short transaction under the workspace lock references the blob (files.blob, a version row) and touches its blobs row. Quota is checked before the upload and again here.

The JSON API (POST /api/entries, PUT /api/files/{id}) works as before for bodies up to 25 MiB; large binaries it receives are offloaded the same way before its transaction.

Reads

GET /api/files/{id}/raw streams from the blob store, with one Range (206, Content-Range; 416 when unsatisfiable), Accept-Ranges: bytes, Content-Length, X-Flockfs-Revision and, for blobs, ETag: "sha256-<hash>". ?revision=N reads a kept version. JSON reads (GET /api/files/{id}, /path) refuse bodies over 25 MiB with 413 use_raw; history inlines version bodies up to 25 MiB and marks larger ones omitted: <size> (read them with raw?revision=). Internally get_raw/raw_row still return whole bodies (reading blobs), so callers see one representation; streaming paths use Store::content / Store::body_at.

Garbage collection

A blob is garbage when no files row and no kept version references it — after a file is overwritten or deleted and retention pruned the versions that held it, or never referenced (an upload that failed half-way). Each maintenance pass (with history retention, hourly) per awake drive:

  1. moves inline bodies above the threshold to the blob store (migration, below), then
  2. collect_blobs: under the workspace lock, selects unreferenced blobs untouched for the grace period FOR UPDATE, deletes their objects, then their rows.

A concurrent write can never lose its blob: its intent made the row recent (not a candidate); its referencing transaction takes the same lock and fails with BlobCollected (the server uploads again, transparently) if the row is gone because the upload outlived the grace period. Uploads always check the object itself, never the row, so a collection that deleted objects but failed to commit cannot leave a dangling reference either.

Not collected: objects with no row at all (a process killed between an upload and its intent's cleanup cannot produce these, since the intent commits first; a bucket shared by two deployments under one prefix can). Use distinct prefixes.

Migration

Turning a blob store on needs no downtime. Existing bodies stay readable in PostgreSQL; the maintenance pass moves them in bounded steps (64 rows per step, 16 steps per pass): non-live files above the threshold, then versions above it. Each row is read, uploaded, then switched in its own short transaction only if unchanged (files: same revision; versions: sha256(content) matches), so the migrator is idempotent and resumes after a restart. Reads resolve both forms throughout. Store::migrate_blobs runs one step on demand.

Turning a blob store off again is not supported while references exist (reads fail with "no blob store is configured").

Quotas and usage

Byte usage counts each file's current size whether inline or in the blob store (files.size holds the blob size; Quota::usage and the plan checks use it). Versions are not counted, as before. Store::blob_usage reports the distinct blob bytes a drive references.

Clients

  • CLI: flockfs put / flockfs get stream (POST /upload, PUT /files/{id}/raw, GET /raw to a temporary file then rename); flockfs export streams the zip to a file and extracts entry by entry (ZIP64 for entries or archives past 4 GiB).
  • Sync daemon: files above 16 MiB on either side stream to and from disk, hashed in chunks; they sync as whole-file binaries (no merge; concurrent changes become conflict copies). A 413 is remembered until the file changes.
  • Mount: buffers past 16 MiB move to a temporary file ($TMPDIR/flockfs-mount) and upload streamed; reads of large files are Range requests, never whole downloads. Partial writes into an existing large file download it to the temporary file first.
  • SDK: upload(path, body), uploadTo(id, body, {revision}) (Blob, bytes or a ReadableStream), rawResponse(id, {revision, range}) for streaming reads; raw/blob accept the same options. The web app uploads files above 16 MiB with upload.
  • Export (GET /api/export): the zip is written to an unlinked spool file with blob bodies streamed into it, then streamed out with Content-Length.

Testing

cargo test --test blobs runs the cases against the filesystem store; cargo test --test blobs_s3 runs the same cases against an S3-compatible endpoint (FLOCKFS_TEST_S3_ENDPOINT, default http://127.0.0.1:59000, bucket flockfs-test, keys flockfstest / flockfstest-secret-123) and skips, saying so, when nothing listens there. A local MinIO:

docker run -d --name flockfs-blobs-minio -p 127.0.0.1:59000:9000 \
  -e MINIO_ROOT_USER=flockfstest -e MINIO_ROOT_PASSWORD=flockfstest-secret-123 pgsty/minio server /data
docker run --rm --network host --entrypoint sh pgsty/mc -c \
  'mc alias set l http://127.0.0.1:59000 flockfstest flockfstest-secret-123 && mc mb l/flockfs-test'

FLOCKFS_TEST_BIG_MB (default 200) sizes the constant-memory round trip.

On this page