Large files
flockfs keeps metadata, the event log and live text in PostgreSQL, and can keep file bodies in object storage: the local filesystem or any S3-compatible service (AWS S3, Cloudflare R2, MinIO). PostgreSQL storage costs roughly 10–20× more per GB than R2 (see vision.md) and caps files at 25 MiB; a blob store removes the cap (5 GiB by default) and makes large drives cheap.
| What | Where |
|---|---|
| Live text (UTF-8, ≤ 2 MiB): current content, CRDT state, live edits | PostgreSQL (small, hot) |
| Other files ≤ the threshold (256 KiB) | PostgreSQL |
| Other files > the threshold (binaries, text > 2 MiB) | blob store |
| Historical versions > the threshold (text included) | blob store (moved in the background) |
| Everything, without a blob store (the default) | PostgreSQL, files ≤ 25 MiB |
Configuration
| Variable | Default | Purpose |
|---|---|---|
FLOCKFS_BLOB_STORE | empty (= postgres) | fs, s3, or postgres/empty (bodies stay in PostgreSQL). |
FLOCKFS_BLOB_DIR | <data>/blobs | fs: where blobs live. Back it up with the database. |
FLOCKFS_S3_ENDPOINT | empty (AWS) | https://<account>.r2.cloudflarestorage.com, http://minio:9000, … (http:// allows plain HTTP). |
FLOCKFS_S3_BUCKET | — (required for s3) | Bucket name. |
FLOCKFS_S3_REGION | us-east-1 | AWS region; auto for R2. |
FLOCKFS_S3_ACCESS_KEY_ID, FLOCKFS_S3_SECRET_ACCESS_KEY | from AWS_* / instance credentials | Credentials. |
FLOCKFS_S3_PREFIX | empty | Key prefix, so several deployments can share one bucket. Never share a prefix between deployments (their garbage collectors would not see each other's references). |
FLOCKFS_S3_VIRTUAL_HOSTED | off | Virtual-hosted-style URLs (default path-style, which MinIO and R2 accept). |
FLOCKFS_BLOB_THRESHOLD | 262144 | Bodies above this many bytes go to the blob store. |
FLOCKFS_MAX_FILE_BYTES | 5368709120 (5 GiB) | Largest file with a blob store (without one: 25 MiB). |
FLOCKFS_BLOB_GRACE_SECS | 3600 | Unreferenced blobs younger than this are never collected. |
FLOCKFS_SPOOL_DIR | fs: <blob dir>/.tmp; s3: <data>/spool | Where uploads are written while they are hashed. Needs free space for the largest concurrent uploads. |
flockfs serve prints which store it uses at startup. Sandboxes (serve --sandboxes) always keep
bodies in PostgreSQL.
Why the default stays in PostgreSQL
An unset FLOCKFS_BLOB_STORE keeps every body in PostgreSQL, as before. Switching an existing
deployment to fs silently would make its pg_dump backups incomplete (bodies would live in the
data volume), and a container run without a volume would lose them on the next recreate. Opting in
is one variable; new installs with large files should set fs (backed up with the data
volume) or s3.
Docker compose
# Filesystem, in the existing `data` volume
echo FLOCKFS_BLOB_STORE=fs >> .env && docker compose up -d
# MinIO next to flockfs (official MinIO images are no longer published; the profile uses the
# maintained pgsty/minio build, override with MINIO_IMAGE / MC_IMAGE)
cat >> .env <<'EOF'
FLOCKFS_BLOB_STORE=s3
FLOCKFS_S3_ENDPOINT=http://minio:9000
FLOCKFS_S3_BUCKET=flockfs
FLOCKFS_S3_ACCESS_KEY_ID=flockfs
FLOCKFS_S3_SECRET_ACCESS_KEY=<openssl rand -hex 24>
EOF
docker compose --profile minio up -d
# Cloudflare R2
FLOCKFS_BLOB_STORE=s3
FLOCKFS_S3_ENDPOINT=https://<account id>.r2.cloudflarestorage.com
FLOCKFS_S3_REGION=auto
FLOCKFS_S3_BUCKET=flockfs
FLOCKFS_S3_ACCESS_KEY_ID=<R2 token access key> # Object Read & Write on the bucket
FLOCKFS_S3_SECRET_ACCESS_KEY=<R2 token secret>
# AWS S3: leave FLOCKFS_S3_ENDPOINT empty, set FLOCKFS_S3_REGION (or rely on instance roles)The credentials need get, put (including multipart), head, list and delete on the bucket.
Layout
Blobs are content-addressed (SHA-256) under a per-drive space, a random name kept in the
drive's own blob_space table:
{FLOCKFS_S3_PREFIX}{space}/{hash[0..2]}/{hash}- Identical bodies in one drive are stored once (a re-upload, a restore, a copy, every version of an unchanged file).
- A drive is one key prefix: purging a drive deletes it (resumable through the control table
blob_purges); an export reads only that drive. files.blob/versions.blobhold the hash (andcontentis empty);blobs(hash, size, touched_at)lists what the drive uploaded.
Writes
Uploads never hold the workspace lock or a whole body in memory:
- The body streams to a spool file while it is hashed (
POST /api/upload,PUT /api/files/{id}/raw), refused pastFLOCKFS_MAX_FILE_BYTES(413file_too_large). - Small bodies and live text (≤ 2 MiB UTF-8) take the ordinary path (CRDT, three-way merges).
- Otherwise: an intent row (
blobs.touched_at = now) is committed, the object is uploaded (filesystem: hard link or copy + fsync + atomic rename; S3: one PUT, or a multipart upload of 8 MiB parts, 4 in flight), skipped if an object of that size already exists. - A short transaction under the workspace lock references the blob (
files.blob, a version row) and touches itsblobsrow. Quota is checked before the upload and again here.
The JSON API (POST /api/entries, PUT /api/files/{id}) works as before for bodies up to
25 MiB; large binaries it receives are offloaded the same way before its transaction.
Reads
GET /api/files/{id}/raw streams from the blob store, with one Range (206, Content-Range;
416 when unsatisfiable), Accept-Ranges: bytes, Content-Length, X-Flockfs-Revision and, for
blobs, ETag: "sha256-<hash>". ?revision=N reads a kept version. JSON reads (GET /api/files/{id}, /path) refuse bodies over 25 MiB with 413 use_raw; history inlines version
bodies up to 25 MiB and marks larger ones omitted: <size> (read them with raw?revision=).
Internally get_raw/raw_row still return whole bodies (reading blobs), so callers see one
representation; streaming paths use Store::content / Store::body_at.
Garbage collection
A blob is garbage when no files row and no kept version references it — after a file is
overwritten or deleted and retention pruned the versions that held it, or never referenced (an
upload that failed half-way). Each maintenance pass (with history retention, hourly) per awake drive:
- moves inline bodies above the threshold to the blob store (migration, below), then
collect_blobs: under the workspace lock, selects unreferenced blobs untouched for the grace periodFOR UPDATE, deletes their objects, then their rows.
A concurrent write can never lose its blob: its intent made the row recent (not a candidate);
its referencing transaction takes the same lock and fails with BlobCollected (the server
uploads again, transparently) if the row is gone because the upload outlived the grace period.
Uploads always check the object itself, never the row, so a collection that deleted objects but
failed to commit cannot leave a dangling reference either.
Not collected: objects with no row at all (a process killed between an upload and its intent's cleanup cannot produce these, since the intent commits first; a bucket shared by two deployments under one prefix can). Use distinct prefixes.
Migration
Turning a blob store on needs no downtime. Existing bodies stay readable in PostgreSQL; the
maintenance pass moves them in bounded steps (64 rows per step, 16 steps per pass): non-live files
above the threshold, then versions above it. Each row is read, uploaded, then switched in its own
short transaction only if unchanged (files: same revision; versions: sha256(content) matches),
so the migrator is idempotent and resumes after a restart. Reads resolve both forms throughout.
Store::migrate_blobs runs one step on demand.
Turning a blob store off again is not supported while references exist (reads fail with "no blob store is configured").
Quotas and usage
Byte usage counts each file's current size whether inline or in the blob store (files.size
holds the blob size; Quota::usage and the plan checks use it). Versions are not counted, as
before. Store::blob_usage reports the distinct blob bytes a drive references.
Clients
- CLI:
flockfs put/flockfs getstream (POST /upload,PUT /files/{id}/raw,GET /rawto a temporary file then rename);flockfs exportstreams the zip to a file and extracts entry by entry (ZIP64 for entries or archives past 4 GiB). - Sync daemon: files above 16 MiB on either side stream to and from disk, hashed in chunks; they sync as whole-file binaries (no merge; concurrent changes become conflict copies). A 413 is remembered until the file changes.
- Mount: buffers past 16 MiB move to a temporary file (
$TMPDIR/flockfs-mount) and upload streamed; reads of large files are Range requests, never whole downloads. Partial writes into an existing large file download it to the temporary file first. - SDK:
upload(path, body),uploadTo(id, body, {revision})(Blob, bytes or a ReadableStream),rawResponse(id, {revision, range})for streaming reads;raw/blobaccept the same options. The web app uploads files above 16 MiB withupload. - Export (
GET /api/export): the zip is written to an unlinked spool file with blob bodies streamed into it, then streamed out withContent-Length.
Testing
cargo test --test blobs runs the cases against the filesystem store; cargo test --test blobs_s3 runs the same cases against an S3-compatible endpoint (FLOCKFS_TEST_S3_ENDPOINT,
default http://127.0.0.1:59000, bucket flockfs-test, keys flockfstest /
flockfstest-secret-123) and skips, saying so, when nothing listens there. A local MinIO:
docker run -d --name flockfs-blobs-minio -p 127.0.0.1:59000:9000 \
-e MINIO_ROOT_USER=flockfstest -e MINIO_ROOT_PASSWORD=flockfstest-secret-123 pgsty/minio server /data
docker run --rm --network host --entrypoint sh pgsty/mc -c \
'mc alias set l http://127.0.0.1:59000 flockfstest flockfstest-secret-123 && mc mb l/flockfs-test'FLOCKFS_TEST_BIG_MB (default 200) sizes the constant-memory round trip.