Storage contract¶
For people running an obsync server.
Dated 2026-09-07. The design decisions behind it: one StorageClass per
storage implementation, a class that other workloads may share, and -- for a
deployment that accepts the risk -- a single copy without backups. The server
therefore assumes nothing about the class behind a path; it assumes only a
POSIX directory that honors fsync. Every size, class name and host path on
this page is guidance for a deployer to size against their own disk, never a
description of any particular installation.
Volumes and roles¶
| Role | Variable | Contents | Typical claim |
|---|---|---|---|
| blobs | OBSYNC_BLOBS_DIR |
ciphertext chunks | a local class, sized to the vault and its history |
| journal | OBSYNC_JOURNAL_DIR |
journal segments, index snapshots, server key | a local class, at least the watermark floor below |
| mirror | OBSYNC_BLOBS_MIRRORS |
optional extra blob copies | any class |
The chart ships a default class name in chart/values.yaml; it is a name
from one cluster and every deployer replaces it with a class their own
cluster offers.
The chart exposes storage.blobs.{className,size},
storage.journal.{className,size}, and storage.mirrors[] with the same
fields. A future HDD class or NAS-backed class is a values change, never a
code change.
On-disk layout¶
<blobs>/v1 chunk volume root, mode 0700
<blobs>/v1/<sid[0..2]>/<sid[2..4]>/<sid> one chunk, mode 0600
<blobs>/v1/tmp/<random> in-flight upload, renamed on success
<journal>/v1 journal volume root, mode 0700
<journal>/v1/journal/<000001>.log append-only segments, 64 MiB each
<journal>/v1/index/<seq>.snap index snapshot, as the journal grows
<journal>/v1/nonces accepted request nonces, 0600
<journal>/v1/server.key only when OBSYNC_SERVER_KEY is unset, 0600
<journal>/v1/setup-token first-boot and recovery login, 0600
<journal>/v1/quarantine/<sid> chunks that failed a scrub
<journal>/v1/quarantine/.tmp-<unique> interrupted quarantine copies
Volume posture¶
Modes above are enforced on every start, not only at creation. A restored
snapshot, a tar -x, a docker cp, or a bind mount hands the server volumes
it did not create, and a server.key or setup-token that arrives readable
by every account on the host is a standing way in: the token is the dashboard
recovery login, and the key with the journal unwraps every stored device
credential.
serve, check, and export therefore run one pass over six classes —
journal_mount and blobs_mount (the directory each configured volume path
names, one per mirror too, and every directory above it), blobs_root (the
primary and each mirror), journal_root, server_key, setup_token —
before anything is read or written through them, and a store cannot be
opened without the completed pass in hand. Every decision is made on an
open handle or on a directory chain, never on a bare name:
- Who may rename a root away (
*_mount), decided before a byte is written. Everything below a root works by name, and a name is only as good as the directories that hold it: whoever can renamev1away can put another tree under the name after the pass, and a later snapshot would write that tree's state into the protected root. So every configured volume path is validated first — absolute, made only of plain names, and free of links down to the last name that exists (not_canonical), because a link is re-pointed by its owner and a link nested inside another link's target is one no walk of the written path would see; nothing behind a refused path is read, created, or removed. Then every directory that already exists on the path, from the filesystem root down, must be owned by root or by the server's user (foreign_owner), and must not be writable by others (writable_by_others) or by its group (writable_by_group): a group number says nothing about who is in the group, andfsGroupexists to share one. The sticky bit satisfies the write conditions, since it narrows rename to the entry's owner, the directory's owner, and root, all of which are already root or the server. A chain that would be refused once complete is refused before anything is created beneath it, so a refused start leaves nothing behind. On Linux, the shipped platform, nothing at all is written before the judging: the user the server runs as is read off/proc/self, which the kernel owns by the process's effective user. On other systems (development only) the user is learned from a uniquely named empty probe file in the deepest existing directory of the journal path, removed at once; a probe a crash leaves behind is a harmless empty file under.posture-probe-and is never swept, since a sweep is a removal by name. A refusal states how many directories up it was found (depth=0is the configured directory itself); it never states a location.
The provisioning precondition. The server takes ownership of
nothing. When there is anything left to create — the configured
directory's tail, or v1 inside it — the deepest directory that exists
must be owned by the server's user and writable by it, or the start is
refused as unwritable. A named Docker volume satisfies this (the mount
point is copied from the image, owned by the server's user), and so does
a hostPath or static local volume the operator created for that user
(65532:65532, mode 0700). A dynamic provisioner that presents a
root-owned 0755 volume root, or a world-writable one, does not: prepare
the backing directory once as the node administrator (chown
65532:65532 and chmod 0700), then start the server. A volume that
already holds the server's v1 under a root-owned mount point is
accepted, since nothing needs creating. The fix for a refusal is always
to give the directory to root or to the server's user and close it,
never to widen the pass. The image smoke's eighth property presents
root-owned 0755 volumes holding no root and requires the unwritable
refusal.
2. Look, open, compare (roots and credential files). The name is
looked at once without following a link (lstat): a link, a type the
server never stores, or a foreign owner is refused before anything is
opened. The name is then opened read-only and non-blocking (a fifo
answers instead of waiting for a writer), and the handle must be the
inode that look saw, or the pass refuses (swapped): whatever the name
is made to say between the look and the open, what the server holds is
the file it looked at. The one thing the look leaves standing that
cannot be opened — a file of the server's own user at a mode that shuts
its user out — is corrected by name only so that a handle can be reached
at all. Every decision from here on is made on that handle. No
O_NOFOLLOW is involved: its value differs between architectures, and
the identity check does not.
3. Type. On the handle. A directory where a file belongs, a file where
a root belongs, or anything the server never stores refuses the start.
4. Owner. On the handle. The file's user must be the user the process
runs as, learned from a file the process creates on the journal volume
and then removes. The process cannot chown, and whoever does own a
file can widen it again, so a foreign owner is a refusal and never a
correction.
5. Mode. On the handle. Roots must be exactly 0700 and credential files
exactly 0600. Anything else is corrected with one fchmod and then
RE-READ off the same handle: a mode that was set is not a mode that
stuck. A correction the volume ignored, or one it refused, refuses the
start.
6. Read through the handle. server.key and setup-token are read
through the handle they were measured on, and a file the server has just
written is measured on the handle it was written through and read back
through it. No name is consulted again after a measurement. Two
credential classes that resolve to one inode (a hard link) are refused,
and a setup-token that does not hold 64 hex characters is refused
rather than minted over: a file that is not a token is not a first boot.
Each decision is one line — event=posture path_class=<class>
decision=<ok|repaired|refused> … — and a correction states from and to
in octal. The modes on the server_key and setup_token_ready lines are the
modes read back off the handle, so a startup line cannot claim a protection a
file does not have. No line carries a filesystem location or any file content.
obsyncd check runs the identical pass, and its policy is repair and
report: an operator who runs it on a restored volume leaves that volume
correct, and the report names every class with its decision
(posture setup_token: repaired 0644 -> 0600; a mount is reported with the
mode it was read at and is never corrected, since it is not the server's to
change). A posture that cannot be corrected exits non-zero, exactly as it
refuses a start.
obsyncd setup-token runs that same pass and then reads the setup token
through the handle the pass measured, printing it on standard output alone
(docs/recovery.md, "Reading the setup token"). It opens no store, so it
takes no journal lock: it is the one verb that answers while serve holds
the journal, which is what makes kubectl exec … -- obsyncd setup-token a
read an operator can perform on a pod that is serving. Every refusal the
pass makes on a credential file it makes here too, and a file that does not
hold 64 hex characters is refused rather than printed.
The chart sets no fsGroup. It is a group-sharing mechanism: the kubelet
would make the mount point and everything under it writable by that group,
and a mount point a group may write is refused above. The image ships
/data/blobs and /data/journal owned by the server's user, and the host
directories behind a local volume are created for that user
(docs/kubernetes.md), so nothing needs sharing. A platform that applies an fsGroup anyway sees the
pass correct the bits below the roots and refuse the mount point; the fix is
to drop the fsGroup, not to widen the pass.
v1/nonces is not one of the six classes and adds no line to the report.
The one thing measured about it is that the name is not a link and not
another type: a restored volume can arrive with it pointing at a file the
server may write, and an append through it would put nonce lines inside
that file. A start refuses one and says why.
It holds no credential: a device id and a nonce are public request values
that open nothing and are only ever compared, so it is server state like the
journal segments beside it, created 0600 under the same measured 0700 root.
What it does hold is the replay window (docs/protocol.md,
"Authentication"): every accepted nonce is appended and fsynced before its
request is answered, a start loads back what the 600 s still covers, the
file is rewritten when it passes twice the cache's ceiling, and a torn final
line costs only itself. Requests that arrive together share one fsync: their
nonces are written as one batch while no lock is held, and each request is
answered only once the batch holding its nonce is durable. A nonce still in
flight is already a replay, and counts against its device's share. A request
refused for the cache's ceiling or for its device's share is refused before
its nonce joins a batch: it holds nothing and costs no fsync.
Nonce log recovery¶
What an operator can rely on after a crash, and what the next start does with what the crash left.
Every accepted nonce is appended to v1/nonces and fsynced before the
request that carried it is answered. A nonce the server has acted on is a
nonce the server has already written down.
The file is rewritten when it passes twice the cache's ceiling. The rewrite is one sequence, in this order:
v1/nonces.tmpis removed by name. A removal by name never follows a link, so a link a restored volume brought is unlinked rather than written through.- The replacement is created at that name exclusively.
O_EXCLrefuses a name of any kind that already exists, so nothing this writes can land in a file that was already there. - The entries the window still covers are written to it.
- The replacement is fsynced.
- It is renamed onto
v1/nonces. - The journal root is fsynced.
The handle that wrote the replacement is the handle that appends to it afterwards. No name is resolved again after step 2.
A crash before step 5 leaves a partial v1/nonces.tmp behind. The next
start ignores it. The live window is read from v1/nonces alone, and
nothing standing at the temporary name is ever loaded, whatever it holds.
The next compaction removes it at step 1 and takes the name back.
A crash after step 5 leaves the replacement standing as the log, and that is the file the next start reads. Step 6 is the only step such a crash can lose, and losing it costs nothing the window promised. The entries the replacement holds were fsynced at step 4. A rename the filesystem has not yet committed can only leave the name on the file the replacement was built from, and every entry in the replacement was appended to that file before it entered the window, so either file answers the window.
The rewrite holds what was durable before the batch that triggered it, and
that batch is appended after it. Any failure inside the sequence, or in the
append, refuses every request in the batch and logs one event=nonce_log
decision=refused batch=<n> line. A volume with no room is 507
storage_full, the code the store gives the same disk. Anything else is 503
nonce_log_unavailable. Nothing already durable changes: v1/nonces holds
what it held, a refused append is cut back off the file before anything else
is written, and the nonces the refused requests carried were never recorded,
so each is unspent and the device may send it again. A signed READ is refused
too, and must be: its nonce is recorded like any other, and a read answered
with its nonce unrecorded is one a crash makes replayable. A cut that itself
fails leaves the log refusing every request, 503 nonce_log_faulted, until
a restart truncates the torn tail. The flush that finds it logs one
event=nonce_log decision=faulted rollback_io=<kind> line, and /readyz
answers 503 not_ready with nonce log faulted; restart to recover until
then. The
compaction threshold is still outstanding, so the next batch attempts the
rewrite again. A flush that panics, which only a bug does, settles the same
way: its whole batch is answered 503 nonce_log_unavailable, cut back and
unspent, with one event=nonce_log decision=refused reason=flush_panicked
batch=<n> line, and no request is left waiting on it.
The boundary is the one "One writer" states below. These steps defend
against what a crash and a restored volume leave behind. A process
already holding the server's own uid inside the 0700 journal root is held
out by none of them: it can write v1/nonces directly, exactly as it can
write the journal segments beside it. Keeping such a process off the
volume is the platform's admission decision, not this file's.
Measured. crates/obsyncd/src/api/auth.rs pins each paragraph above with
a test that drives a real compaction over a real volume: a partial
temporary file present at start, a directory standing at the temporary
name, the state step 5 leaves before step 6, and a replacement that
cannot be created because the root is read-only.
Not measured. Nothing here is tested by cutting power or by making a
syscall that succeeded report a failure. The states a crash would leave
are built by hand and then opened. The ordering claim — that a file
fsynced before its rename is on the volume once that rename is visible —
rests on the POSIX fsync contract and on the class honoring it, which
this document assumes of every class.
One writer¶
The journal has exactly one writer, and the access mode is not what makes
that true. ReadWriteOnce keeps other nodes off a volume and nothing more:
Kubernetes lets a second pod on the same node mount it, and this is a
single-node cluster. So obsyncd holds an exclusive advisory lock on
v1/lock of the journal volume for the life of the store, taken before a
byte of the journal is read: a second obsyncd on the same volumes — a
second pod, a rolling surge, a check or export while serve runs —
refuses to start with event=store_open decision=refused reason=journal_locked
rather than share the journal. The lock goes with the process, so a crash
leaves nothing to clean, and a stopped server frees it at once. It is an
advisory lock, and its boundary is stated exactly: it refuses cooperative
duplicate starts of this server. A process running as the server's own
user is the server by every test the filesystem offers — owning v1 is
the authority to rename it and lock a fresh v1/lock on a new inode — so
an arbitrary Pod that could mount the claim and run as that user is not
excluded by any file mode; keeping such a Pod off the claim is the
platform's admission decision (who may create Pods mounting these claims),
together with the chart's replicas: 1 and strategy: Recreate. The
volume-directory conditions above constrain other accounts, not that one. Run check
and export with the server stopped. The chart's replicas: 1 and
strategy: Recreate are the rendered half of the same boundary, so a rollout
never asks for a second writer; ReadWriteOncePod is not available on the
non-CSI local class and is not relied on. The image smoke's seventh property
starts a second container on the same volumes while the first serves and
requires the refusal.
Durability rules¶
- Chunk write: stream to
tmp, hashing; on completionfsync(file),renameinto place,fsync(dir); only then respond. A crash leaves either the complete chunk or atmpfile that startup removes. Each of the three points refuses differently and leaves a different residue: a refusal while streaming removes the temp on the way out; a refusal at thefsyncleaves an unsynced temp; a refusal at therenameleaves a synced one. Both leftovers are removed at the next start, which counts them on itsstore_openSUMMARY, and in every case the chunk is simply absent and the client re-uploads. A fan-out directory therenameneeds (v1/<ab>/,v1/<ab>/<cd>/) is created first, top down, and each new one's parent is fsynced before going deeper, so every directory on an acknowledged chunk's path is durable (1.1.5, #273). A directory whose parent'sfsyncfailed is remembered, and the next chunk under it fsyncs that parent again. A start that finds a leftover temp fsyncsv1/and each first-level directory once (fanout_synced=on the same SUMMARY): the cut may have fallen between a directory's creation and its parent'sfsync. The temp goes only after that repair succeeds, so a start that fails or stops part way leaves it, and the next start repairs again. Every start fsyncs the volume andv1/, which namev1/andv1/tmp/, whether or not it created them. On a fresh store most early chunks open a new leaf, so they cost one more flush (two for the first of each 256 first levels). - Journal append: frame =
u32 len | u32 crc32 | payload;write,fsync(segment); only then respond. Startup replay stops at the first torn or CRC-failing frame, truncates the segment there, and logs the count of frames recovered. Version posts that arrive while anfsyncis in flight are appended together and share the next one (group commit, 1.1.5): at most 64 posts or 4 MiB of manifests a turn, one post per file, each checked against the index as it stands and answered only after thatfsyncreturns and its frames are applied. A turn the watermark or the volume refuses is rolled back whole under rule 3, and each of its posts is then tried alone, as it would have been without the turn: refused only when its own frames do not fit or its own write fails. Eachversion_appendline says how many posts itsfsynccarried (batch=). - Failed journal append: the journal records the length it has made durable
before it writes, and any failure of the write or of its
fsynccuts the segment back to that length and fsyncs the cut before the error returns. Nothing was acknowledged, and the next frame starts clean. This is not a nicety: segments are openedO_APPEND, so without the rollback the next successful frame would land AFTER a torn one, and the next start would truncate at the torn frame and discard every write acknowledged since. Each failure logs one line,event=journal_append_failed decision=truncated io=<kind> segment=<n> torn_bytes=<n>— the error's kind, never its message, which can carry a path. - Faulted journal: if that rollback ITSELF fails, at the truncation or at
its
fsync, the journal is faulted. The line saysdecision=faultedand addsrollback_io=<kind>, and the refusal carries both kinds. The second kind is a second fact and the state's actual cause: a truncation refused by a full volume and one refused by a read-only mount both read asfaultedand need different repairs. From then on every append refuses withjournal_faultedwithout touching the volume,/readyzanswers503 not_readywithjournal faulted; restart to replay, and the state clears only at the next start, which replays and truncates the tail as rule 2 describes. Chunk uploads are unaffected: the blob volume is its own record. - Snapshot: written to
tmp, fsynced, renamed; replay starts from the newest valid snapshot and applies later frames. One is written once the journal has grown 16 MiB since the last, and by at least that snapshot's own size, so an idle journal is not snapshotted again and every start, after a stop or a crash alike, replays at most about one snapshot's worth of frames; a stop writes one only when one is due. The index is copied under its lock and the copy is encoded and written with no lock held, so writes and reads go on while it lands; its size counts toward the journal volume from the moment it is admitted. - No write is acknowledged before it is durable. This is not configurable (AGENTS.md requirement 4).
Journal frames¶
account (setup, and every change to the account's recovery verifier:
its registration, and the operator's obsyncd recovery reset apply, which
writes the account again without one and with recovery_cleared_at, the
reset's time, arming one re-enrolment until a verifier is registered),
device (create, update, activate,
revoke, delete, wrap),
version (which carries its file's domain_id, so replay reaches the same
domain the post named), gc (a list of sids collected), scrub (a step
that found a mismatch, or completed a pass), seen (device sign-in and edit
events, retention-bounded; an accepted version post appends its version
frame and its seen edit frame together, with one fsync, which posts queued
behind the previous fsync share; durability rule 2). There
is no domain frame: a domain exists because a file record names it
(docs/architecture.md 5.1 item 4). Pairings live in memory only, so a
start destroys every pending device no pairing is holding any more, through
the device delete frame expiry uses (docs/architecture.md 4.2). Frames
carry account_id.
The registration time (1.1.5). Since 1.1.5 the account frame and the
snapshot's account carry recovery_at, the Unix milliseconds at which the
verifier in recovery was registered, which the last-device rule reads
(docs/protocol.md, "Devices"). Both directions of a version change load:
- Older state on 1.1.5. A frame or snapshot without
recovery_atreads as a verifier registered before times were kept, which keeps the older rule. A time that is not a number is refused as corrupt rather than read as absent, because absent is the permissive reading. - 1.1.5 state on an older server. A server before 1.1.5 reads an account's
members by name, so it ignores
recovery_at, loads every 1.1.5 frame, and applies no hold. A reset is an ordinaryaccountframe without a verifier, which it reads the same way. A snapshot it writes dropsrecovery_at, so a verifier that passes through one reads as registered before 1.1.5 when a 1.1.5 server starts on that journal again. - The reset's arm.
recovery_cleared_atis read only beside no verifier; beside one it describes nothing, and a value that is not a number is refused as corrupt. A server before 1.1.5 ignores it and refuses recovery on an account with no verifier, and a snapshot it writes drops it, so the arm fails closed: the account answersrecovery_unavailableuntil the next reset.
Memory¶
The index lives in memory: the account, the devices, and every retained
version of every file with its encrypted manifest. Chunk bodies never do.
So memory grows with retained history, not with the size of the vault:
about 1.3 KiB for each retained version of a small note, plus about 1 MiB.
A version of a larger file costs more, because its manifest lists more
chunks. Retention bounds history (OBSYNC_RETENTION_VERSIONS per file,
kept at least OBSYNC_RETENTION_DAYS), so a vault edited for years holds
what retention keeps, not everything it ever was.
A start rebuilds the index from the newest snapshot and the journal after it. The snapshot is read one file at a time, with its CRC computed as it streams, so no copy of the whole snapshot is ever in memory, and a start peaks at about what the server holds at rest. Measured on linux/arm64 (2026-09-26), small notes posted through the API, resident memory in MiB:
| Retained versions | At rest | Start, peak | Start, peak before 1.1.4 |
|---|---|---|---|
| 1 000 | 3 | 4 | 4 |
| 10 000 | 15 | 14 | 14 |
| 50 000 | 65 | 62 | 146 |
| 100 000 | 131 | 115 | 285 |
| 200 000 | 238 | 221 | 1043 |
The same run on linux/amd64, to 50 000 versions, grew by the same 1.3 KiB per version, and its start peaked at half what it did before 1.1.4. Before 1.1.4, a start parsed the whole snapshot at once. It needed about four times the index, which put 200 000 versions past a 1 GiB limit at every start.
Integrity¶
- Every upload is verified against its
sidwhile streaming. - The scrub thread re-hashes blobs one pass at a time in sid order, repairs
a mismatch from a healthy mirror when available, and otherwise preserves
it in
quarantine/using the sequence below. It reads no faster thanOBSYNC_SCRUB_RATEand works no more than one minute in each hour: a step that tooktrests at least 59t, because on a store of small notes a chunk costs a file open, not its bytes, and a byte rate alone let the scrub work most of every minute. A pass begins no sooner than a day after the last one began, and at once when the dashboard asks for one: a bad chunk is only worth finding while a mirror or a device can still repair it, and the shortest window this contract gives either is a day (OBSYNC_RETENTION_DAYSis at least 1, a newborn chunk is protected for 24 h). A store the scrub cannot walk in a day at that pace is scrubbed continuously at it. The day and the minute are constants, not settings. A step journals only a mismatch and what became of it, or a completed pass; each walk logs one START and one SUMMARY, whoseworked_msbeside itsduration_msshows the pace. - Every read verifies size; the client verifies the plaintext hash from the manifest after decryption, so a corrupted chunk can never be written into a vault.
Quarantine works when the blob and journal roots are separate mounted volumes. Under the chunk's SID lock and the journal guard, the server opens the primary, corrects its indexed length to the observed length, surveys journal usage, and checks room for a complete additional copy above the journal watermark. An existing quarantine copy still counts during this check; replacing it is not a credit against peak space. Capacity remains declared capacity minus tracked bytes, not a physical free-space probe. If the source cannot be opened or measured before mutation, the scrub logs the failure and retains its last observed inventory for retry. That value is not a fresh filesystem measurement; this repair does not add a new blob usage survey or readiness state.
The primary is copied to an exclusive .tmp-<unique> file inside the
quarantine directory. The copy is fsynced, renamed to <sid> within that
same directory, and both the quarantine directory and its parent are
fsynced. Only then is the primary unlinked and its parent fsynced. Mirrors
are left alone. Every failed stage logs its error kind. Before a durable
destination exists, failure preserves the primary. Failed temporary cleanup
leaves counted residue; startup removes only regular quarantine temporary
files, syncs the directory and logs the number removed.
A failed quarantine never appears in that step's quarantined list. If
the primary remains, its inventory and observed bytes remain, and a later
scrub retries. Failure after unlink but before its directory fsync retains
conservative inventory; the next scrub or a recovery upload must sync the
absence before forgetting that entry. An upload cannot return existed
for an absent primary. On restart, the primary-volume scan rebuilds chunk
inventory; it counts both copies when both remain. A retained primary can
need extra destination headroom for a retry even when an earlier complete
copy already exists in quarantine.
After durable removal, inventory is updated before the SID lock is released, even if the later summary frame cannot be appended. The frame format is unchanged: scrub summaries are historical reports and do not remove a subsequent reupload from inventory. Old recorded summaries are preserved as recorded claims; they are not fresh evidence that files were moved. An operator can reupload the verified original ciphertext after quarantine; the normal upload verification still applies. Unchanged client files do not automatically trigger that reupload.
Uploads and scrub use a fixed table of 256 SID locks before taking journal
or index locks. Streaming and hashing hold no journal or index lock. A
hash-prefix collision can serialize unrelated uploads; it does not change
their identities or admission rules. GC tries the whole table once in
order, releases every acquired lock and logs gc_skipped
decision=chunks_busy if any is busy, then retries on its next scheduled run.
It holds the table, the journal and the index only to decide and journal a
collection; it then unlinks each chunk under that chunk's lock alone, and
leaves one that was uploaded again since the frame (counted as reuploaded
on its SUMMARY).
Garbage collection¶
Refcounts come from the index (sids referenced by retained versions). A
chunk is collectable when no retained version references it, it is older
than 24 h (protects an upload whose version post has not landed), and the
retention window has passed for the version that last referenced it.
Retention keeps at least OBSYNC_RETENTION_VERSIONS versions per file and
everything younger than OBSYNC_RETENTION_DAYS. GC runs hourly, logs a
SUMMARY with counts and bytes, and needs no device ceremony.
Free-space watermark and quota¶
Free space is declared capacity minus tracked usage: the standard library
exposes no filesystem statistics, so OBSYNC_BLOBS_CAPACITY and
OBSYNC_JOURNAL_CAPACITY are required and the chart sets them from the
claim sizes. Writes are refused with 507 when free space on the blob
volume is below the larger of OBSYNC_FREE_WATERMARK's two terms, or when
the account's quota is exceeded. The dashboard shows both thresholds and the current
values. Each declared capacity must exceed its own watermark: at or below it
every write would be refused from the first, so the server refuses to start
and names the variable.
A chunk upload reserves what its check admits (server 1.1.5, issue #301). Its
declared length is measured against the tracked usage PLUS every upload that
has passed its own check and is not yet counted, under the one lock that
checks. The admitted upload then holds its bytes until they are counted. A
refused upload holds nothing. An upload that does not land (a body that does
not verify, a write the volume refuses, a panic) gives its bytes back on the
way out. Before, uploads of different chunks that arrived together were each
measured against the same total and all passed: an 8 MiB volume with a 1 MiB
reserve stored a 12.6 MB note and ran down to nothing free. The refusal's free
and used count those reserved bytes, so a volume_full line can name less
free space than the dashboard shows while uploads are in flight.
A refused chunk is read to its end before the 507 is sent (server 1.1.5,
issue #304). The watermark and the quota refuse before a byte of the chunk is
read, and the HTTP layer drains at most 1 MiB of a body a handler left
unread. Up to 1.1.4 a larger refused chunk was answered over the rest of its
upload and the connection closed on it. A proxy or tunnel still writing that
upload was reset, and it answered the device with a bare 502, which the
device reads as the server being unreachable, not full. The rest of the body
is read only for a request that proved a device's credential, and only up to
its declared length, which the chunk limit (8 MiB plus the 16-byte tag)
already bounds. A sender below the rate floor is still ended as 503
slow_body.
The journal volume has the same watermark applied to its own capacity, and
refuses a frame that would take it below with 507 journal_full before the
volume is asked. Tracked journal usage is everything under the journal root —
the open segment's durable length plus every other file on the volume: the
other segments, the index snapshots, the quarantine, the nonce log, and the
small fixed files. Snapshots are why this is not a segment count: they rest on
the same volume and reach tens of megabytes for a large vault. VolumeStatus
reports the same number, so a refusal and the dashboard never disagree about
how full the volume is, at any moment rather than at the last roll. The two
volumes have separate refusal codes, volume_full and journal_full, because
they are provisioned and filled independently and a refusal that did not say
which one it measured would send an operator to the wrong disk. A journal
volume that fills anyway, below the declared capacity, is rule 3 above: the
append is refused, rolled back, and never acknowledged.
A volume the FILESYSTEM fills first -- a declared capacity larger than the
disk, or a disk something else filled -- answers 507 storage_full on
whichever volume ran out (ENOSPC, or EDQUOT from a filesystem quota), so a
device reads a full server rather than a fault it retries as absence. A
refused chunk leaves nothing a start would not remove (rule 1). A refused
journal append whose rollback succeeds leaves the journal unfaulted: the
refused frame is cut away, and the next append lands once there is room. Only
a rollback that fails as well faults the journal; that append still answers
with the volume's own refusal, and every write after it is 503
journal_faulted until a restart.
Where the measure runs¶
The append path may not walk a directory, and everything that writes to this volume between two walks would otherwise be invisible to the refusal. So the total is four numbers, each owned by exactly one writer and each ABSOLUTE rather than a delta, so they cannot double-count and a missed update cannot accumulate: the open segment's durable length, a survey of everything the other three do not own, the quarantine, and the nonce log.
Every writer of this volume, what it leaves behind when it fails, and how the number stays right. Two properties, never one: the residue is ACCOUNTED FOR, and the original durability error is still RETURNED. A failing path that quietly balanced the books would be worse than one that did not.
| Writes | When | Residue on failure | Accounted by |
|---|---|---|---|
journal/<n>.log, the open segment |
every journalled write | a torn tail, rolled back in the same call; kept for the next replay if the rollback also fails | the open segment's durable length, advanced only after a successful fsync |
journal/<n>.log, a new segment (roll) |
first append after a start or replay, and at 64 MiB | an empty segment | the survey, re-run by the roll itself, which then excludes the segment it opened |
| the last segment, truncated (replay) | every start | none: it removes bytes | the survey, re-run before replay returns |
index/<seq>.tmp → <seq>.snap |
each snapshot | a .tmp a failed write or rename left, which nothing later removes |
the snapshot's own size from its admission, while it is written with the journal unlocked; then the survey, re-run by the prune a successful call ends in AND on every failing exit, before the original error is returned |
index/<seq>.snap and covered segments, removed (prune) |
end of each snapshot | whatever was removed before one removal failed | the survey, re-run at the end and on every failing exit, before the original error is returned |
quarantine/<sid> and .tmp-<unique> |
a scrub mismatch no mirror can repair | partial temporary bytes, a published copy beside the retained primary, or a durable copy after primary removal | a complete survey before admission and on every operation exit, under the same journal guard as the copy and removal; a refused survey marks usage unverified |
nonces |
every batch of authenticated requests | a partial batch from a short write, cut back at once; kept only if that cut fails | the log publishes the absolute size of both its names after every write, the failing ones included |
nonces.tmp → nonces (compaction) |
when the log passes twice the nonce ceiling | a nonces.tmp a failed compaction left |
the same publish, which counts the temporary BY NAME so that leftover is seen |
server.key |
first boot | a partial key refuses the start | the survey at Journal::open, which runs after it |
setup-token |
first boot | a partial token refuses the start | a survey cli::serve runs after it, being the last write the volume takes before the server serves |
lock |
every start | none: it is empty | the survey at Journal::open, which runs after it |
The survey walks journal entries, snapshots and quarantine residue, never a
vault, and does not follow symlinks. Quarantine holds the journal guard
across peak-space admission, copy, removal and the final survey. No watermark
reader or other survey can observe the operation between its file changes
and accounting. On any refused survey, usage stays at its last measured
value with usage_unverified=true; unreadable metadata is never zero bytes.
The append path retries that survey and refuses admission if it still fails.
When the survey itself fails¶
An accounted residue is only as good as the walk that measured it, and a walk
can be refused. A survey that fails is a DIFFERENT fact from a survey that
succeeded, and the journal records it as one: unverified holds the
io::ErrorKind that refused the walk, set on failure and cleared only by a
COMPLETE later one. Nothing partial is ever published — the two walks the
survey makes are stored together or not at all, because one fresh number
beside one stale one is a total that was never true of the volume at any
instant.
While it is set the tracked total is known to be stale, so admission is
fail-closed. The next append retries the survey once: if that succeeds, the
state clears and the watermark is applied to the total it just read; if it
fails, the frame is refused with 503 journal_unverified and nothing reaches
the volume. Readiness retries it too, which is what makes the recovery visible
without a write: /readyz answers 503 not_ready with journal usage
unverified; survey failed: <kind> while it stands, and 200 on the first
probe after the volume can be walked again. The verdict cache bounds how often
an unauthenticated prober can make it walk. VolumeStatus carries
usage_unverified beside bytes_used, so a dashboard shows the figure as the
last one read successfully rather than as a current one.
Two states, and the distinction matters to an operator: journal_faulted is
about the segment's CONTENTS — bytes no frame owns — and clears only at a
restart that replays and truncates; journal_unverified is about the
ACCOUNTING and clears the moment a walk succeeds. A journal that is both stays
faulted: no survey can speak to a torn tail. Each transition says so once,
event=journal_survey_failed io=<kind> at=<where> and
event=journal_survey_recovered by=<where>, and never on the retries in
between.
The original operation error is still what its caller gets, in every case above. The accounting is a consequence of the failure, never a replacement for reporting it.
Replication and propagation (design hooks, phased)¶
- Mirrors (v0.1): every chunk write goes to the primary and each mirror before the response; reads come from the primary; scrub cross-checks mirrors and repairs a bad copy from a good one. This gives a second copy on a different class or drive on the same node today.
- Replica server (v0.3): a second
obsyncdin replica mode follows the primary's change feed and fetches chunks, giving a warm copy on another node. Promotion is an operator action. - Export (v0.1):
obsyncd export --domainwrites that domain's stored CIPHERTEXT, which the operator decrypts on a device holding the key; the server implements no AES and never could write plaintext (docs/architecture.md3 and 5.1). Its payload assembles a selected newest head; retained-version metadata in the manifest is not a backup of every historical payload. The key it accepts (--key-file) does not decrypt content. Both export and check are offline recovery tools, not a full restore proof.
Offline check and recovery verdicts¶
obsyncd check hashes the primary copy of every inventoried chunk and every
SID referenced by a retained version, including older non-head versions.
It deduplicates SIDs, so repeated references count once, and counts verified
bytes from the actual hash read. A missing referenced chunk fails even when
the volume scan cannot inventory it. Corrupt chunks fail; other read errors
refuse the command. The additional SID set is bounded by the distinct
inventory and retained references already described by the in-memory store;
chunk contents are streamed, never collected in that set.
Both commands print their report before exiting. Check exits zero only with
no failed chunks or frames; export exits zero only with no missing selected
chunks. An incomplete report prints result: FAILED, emits
cli_failed decision=integrity_failed, and exits 1. I/O or storage refusals
also exit 1; invalid configuration or arguments exit 2. Require successful
completion and result: ok, then compare counts and recovered state with the
backup baseline. An empty, internally consistent store can legitimately
return zero; that does not prove that the expected backup was restored.
Run either command with the server stopped and on a dedicated restored copy,
preserving the pristine backup. Opening storage enforces posture, creates
missing layout or a volume-backed server key, removes temporary leftovers,
may fall back to an older snapshot, and truncates a torn final journal tail
before checking the remaining frames. Check is not a read-only forensic scan.
For a lossless restore drill, separately reconcile store_open sequence,
skipped snapshots, truncated bytes, removed temps and strays against the
baseline; result: ok does not waive unexpected recovery changes.
Back up both complete volume trees from one quiescent application state or
coordinated filesystem snapshot, including journal segments, snapshots, nonce
log and recovery credentials. When OBSYNC_SERVER_KEY is supplied externally,
the server does not persist it in server.key; preserve that same external
key through the operator's protected backup custody. Hash/CRC checking does
not verify this key. Full existing-device recovery also needs an authenticated
read using the restored device state, and plaintext recovery needs the client
vault key or recovery phrase. The server key cannot decrypt vault contents.
Encryption at rest¶
Chunks and manifests are already ciphertext. Device secrets are wrapped under the server key. Host-level disk encryption is a host decision outside this repository.
A local-volume deployment, end to end¶
Static local PersistentVolumes under a host directory per role --
/mnt/<disk>/obsync-blobs and /mnt/<disk>/obsync-journal is the shape
docs/kubernetes.md builds -- on a local class, Retain,
WaitForFirstConsumer, ReadWriteOnce (which excludes other nodes; the
one-writer boundary on the node is the server's own lock, above), node-affine
to the node that holds the disk, claimed by obsync-blobs and
obsync-journal in the namespace the chart is installed into. The claim names
come from the chart, which names every object for the application (obsync)
and never for the namespace it happens to be installed into;
scripts/ci/chart_pins.py refuses a name here that the render does not
create. Growing a volume is a PV capacity edit and a claim resize. A platform
with a storage-exposure policy has to admit the class, the provisioner and the
host root before any of this binds.