Speed: where the time goes, and what 1.1.5 changes — 2026-09-29¶
The user asked for speed ("we're using Rust, make it blazing fast"). This run
measures obsync end to end on real Obsidian at the 1.1.4 release (afbf7e7),
finds where the time goes, and checks the 1.1.5 server changes against it.
Setup¶
- Two desktop Obsidian 1.13.4 instances on macOS 27.0 (Apple M1 Max, ten
cores), each with its own profile and disposable vault, driven through their
DevTools ports; one loopback
obsyncd(plain HTTP,OBSYNC_EDGE=none). - The laptop was shared with seven other lanes of work: its load average ran from 25 to 240 during these runs. Absolute times below are therefore shapes, not floors; every comparison between two builds is interleaved runs under the same load, and the load-independent counts (requests, fsyncs per post) carry the weight.
- Fixture vault: 7,700 Markdown notes in 60 folders (median about 1.5 KiB) plus three 100 MiB attachments: 7,703 files, 317 MiB.
- Route comparisons: a signed load generator (a lab tool built on
obsync-core's signing, not shipped) holding N requests in flight against a fresh loopback server with four enrolled devices, 5 to 8 s a scenario, the two builds alternated and the first of each round swapped. - Builds: the 1.1.4 server and plugin; the 1.1.5 server from this branch with
the same plugin.
main.jsSHA-256, both runs:19d3202059551d72f53f3f3a0deaf3eb159971604c1d9fb481f756ac56c28d0c.
Typing on one desktop, reading on the other (1.1.4)¶
A token typed as trusted input at the end of an open note on A, twenty times, four seconds apart; times after the keystroke, p50 / p95 ms (p95: the 19th of 20):
| Step | p50 | p95 |
|---|---|---|
| Obsidian writes the note to A's disk (its own save delay) | 2013 | 2018 |
obsync on A settles the save (EDITOR_SETTLE_MS, 150 ms) |
2159 | 2179 |
PUT of the chunk (server: nonce, file and directory fsyncs) |
21 ms long | |
| version post (server: nonce and journal fsyncs) | 23 ms long | |
| B's waiting long poll answers (the moment the post is durable) | 2193 | 2225 |
| B fetches the chunk | 8 ms long | |
| the note on B's disk (decrypt, temp write, two fsyncs, rename) | 2240 | 2275 |
| B's open editor shows it | 2249 | 2288 |
Nine tenths of the 2.25 s is Obsidian's own two-second save delay before obsync can see the keystroke at all; obsync's settle adds 0.15 s; the server's three requests together take about 52 ms, the long poll wakes within a millisecond of the post, and B's write takes 26 ms.
First sync and Sync now (1.1.4)¶
A holds the fixture vault and is set up first; B is paired once A is done.
- Up: first push to last push 224 s (34 files/s), load 25 to 57; the
server spent 14.4 s of CPU. A push took p50 102 ms (p90 155): the chunk
PUT46 ms and the version post 29 ms as the plugin saw them, 41 and 26 ms inside the server. Pushes per second fell from 48 in the first thousand to 28 in the seventh while a push grew from 80 to 125 ms (p50), because each push waits for the whole state file to be rewritten, and that file grows by about 365 bytes a note. - A's own long poll answered 3,831 times while A made 7,771 posts: a quarter of the uploading device's requests are its own echo.
- Down: first pull to last pull on B 148 s (52 notes/s). Applying a note took p50 14 ms, 134 s in all; the server's part, 165 batched reads at p50 13 ms, was 2.7 s.
- Sync now with nothing changed, 7,700 notes examined: 4.0, 3.8 and 3.9 s on A; 4.2, 4.1 and 3.7 s on B.
Where the server's time goes¶
sample of the 1.1.4 server (symbols kept) for about 13 s under the plugin's
upload shape, a PUT and a version post per new note, four notes in
flight: of the request threads' blocked samples in named calls (40,708),
F_FULLFSYNC took 64 % (nonce log 8,313, blob file 6,314, blob directory
6,278, journal 5,240), write 23 % (the log 3,164, the nonce log 3,157,
the blob 1,716, the journal 1,305), mkdir of a new fan-out directory 5 %,
rename 4 %, opening the temporary file 3 %. The server's own code was
too small to rank: SHA-256 was the top frame in 24 samples.
The device serializes full flushes: a probe with N threads each appending
100 bytes to its own file and calling sync_all for 3 s made 142, 178,
220, 293 and 333 flushes/s at 1, 2, 4, 8 and 16 threads.
1.1.5 on real Obsidian¶
The same rigs against the server built from this branch at 107ded0 (the
plugin unchanged, main.js as above). The route-by-route and Linux
comparisons of the two servers are in
benchmarks.
The first attempt used the 7,703-file vault while the laptop's load average ran from 60 to 240. It was stopped after 35 minutes with 2,349 notes up: the saves of the plugin's state, not the server, set the pace (below). Up to then the server made 2,417 version posts durable with 2,140 journal flushes, 0.89 per post: 79 % of posts had a flush to themselves, 16 % shared one with another post, 5 % with two more.
The pass that completed used a 1,501-file vault made the same way (1,500 notes, one 100 MiB attachment, 103 MiB), load 80 to 160:
- Up: first push to last push 135 s (11 files/s). 1,569 posts (the notes and their folders), 1,394 journal flushes: 0.89 per post again.
- Down: 1,501 files on B in 40 s; applying them took 36 s of it (p50
20 ms a note). Every file on B matched A byte for byte (SHA-256 of all
1,501, outside
.obsidian). - Sync now with nothing changed, 1,500 notes: 2.1, 1.9, 1.6 s on A; 2.0, 1.7, 1.6 s on B.
- Typing, 20 rounds: B's editor showed the word p50 2,278 ms, p95 2,318 ms after the keystroke (1.1.4, lighter load: 2,249 and 2,288).
- Visual sweep, both desktops: no notices; the status item
idle; Sync now answered and left itidle; Show sync status listed 1,502 files tracked, nothing remote-only, the same feed sequence on both; the obsync settings tab rendered; no warning or error in either console during the sweep. B's sync status offers the recovery phrase, not yet confirmed there: expected for a paired device.
The plugin's state saves¶
Each push waits for the plugin to rewrite its whole data file, which grows
by about 365 bytes a note. With saveData wrapped to count calls and bytes:
- the completed pass: 1,829 saves for 1,501 notes, 478 MB of JSON, 66 s spent saving (longest 143 ms);
- the stopped 7,703-file pass: 1,262 saves for 1,550 notes, 720 MB of JSON by the time 2,349 notes were up, 380 s spent saving (longest 41 s, with the laptop starved of I/O).
At that rate a first sync of the whole 7,703-file vault writes on the order of 10 GB of state, and it is why pushes slow as the vault grows. The second round below changes it (#274).
Second round: #273, #274, #275¶
Builds from this branch: the server with #273 (9cd404b), the plugin with
274 (0ef6146, main.js¶
e8ee0a2ecb62fa01a3eac15833470678ae116ffaeed3a083a290605457956695). The
"before" of #274 is the same server with the 1.1.4 plugin, so the
comparison is the plugin's alone. Each rig ran with its own HOME (so
Obsidian's CLI socket was its own), its starter and Settings windows closed
after each step, and the rig not being measured minimized; the rig being
measured stayed shown without focus, for the reason in the next paragraph.
A minimized window. Mid-upload on the 1.1.5 plugin, rig A made 1.35
notes/s over 20 s minimized and 37.05 over the next 20 s shown without
focus (electronWindow.showInactive()), load 27. Minimized, its renderer
sat at 0.1 % CPU, setTimeout(0) took 0 to 224 ms and a vault stat 12
to 77 ms. A person who minimizes Obsidian during a first sync waits for
the window, not for obsync.
#274, first sync up, 7,703 files. With saveData wrapped to count:
the 1.1.4 plugin wrote its data file 7,451 times, 9.29 GB of JSON, and
spent 159 s writing it (longest 197 ms), first push to last 302 s, load 23
to 39. The 1.1.5 plugin, three runs: 211, 233 and 253 writes; 0.22, 0.23
and 0.27 GB; 7.7, 7.9 and 11.2 s writing; 233, 250 and 281 s, load 14 to
50. Pushes per second by thousand, 1.1.4: 24.8, 36.6, 28.4, 24.1, 25.7,
30.4, 20.5, 19.0; 1.1.5: 30.6 to 27.5, 33.6 to 23.2, 33.2 to 20.2. The
server made the 7,771 posts durable with 6,237 journal flushes before
(0.80 a post) and 7,363 to 7,455 after (0.95): with no save between them,
the posts arrive spread out and share a flush less often.
Downloads. B's first sync down: 140 s before; 161 s and 135 s after, each byte-identical to A (SHA-256 of all 7,703 files). The first download after the change, one run earlier, stalled: seconds in, B had applied 82 entries of its first 1,000-entry page, parked two of the 100 MiB attachments for the background lane, and staged one of them whole in its temporary file; the lane was waiting for the pull lock the page held, the page was waiting on something that made no request, no connection to the server was open, and the renderer was idle. It was stopped after 18 minutes. B's own log from before that moment was not captured in that run; the two runs after it logged B from before pairing and did not stall. The cause is open.
#273's cost. scripts/ci/bench.sh at smoke scale, images from this
branch before and after the change, two runs each, alternated, load 36 to
116: B1's fsyncs per note on a fresh store went from 4.88 and 4.92 to 5.88
and 5.93, and its wall time from 6.7 and 6.1 s to 7.0 and 5.6 s.
#275. The same runs used the corrected harness: two requests per note, as the plugin sends a note of at most 1 MiB.
Whole-app sweep at the head. A last pass on the 1,501-file vault: up in
50 s with 94 writes of the data file (9.5 MB), down in 25 s, byte-identical.
Then both rigs, shown without focus: no notices; idle before and after
Sync now; Show sync status with 1,501 files tracked, nothing remote-only
and the same feed sequence on both; the obsync settings tab rendered; no
warning or error in either console. B offers its recovery phrase, not yet
confirmed there, as a paired device does.
Third and fourth rounds: P10, #283, #288¶
Builds: the train at 9963bf0 ("before") and this branch ("after"), with
the same server binary; each rig has its own HOME, its starter and
Settings windows are closed after each step, and the rigs are quit at the
end of each run. Other lanes kept the laptop at load 5 to 57, and each
number below is given with the load it was measured at.
P10. Main-thread time per new-note push, measured as the event loop's active time over 300 pushes through the real push path over the test fakes (saves deferred, as the engine's drain defers them). Median of five rounds, interleaved three times, at load 18 to 20: 1.15 ms before and 0.38 ms after at 7,700 to 9,500 records; 0.43 and 0.50 ms at 0 to 1,800.
#283, the cause. Rig A alone, with no sync running, a timer probe in its page (median of up to ten samples per row):
- Minimized, a page timer of 10 ms or 100 ms took 994 to 1,021 ms, 1,000 ms took 2,000 ms, and ten chained 10 ms timers took 9.7 to 37 s. Hidden for more than five minutes, the chain managed only six steps in 21 s.
- obsync's worker timer of 100 ms took 295 to 304 ms minimized.
- With
electronWindow.webContents.setBackgroundThrottling(false), every row was back to its shown value. The page stayedhiddenand raised novisibilitychange. - Putting the throttling back brought the one-second wake-ups back.
- Each call took 0.4 to 0.6 ms.
#283, the upload and the download, in alternating 30 s windows, 7,723 files. Before, at load 5 to 8: up 30.4, 27.2 and 36.0 notes/s minimized against 37.1, 31.2 and 31.2 shown; down 33.5 and 31.2 against 64.1 and 56.7. After, at load 9 to 37: up 36.8, 31.6 and 22.2 against 37.9, 36.0 and 30.2; down 52.9, 26.4 and 24.7 against 39.2, 53.8 and 23.6. The upload's 27x slowdown of the second round (load 27) did not appear at load 5, while the timers' one-second wake-ups appear at any load. Both vaults ended byte-identical.
#283, a folder rename of 20 notes on A, timed until B's disk held all of them under the new name (B shown):
| Build, vault | A shown | A minimized 30 s | A minimized more than 5 min, then twice more |
|---|---|---|---|
| Before, 320 notes, load 11 to 27 | 0.5, 0.5, 0.7 s | 0.6, 2.1 s | 3.6, 1.2, 0.4 s |
| After, 320 notes, load 22 to 41 | 0.6, 0.4, 0.5 s | 0.4, 0.6 s | 0.7, 0.4, 0.4 s |
| After, 7,723 files, load 19 to 32 | 1.5, 1.2, 1.3 s | 1.4, 1.6 s | 1.3, 1.4, 1.4 s |
The 7,723-file run before the change lost both rigs at its minimized
renames, for a reason not found here, so it has no row. After the change,
the two rigs lifted and restored the throttling 18 times over the whole
7,723-file session: once per span of work. A's renderer used 16 ms of CPU
in a minute of idle, minimized, with the throttling restored, and the
same 16 ms with it lifted by hand. A window that was lifted before it was
minimized kept reporting itself visible after the restore.
#288, a device woken after a stall. This uses lane D's method on two
rigs paired on a five-note vault. Both windows were minimized. Both
renderers were stopped (SIGSTOP) 2 s into a long poll of A's, for 25 s,
then continued, and both windows were shown again. Each request's start
and end were recorded in the page, beside the server's log of each
request's end and duration.
- Before. The window's return dropped A's poll, which had waited
27.9 s (
dropped_poll=1). The quick read was answered in 30 ms. The new long poll left the plugin at +0.57 s and reached the server at +20.58 s, 20.01 s later. The server answered it at +75.59 s, but the client's 70 s budget ran out at +70.58 s:engine decision=offline reason=unanswered. A readoffline — retryingfrom +71.4 s to +127.1 s. - After, at the head. The poll was kept (
dropped_poll=0 read_beside=1 waited_ms=27767), and a read beside it was answered in 17 ms. The kept poll was answered on time at +27.8 s, and the next one went out at once and reached the server with no gap. There was noofflineon either rig, and both readidlefrom +1.35 s to the end of the 150 s sampled. - In an earlier run of the change. B's own poll ran out of its
budget while B was stopped. It was retried with
answered_meanwhile=1and was not called offline. The feed's stall watch then named the kept poll a stall, so a read beside the poll now restarts that watch. - Where the 20 s goes. On one rig, two signed 30 s polls were sent
200 ms apart. To the same URL, the second waited 20.00 s on the device
before it reached the server; it was answered 50 s after it left,
outside its 45 s budget. To URLs that differ only in
limit, there was no wait. The wait is per URL and on the device, for a request whose URL matches one still in flight. That fits Chromium's HTTP cache, which holds a second request for one URL behind the first. The same pair before the change logged a falseoffline; after, it loggedanswered_meanwhile=1. - A press or a focus, no stop. The same rigs, windows shown; three rounds of a Sync now press on A 8 s into a long poll, then a focus event 8 s into the next.
- Before: every one of the six dropped the poll (
dropped_poll=1), and the poll that replaced it waited 20.00 to 20.01 s on the device. - Before: five of those replacement polls ran out of their budget,
each followed by
engine decision=offline reason=unanswered, and eachofflinelasted until the next wake. - After: all six kept the poll and read beside it in 12 to 27 ms
(
dropped_poll=0 read_beside=1). No request waited on the device, and noofflineappeared. - An
offlineline now names the request it gave up on, and why.
Whole-app sweep at the head, after the frozen run, on the five-note
vault: no notices; idle before and after Sync now; Show sync status
with five files tracked, nothing remote-only and feed sequence 25 on both;
the obsync settings tab rendered; no warning or error in either console.
Fifth round: #297, a run of 5xx answers¶
Builds: the train at 38fedadd ("before") and this branch ("after"),
with the same server binary. One rig, A, shown, whose Server URL is a
loopback hop in front of obsyncd (lane K's fault-proxy.mjs). While a
flag file exists, the hop answers every chunk upload 500, and it
forwards everything else unchanged. A note was written with the fault on,
and the hop kept faulting for 240 s: the note's push tried its chunk
eight times into the 500s and gave up. Each request's start and end were
kept in the page, beside the hop's log of each request's end and
duration (lab-H exp297.sh, analyze297.py).
- Before, at load 8 to 12. Four answers came after a 500 while the
long poll had waited, and each dropped it (
woken reason=answered dropped_poll=1, after 9.3, 10.1, 32.6 and 58.1 s).requestUrlcannot withdraw a request: the device held up to four long polls at once, and the hop up to three. Two of the polls sent in their place reached the hop 20.00 s after they left the plugin. - After, at load 6 to 10. Five answers came after a 500, and each
kept the poll and read beside it in 9 to 18 ms (
dropped_poll=0 read_beside=1 polls_in_flight=1). The device and the hop held one long poll throughout, and every poll reached the hop within 2 ms. No feed request was cancelled, timed out or retried. - In the first run of the change (load 64 to 95), two of those
answers were the poll's own. The transport had let the poll's count go
before the wake its answer raised, so that wake dropped a poll whose
answer had come (
dropped_poll=1 polls_in_flight=0 waited_ms=55018). A poll now counts until its attempt has reported its answer. The run above is at that fix: two of its five kept polls had waited 55.0 s. - What the status said. In both runs, each of the eight
offlinelines named the chunk upload (request=PUT /v1/chunks/<id> status=500), and none a read of the feed. The status readoffline — retryingfrom each 500 to the next answer: in 21 of 90 samples before, longest 20 s, and in 55 of 111 after, longest 48 s. The difference is in how often something is answered: before, the dropped polls, held side by side, answered more often; after, the one poll answers once in 55 s. A server that answers 5xx reading as absence is #291, #292 and #295. - Once the fault ended and a focus started a walk, the note landed and
the status read
idle: 7 s later before (at load 61), and within the first 2 s sample after.
Whole-app sweep at the head, after the run: no notices; idle before
and after Sync now; Show sync status with one file tracked, nothing
remote-only and feed sequence 10; the obsync settings tab rendered; no
warning or error in the console.
Sixth round: #298, the server's own 5xx¶
Builds: the train at 52e4831e ("before") and this branch ("after"), with
the same server binary. The rig and hop are those of the fifth round. The
hop answered every chunk upload 500 {"error":"io_error"}, obsync's own
error, for 240 s; in a third run it answered a bare 502 with no body, as
a proxy does when obsync behind it is gone, for 120 s.
- Before, at load 6 to 13. Each 500 was taken for the server's
absence. The status read
offline — retryingin 27 of 110 samples, longest 38 s. Each of the 8offlinelines named the upload (request=PUT /v1/chunks/<id> status=500). - After, at load 4 to 8. No
offlineline, and no sample readoffline: the status readsyncing 1 filefrom the upload's start to the end of the fault. Each retry line named the server's code (status=500 code=io_error); the push gave up after 8 attempts and 96 s, and the note was sent again and landed at the first pass after the fault. - After, a bare 502, at load 6 to 9. Still absence: the status read
offline — retryingin 35 of 55 samples, longest 46 s, and each of the 8offlinelines namedPUT /v1/chunks/<id> status=502. - In all three runs one long poll was held at a time, and no feed request was cancelled, timed out or retried.
Whole-app sweep at the head, after the 500 run: no notices; idle
before and after Sync now; Show sync status with one file tracked, nothing
remote-only and feed sequence 10; the obsync settings tab rendered; no
warning or error in the console.
Not covered here¶
- Mobile, Windows and Linux desktops.
- Several devices uploading at once against a server on real hardware, the case the group commit is for: measured here only with the signed load generator.
- A clean-machine absolute baseline: the laptop was never idle.