PERF: Profile small-file WebDAV durable-write latency #429

Open
opened 2026-09-29 11:21:42 +00:00 by kayg · 6 comments
Owner

Evidence

On the isolated perf-test VM, one completed 10,000-file upload of 30,000-byte files took 1,477.205 s at 6.77 files/s (--transfers 16). The matching rclone serve webdav sample took 348.843 s at 28.666 files/s. A second baseline repeat was 378.109 s at 26.447 files/s. Build-host load was high and changed during these runs; these are not medians.

A diagnostic 30,000-byte PUT took 330.97 ms from the client. Its profile spans were: atomic_write_and_fsync 131.618 ms, staging_write_and_fsync 25.223 ms, upload_create 58.079 ms, quota_reservation 28.370 ms, and app_password_verifier 17.554 ms. Server CPU during the completed Calternal upload averaged 85.29% of one core, peaking at 170.94%; RSS ranged from 289,608 to 707,804 KiB. Raw commands, loads and limits are in issue #409 and docs/perf/webdav-2026-09-29.md.

Follow-up

Profile the repeated per-file WebDAV write path on the low-spec host. Check the cost of staging and atomic write durability, and reduce overhead only if file and database durability remain unchanged. Re-run the full A/B protocol before using this sample as a performance decision.

## Evidence On the isolated perf-test VM, one completed 10,000-file upload of 30,000-byte files took 1,477.205 s at 6.77 files/s (`--transfers 16`). The matching `rclone serve webdav` sample took 348.843 s at 28.666 files/s. A second baseline repeat was 378.109 s at 26.447 files/s. Build-host load was high and changed during these runs; these are not medians. A diagnostic 30,000-byte PUT took 330.97 ms from the client. Its profile spans were: `atomic_write_and_fsync` 131.618 ms, `staging_write_and_fsync` 25.223 ms, `upload_create` 58.079 ms, `quota_reservation` 28.370 ms, and `app_password_verifier` 17.554 ms. Server CPU during the completed Calternal upload averaged 85.29% of one core, peaking at 170.94%; RSS ranged from 289,608 to 707,804 KiB. Raw commands, loads and limits are in issue #409 and `docs/perf/webdav-2026-09-29.md`. ## Follow-up Profile the repeated per-file WebDAV write path on the low-spec host. Check the cost of staging and atomic write durability, and reduce overhead only if file and database durability remain unchanged. Re-run the full A/B protocol before using this sample as a performance decision.
Author
Owner

Baseline durability evidence for the #409 measurement: the installed comparison is rclone v1.75.1 serve webdav on its local filesystem, with no VFS cache override. In the tagged source, the WebDAV PUT handler opens the final path with O_CREATE|O_TRUNC, streams the body, then closes it; the local backend also writes directly to that pathname. The handler does not call Sync or rename a staging file, and Rclone's VFS Sync method is a no-op. There is no per-file fsync in this path, so its throughput does not carry calternal's DESIGN §2 crash-durability guarantee. Sources: https://github.com/golang/net/blob/v0.57.0/webdav/webdav.go#L258-L296, https://github.com/rclone/rclone/blob/v1.75.1/vfs/vfscommon/options.go#L51-L55, https://github.com/rclone/rclone/blob/v1.75.1/backend/local/local.go#L1538-L1655, https://github.com/rclone/rclone/blob/v1.75.1/vfs/write.go#L362-L366.

The current fsync(temp) → rename → fsync(parent) contract and the guarantees preserved by group commit, batched parent-directory fsync, and syncfs are recorded in docs/perf/webdav-2026-09-29.md. Each proposed barrier must complete before the server acknowledges a write; new directory entries also need durable parent entries.

Baseline durability evidence for the #409 measurement: the installed comparison is rclone v1.75.1 `serve webdav` on its local filesystem, with no VFS cache override. In the tagged source, the WebDAV PUT handler opens the final path with `O_CREATE|O_TRUNC`, streams the body, then closes it; the local backend also writes directly to that pathname. The handler does not call `Sync` or rename a staging file, and Rclone's VFS `Sync` method is a no-op. There is no per-file `fsync` in this path, so its throughput does not carry calternal's DESIGN §2 crash-durability guarantee. Sources: https://github.com/golang/net/blob/v0.57.0/webdav/webdav.go#L258-L296, https://github.com/rclone/rclone/blob/v1.75.1/vfs/vfscommon/options.go#L51-L55, https://github.com/rclone/rclone/blob/v1.75.1/backend/local/local.go#L1538-L1655, https://github.com/rclone/rclone/blob/v1.75.1/vfs/write.go#L362-L366. The current `fsync(temp) → rename → fsync(parent)` contract and the guarantees preserved by group commit, batched parent-directory fsync, and `syncfs` are recorded in `docs/perf/webdav-2026-09-29.md`. Each proposed barrier must complete before the server acknowledges a write; new directory entries also need durable parent entries.
Author
Owner

New #409 evidence for #429: Mac Rclone v1.75.1, 200 files / 1,673,945 bytes, 16 transfers, three interleaved repeats. Calternal median upload was 18.236 s (10.967 files/s) and download 5.324 s (37.566 files/s). Rclone WebDAV baseline medians were 8.447 s (23.677 files/s) and 3.121 s (64.082 files/s). The baseline uses direct final-path writes and no per-file fsync/rename barrier, so this is not an equal-durability comparison. Raw repeats are being committed under docs/perf/ in #409.

New #409 evidence for #429: Mac Rclone v1.75.1, 200 files / 1,673,945 bytes, 16 transfers, three interleaved repeats. Calternal median upload was 18.236 s (10.967 files/s) and download 5.324 s (37.566 files/s). Rclone WebDAV baseline medians were 8.447 s (23.677 files/s) and 3.121 s (64.082 files/s). The baseline uses direct final-path writes and no per-file fsync/rename barrier, so this is not an equal-durability comparison. Raw repeats are being committed under `docs/perf/` in #409.
Author
Owner

Additional #429 evidence from three interleaved Mac Rclone repeats on 2,000 × 1–16 KiB files: Calternal median upload was 199.159 s (10.042 files/s) and download 42.010 s (47.608 files/s). Rclone WebDAV baseline medians were 75.892 s (26.353 files/s) and 24.431 s (81.865 files/s). Calternal upload repeats were 194.194, 199.159 and 215.106 s. The baseline does not fsync each file or stage and rename writes, so it does not keep the same durability guarantee. Raw per-run numbers and load gates are in #409 docs/perf/raw/webdav-2026-09-29-mac-rclone-small-2000.jsonl.

Additional #429 evidence from three interleaved Mac Rclone repeats on 2,000 × 1–16 KiB files: Calternal median upload was 199.159 s (10.042 files/s) and download 42.010 s (47.608 files/s). Rclone WebDAV baseline medians were 75.892 s (26.353 files/s) and 24.431 s (81.865 files/s). Calternal upload repeats were 194.194, 199.159 and 215.106 s. The baseline does not fsync each file or stage and rename writes, so it does not keep the same durability guarantee. Raw per-run numbers and load gates are in #409 `docs/perf/raw/webdav-2026-09-29-mac-rclone-small-2000.jsonl`.
Author
Owner

New first-repeat Mac large-file measurements are in docs/perf/webdav-2026-09-29.md and docs/perf/raw/webdav-2026-09-29-mac-rclone-large-files.jsonl. For 5 GB, Calternal took 458.750 s up and 114.359 s down; the Rclone baseline took 290.696 s up and 101.015 s down. Service CPU/RSS was captured: Calternal upload averaged 20.722% of one core at 756,444–756,504 KiB RSS; baseline upload averaged 31.296% at 97,544–98,972 KiB. Calternal download averaged 14.361% at 762,956–828,984 KiB; baseline download averaged 5.206% at 97,628–157,808 KiB.

This remains a different-durability comparison. The baseline Rclone WebDAV handler writes directly to the final path and closes it without a per-file fsync or staged rename. Calternal's durable write path still requires file fsync, same-directory rename and parent-directory fsync before success. The prior stage profile measured atomic_write_and_fsync at 131.618 ms in a 30 KB PUT. Group commit or batched directory fsync can preserve the acknowledged-write guarantee if file fsyncs and all affected directory barriers complete before success; syncfs can preserve it with a broader filesystem-wide barrier. These are one-sample timings, not medians.

New first-repeat Mac large-file measurements are in `docs/perf/webdav-2026-09-29.md` and `docs/perf/raw/webdav-2026-09-29-mac-rclone-large-files.jsonl`. For 5 GB, Calternal took 458.750 s up and 114.359 s down; the Rclone baseline took 290.696 s up and 101.015 s down. Service CPU/RSS was captured: Calternal upload averaged 20.722% of one core at 756,444–756,504 KiB RSS; baseline upload averaged 31.296% at 97,544–98,972 KiB. Calternal download averaged 14.361% at 762,956–828,984 KiB; baseline download averaged 5.206% at 97,628–157,808 KiB. This remains a different-durability comparison. The baseline Rclone WebDAV handler writes directly to the final path and closes it without a per-file fsync or staged rename. Calternal's durable write path still requires file fsync, same-directory rename and parent-directory fsync before success. The prior stage profile measured `atomic_write_and_fsync` at 131.618 ms in a 30 KB PUT. Group commit or batched directory fsync can preserve the acknowledged-write guarantee if file fsyncs and all affected directory barriers complete before success; `syncfs` can preserve it with a broader filesystem-wide barrier. These are one-sample timings, not medians.
Author
Owner

Round 2 update (2026-09-29): this run collected no new quiet-VM WebDAV measurements. The Files profile gate was still compiling when the four-hour job limit was reached; the pending helper edit was discarded, and no profile ran under /root/perf.lock. Keep the existing sample above unchanged; this update adds no throughput or durable-write numbers. Parent issue #367 records the exact gates and remaining work.

Round 2 update (2026-09-29): this run collected no new quiet-VM WebDAV measurements. The Files profile gate was still compiling when the four-hour job limit was reached; the pending helper edit was discarded, and no profile ran under `/root/perf.lock`. Keep the existing sample above unchanged; this update adds no throughput or durable-write numbers. Parent issue #367 records the exact gates and remaining work.
Author
Owner

#367 quiet-VM numbers for the durable write path (perf-test VM, 4 vCPU; release 369ab6a2f; one client, sequential Tus uploads; load < 1 at start under /root/perf.lock). These are the same per-file write primitives as the WebDAV path in this issue.

Workload Files Wall Rate Upload p50 / p95 Server CPU per file
~120 B .txt into one folder (files-50k) 50,000 5,529.6 s 8.26 /s (10.78 /s for the first 1k) 111.1 / 195.1 ms (80.7 / 165.7 ms at 1k) 37.7 ms (28.9 ms over the first 1k)
64 KiB PDF (pdf-batch) 10,000 4,128.2 s 2.42 /s (155 KiB/s) 293.9 / 358.1 ms 32.2 ms
10k Items for the change feed (sync-feed-10k) 10,000 941.4 s 10.6 /s — 15.6 ms
Client daemon, 10k-file Home, 200 remote replaces 200 — — server replace upload 105 / 216 ms —
  • The PDF import uses 32 ms CPU per 294 ms upload, so about 90% of each request is waiting, not computing. That matches this issue's fsync-heavy spans.
  • In the .txt import, 40% of perf samples were ONNX (semantic embedding of each new text file), so part of the growing CPU per file is embedding backlog, not the write path.
  • The PDF import also got 33 503 Index is busy; retry shortly answers from one sequential client (the probe now retries them; before that the run stopped at document 2,009).

Raw data: docs/perf/runs/2026-09-29T224501Z-369ab6a2/raw/{files-50k,pdf-batch,sync-feed-10k}.json, client-sync.txt; profile profiles/files-early on the VM.

#367 quiet-VM numbers for the durable write path (perf-test VM, 4 vCPU; release `369ab6a2f`; one client, sequential Tus uploads; load < 1 at start under `/root/perf.lock`). These are the same per-file write primitives as the WebDAV path in this issue. | Workload | Files | Wall | Rate | Upload p50 / p95 | Server CPU per file | |---|---:|---:|---:|---:|---:| | ~120 B `.txt` into one folder (`files-50k`) | 50,000 | 5,529.6 s | 8.26 /s (10.78 /s for the first 1k) | 111.1 / 195.1 ms (80.7 / 165.7 ms at 1k) | 37.7 ms (28.9 ms over the first 1k) | | 64 KiB PDF (`pdf-batch`) | 10,000 | 4,128.2 s | 2.42 /s (155 KiB/s) | 293.9 / 358.1 ms | 32.2 ms | | 10k Items for the change feed (`sync-feed-10k`) | 10,000 | 941.4 s | 10.6 /s | — | 15.6 ms | | Client daemon, 10k-file Home, 200 remote replaces | 200 | — | — | server replace upload 105 / 216 ms | — | - The PDF import uses 32 ms CPU per 294 ms upload, so about 90% of each request is waiting, not computing. That matches this issue's fsync-heavy spans. - In the `.txt` import, 40% of `perf` samples were ONNX (semantic embedding of each new text file), so part of the growing CPU per file is embedding backlog, not the write path. - The PDF import also got 33 `503 Index is busy; retry shortly` answers from one sequential client (the probe now retries them; before that the run stopped at document 2,009). Raw data: `docs/perf/runs/2026-09-29T224501Z-369ab6a2/raw/{files-50k,pdf-batch,sync-feed-10k}.json`, `client-sync.txt`; profile `profiles/files-early` on the VM.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#429
No description provided.