PERF: CalDAV at scale: 10k-change sync p95 35 s, area moves ~5 min, 50k listing times out #573

Open
opened 2026-10-01 05:00:18 +00:00 by kayg · 28 comments
Owner

Found by caldav-stress round 3 (#457, 2026-10-01; profiles in job/caldav-stress docs/perf/caldav-baseline.md)

Correctness held (50-client convergence, If-Match race with 1 winner + 7×412, area moves consistent), but the scale numbers are far off:

  • sync-collection after 10,000 changes: p50 193 ms, p95 34.6 s, p99 76.9 s, peak RSS 527 MB;
  • live area moves (5 writers, 50 clients, 1,100 resources): p50 289 s, p95 303 s;
  • 50,000-entry fixture listing times out at the 120 s request bound (no valid rows).
    Job (root cause first; Luna max, Sol if the trace is unclear):
  1. Profile the 10k sync-collection and the area-move path on the perf VM (flock /root/perf.lock; HDD emulation on if #549 left it set up). Use Server-Timing (from #549, if merged) or tracing spans. Find the per-change cost: per-item file reads? a recompute of every ETag? the CTag/sync-token computed by a full scan? N+1 SQLite queries? a global lock held across a move?
  2. Fix the causes, one commit each, with before/after numbers.
  3. Budgets: sync-collection after 10k changes p95 ≤ 2 s; an area move visible to all clients ≤ 5 s; a 50k-entry listing (PROPFIND depth 1 / calendar-query) streams within 10 s with bounded RSS.
    Correctness must stay: rerun the convergence + If-Match race + Apple replay fixtures. Per-crate gates (calternal-dav, plugins/calendar, plugins/notes as touched). Time limit 5 h.
## Found by caldav-stress round 3 (#457, 2026-10-01; profiles in job/caldav-stress `docs/perf/caldav-baseline.md`) Correctness held (50-client convergence, If-Match race with 1 winner + 7×412, area moves consistent), but the scale numbers are far off: - sync-collection after **10,000 changes**: p50 193 ms, **p95 34.6 s, p99 76.9 s**, peak RSS 527 MB; - live **area moves** (5 writers, 50 clients, 1,100 resources): **p50 289 s**, p95 303 s; - **50,000-entry** fixture listing times out at the 120 s request bound (no valid rows). **Job (root cause first; Luna max, Sol if the trace is unclear):** 1. Profile the 10k sync-collection and the area-move path on the perf VM (`flock /root/perf.lock`; HDD emulation on if #549 left it set up). Use `Server-Timing` (from #549, if merged) or tracing spans. Find the per-change cost: per-item file reads? a recompute of every ETag? the CTag/sync-token computed by a full scan? N+1 SQLite queries? a global lock held across a move? 2. Fix the causes, one commit each, with before/after numbers. 3. **Budgets:** sync-collection after 10k changes p95 ≤ 2 s; an area move visible to all clients ≤ 5 s; a 50k-entry listing (PROPFIND depth 1 / calendar-query) streams within 10 s with bounded RSS. Correctness must stay: rerun the convergence + If-Match race + Apple replay fixtures. Per-crate gates (calternal-dav, plugins/calendar, plugins/notes as touched). Time limit 5 h.
Author
Owner

Started issue #573 on branch job/davscale-573, based at cc25c441b7a974185622a1dee853cf38686d2b67 (Use valid Appearance writes in contention probe). The branch is currently 326 commits behind origin/dev; I will merge origin/dev once before final gates. I am tracing sync-collection, area moves, and collection listing before changing code.

Started issue #573 on branch `job/davscale-573`, based at `cc25c441b7a974185622a1dee853cf38686d2b67` (`Use valid Appearance writes in contention probe`). The branch is currently 326 commits behind `origin/dev`; I will merge `origin/dev` once before final gates. I am tracing sync-collection, area moves, and collection listing before changing code.
Author
Owner

Current trace after the one origin/dev fast-forward (3f258302a0f2d6418ff60c9ce22cbb33e008ca99): crates/calternal-dav/src/protocol.rs::area_report reads JournalProvider::changes and then calls JournalProvider::list for both initial and incremental reports (around lines 1167–1222). NotesJournalProvider::list calls scan, which walks Home and reads every visible Markdown file (crates/plugins/notes/src/lib.rs:345–378; crates/plugins/notes/src/store.rs:1802–1845). Incremental sync therefore reparses the whole Home for each delta. Calendar collection PROPFIND calls changes(user, None) only to get a token (protocol.rs:1049–1051), while changes(None) selects and materializes every resource ID (plugins/notes/src/lib.rs:976–982). Resource lookup for an area copy also fetches every resource ID before hashing them (plugins/notes/src/lib.rs:860–880). The existing note_calendar_logs projection has event fields, tags, date and attachments; I will use it with journal_resources to query only the requested collection and changed IDs, and add an O(1) token read.

Current trace after the one `origin/dev` fast-forward (`3f258302a0f2d6418ff60c9ce22cbb33e008ca99`): `crates/calternal-dav/src/protocol.rs::area_report` reads `JournalProvider::changes` and then calls `JournalProvider::list` for both initial and incremental reports (around lines 1167–1222). `NotesJournalProvider::list` calls `scan`, which walks Home and reads every visible Markdown file (`crates/plugins/notes/src/lib.rs:345–378`; `crates/plugins/notes/src/store.rs:1802–1845`). Incremental sync therefore reparses the whole Home for each delta. Calendar collection PROPFIND calls `changes(user, None)` only to get a token (`protocol.rs:1049–1051`), while `changes(None)` selects and materializes every resource ID (`plugins/notes/src/lib.rs:976–982`). Resource lookup for an area copy also fetches every resource ID before hashing them (`plugins/notes/src/lib.rs:860–880`). The existing `note_calendar_logs` projection has event fields, tags, date and attachments; I will use it with `journal_resources` to query only the requested collection and changed IDs, and add an O(1) token read.
Author
Owner

Replay finding: after the area query began accepting source block IDs, the existing apple_calendar_multiget_accepts_100_hrefs_over_legacy_report_cap fixture failed because calendar-multiget hrefs carry area copy IDs. Added a copy-aware alias-index lookup; the unchanged fixture now passes. Legacy journal_resources rows have no cached Log fields, so the first DAV read reindexes those Daily notes through the normal Markdown index transaction. The regression checks cached ETag, iCalendar data, linked Note URL, attachments, alarms, untagged membership, copy lookup, and that backfill leaves the sync token unchanged.

Replay finding: after the area query began accepting source block IDs, the existing `apple_calendar_multiget_accepts_100_hrefs_over_legacy_report_cap` fixture failed because calendar-multiget hrefs carry area copy IDs. Added a copy-aware alias-index lookup; the unchanged fixture now passes. Legacy `journal_resources` rows have no cached Log fields, so the first DAV read reindexes those Daily notes through the normal Markdown index transaction. The regression checks cached ETag, iCalendar data, linked Note URL, attachments, alarms, untagged membership, copy lookup, and that backfill leaves the sync token unchanged.
Author
Owner

Profile review finding: ensure_ids now checks for legacy Journal rows whose cached log_line is empty. That check runs on DAV reads; without a matching index it would scan all journal_resources on every request, including after the one-time backfill. I am adding a partial (user_id,path) index for only log_line='' rows and a query-plan regression check. After migration backfill, the hot request checks an empty partial index instead of scanning the 50k-resource table.

Profile review finding: `ensure_ids` now checks for legacy Journal rows whose cached `log_line` is empty. That check runs on DAV reads; without a matching index it would scan all `journal_resources` on every request, including after the one-time backfill. I am adding a partial `(user_id,path)` index for only `log_line=''` rows and a query-plan regression check. After migration backfill, the hot request checks an empty partial index instead of scanning the 50k-resource table.
Author
Owner

Progress on job/davscale-573 (HEAD d714380fc2): built the real release server after generating the required ignored apps/web/build bundle. A local smoke profile (20 area resources, 10 Log writes, 2 sync clients) completed: sync p95 47.805 ms; five-move visibility p95 842.483 ms; per-client sync caches matched final full listings; Calendar Log projection passed; If-Match race was 1×201 + 7×412; all five hostile REPORT inputs returned their expected 400/413 statuses. Fixed two profile assumptions exposed by the run: PROPFIND must request displayname to discover area collections, and Calendar range exposes Journal rows in logs. The full local scale profile is running next; host load was high and perf VM lock was busy, so results will be labeled local.

Progress on job/davscale-573 (HEAD d714380fc2da5f7006d8429ab92e887d9b8b6b74): built the real release server after generating the required ignored apps/web/build bundle. A local smoke profile (20 area resources, 10 Log writes, 2 sync clients) completed: sync p95 47.805 ms; five-move visibility p95 842.483 ms; per-client sync caches matched final full listings; Calendar Log projection passed; If-Match race was 1×201 + 7×412; all five hostile REPORT inputs returned their expected 400/413 statuses. Fixed two profile assumptions exposed by the run: PROPFIND must request displayname to discover area collections, and Calendar range exposes Journal rows in `logs`. The full local scale profile is running next; host load was high and perf VM lock was busy, so results will be labeled local.
Author
Owner

Finding from the first full-profile setup attempt: with 50k Log rows written as 3,653 raw Daily Note files before startup, the local server did not open /readyz within the 900 s bound. At timeout the local SQLite index was about 163 MB and /proc reported about 5.3 GB of cumulative writes for that server process; no DAV request samples were produced. The profile now creates the same real Journal resources through the bounded Notes batch API after a clean server is ready, then checks the indexed resource count before timing DAV reads. This records fixture seeding separately from request latency and avoids misreporting cold Home import as listing latency. Reduced smoke run passes with the live seeding path.

Finding from the first full-profile setup attempt: with 50k Log rows written as 3,653 raw Daily Note files before startup, the local server did not open `/readyz` within the 900 s bound. At timeout the local SQLite index was about 163 MB and `/proc` reported about 5.3 GB of cumulative writes for that server process; no DAV request samples were produced. The profile now creates the same real Journal resources through the bounded Notes batch API after a clean server is ready, then checks the indexed resource count before timing DAV reads. This records fixture seeding separately from request latency and avoids misreporting cold Home import as listing latency. Reduced smoke run passes with the live seeding path.
Author
Owner

Profile adjustment (HEAD 6368ee4d7a): API-seeding all 50k resources into one Daily Note caused batch setup to reindex the growing file on every batch. I stopped that setup at 14k (about 394 MB process RSS, ~1 CPU core, no timed DAV requests yet) and changed the fixture to put each bounded batch on its own real Daily Note. This keeps each index transaction bounded while preserving 50k area resources. The reduced profile passes again with this layout; starting the full run now.

Profile adjustment (HEAD 6368ee4d7a219ce8d02556411cd3ab47e25da427): API-seeding all 50k resources into one Daily Note caused batch setup to reindex the growing file on every batch. I stopped that setup at 14k (about 394 MB process RSS, ~1 CPU core, no timed DAV requests yet) and changed the fixture to put each bounded batch on its own real Daily Note. This keeps each index transaction bounded while preserving 50k area resources. The reduced profile passes again with this layout; starting the full run now.
Author
Owner

Another profile setup finding (HEAD c4b9334325): after 50k live Log writes, the collection Home omitted Bench01, so that run produced no timed request samples. A reduced run passed after I seeded Bench00 and Bench01 as real Note Tags and explicitly reconciled the small Tags catalog before generating the large Journal fixture. The latest 20-resource profile completes with this setup and retains the cross-area move checks. I am rerunning the full profile once with the corrected catalog setup.

Another profile setup finding (HEAD c4b9334325b234e80b8e98ecf31b0699037839e5): after 50k live Log writes, the collection Home omitted `Bench01`, so that run produced no timed request samples. A reduced run passed after I seeded `Bench00` and `Bench01` as real Note Tags and explicitly reconciled the small Tags catalog before generating the large Journal fixture. The latest 20-resource profile completes with this setup and retains the cross-area move checks. I am rerunning the full profile once with the corrected catalog setup.
Author
Owner

Final profile-fixture correction (HEAD b38f3bacd1): the prior 50k PROPFIND verified all 50,000 resources, but calendar-query returned 2,000 because the fixture started on 2026-01-01 and the query ended on 2026-01-03. The resource index was complete; the synthetic dates exceeded the test window. I moved the 50k fixture back to 2016-01-01 so all 50 daily batches fall within the query range. The reduced run now completes with the expected response counts and all correctness checks. Starting the full measurement now.

Final profile-fixture correction (HEAD b38f3bacd17c3c91ddc0e360d6eeb8883dd5e7da): the prior 50k PROPFIND verified all 50,000 resources, but calendar-query returned 2,000 because the fixture started on 2026-01-01 and the query ended on 2026-01-03. The resource index was complete; the synthetic dates exceeded the test window. I moved the 50k fixture back to 2016-01-01 so all 50 daily batches fall within the query range. The reduced run now completes with the expected response counts and all correctness checks. Starting the full measurement now.
Author
Owner

Sync profile finding (HEAD 1e4a749bef): the Notes Log batch API durably writes Markdown and queues Journal projection work, so an immediate DAV REPORT can race that background job. The 10k profile saw 9k rows immediately after ten accepted batches; the 1k diagnostic showed a 1.528 s projection settle and then returned all 1,000 rows, with a 1.039 s sync request. The harness now waits (bounded by the request timeout) for all expected rows before timing the stable sync request and records projection_settle_ms separately. No DAV API code changed for this timing artifact.

Sync profile finding (HEAD 1e4a749bef8e7fd5e5635bc5f97eb3fa75fb53bf): the Notes Log batch API durably writes Markdown and queues Journal projection work, so an immediate DAV REPORT can race that background job. The 10k profile saw 9k rows immediately after ten accepted batches; the 1k diagnostic showed a 1.528 s projection settle and then returned all 1,000 rows, with a 1.039 s sync request. The harness now waits (bounded by the request timeout) for all expected rows before timing the stable sync request and records `projection_settle_ms` separately. No DAV API code changed for this timing artifact.
Author
Owner

Local diagnostic for #573: with 1,000 existing entries, 100 changes, 50 sync clients and five MOVE writers, convergence p95 was 12,800.486 ms and MOVE writer p95 was 8,917.299 ms. Server peak RSS was 284,295,168 bytes. Host load rose from 14.625 to 23.263. This smaller fixture reproduced the move latency, so the 50k listing size was not its only cause.

Code inspection found every healthy REPORT took the Notes per-User writer lock to run ensure_ids, even when both indexed repair queues were empty. Each poll calls changes and list_area; that serialized the poll storm against MOVE writes. I committed a fast path in 6f23ec9ce that checks indexed pending queues before locking and rechecks after acquiring the lock when repair is needed. Notes clippy passed; Notes tests passed (163 unit, one Apple replay integration). These are local diagnostics; the perf VM was unreachable (No route to host). The after measurement is in progress.

Local diagnostic for #573: with 1,000 existing entries, 100 changes, 50 sync clients and five MOVE writers, convergence p95 was 12,800.486 ms and MOVE writer p95 was 8,917.299 ms. Server peak RSS was 284,295,168 bytes. Host load rose from 14.625 to 23.263. This smaller fixture reproduced the move latency, so the 50k listing size was not its only cause. Code inspection found every healthy REPORT took the Notes per-User writer lock to run `ensure_ids`, even when both indexed repair queues were empty. Each poll calls `changes` and `list_area`; that serialized the poll storm against MOVE writes. I committed a fast path in `6f23ec9ce` that checks indexed pending queues before locking and rechecks after acquiring the lock when repair is needed. Notes clippy passed; Notes tests passed (163 unit, one Apple replay integration). These are local diagnostics; the perf VM was unreachable (`No route to host`). The after measurement is in progress.
Author
Owner

Tracing evidence from a local 1,000-entry / 100-change / 50-client run (host load 22.55 to 23.27): 248 App Password verifier calls had p50 1,808 ms and p95 5,000 ms; 6 requests retried HTTP 429. The DAV stages had indexed area-list p95 14.612 ms, MOVE Journal put p95 3,218.752 ms, and per-User area write-lock wait max 5,202.262 ms.

The lock span confirms that five independent MOVE requests wait on the global per-User area mutex, which covers source lookup, target collision check and the conditional Notes write. The Notes write already uses the source ETag as a compare-and-swap. I am removing this outer mutex from MOVE only and adding a same-source concurrent-move regression that requires one winner and one 412. PUT and DELETE keep their duplicate/name serialization. The profile now retries and counts bounded 429 auth backpressure, and stores only aggregate timings.

Tracing evidence from a local 1,000-entry / 100-change / 50-client run (host load 22.55 to 23.27): 248 App Password verifier calls had p50 1,808 ms and p95 5,000 ms; 6 requests retried HTTP 429. The DAV stages had indexed area-list p95 14.612 ms, MOVE Journal put p95 3,218.752 ms, and per-User area write-lock wait max 5,202.262 ms. The lock span confirms that five independent MOVE requests wait on the global per-User area mutex, which covers source lookup, target collision check and the conditional Notes write. The Notes write already uses the source ETag as a compare-and-swap. I am removing this outer mutex from MOVE only and adding a same-source concurrent-move regression that requires one winner and one 412. PUT and DELETE keep their duplicate/name serialization. The profile now retries and counts bounded 429 auth backpressure, and stores only aggregate timings.
Author
Owner

Local release profile after 1f764f5ff (perf VM still returns No route to host): 50k PROPFIND p95 5.97s, 50k calendar-query p95 4.81s, 10k-change sync p95 1.96s. The listing and sync budgets pass. Five area moves observed by 50 clients take 18.35s p95; all 5 moves and all 50 sync caches converge, the If-Match race still has 1 winner and 7×412, and hostile REPORTs return 400/413 as expected. The DAV name-lock wait is now 0.015ms p95, but move_journal_put is 8.74s p95 and App Password verification is 4.61s p95. Load at start/end was 15.13/17.57/17.22 and 13.17/19.40/19.30. I am tracing the Notes write stages next; the profile is saved at docs/perf/runs/2026-10-01-davscale-573-local-after-move-lock.json.

Local release profile after `1f764f5ff` (perf VM still returns `No route to host`): 50k PROPFIND p95 5.97s, 50k calendar-query p95 4.81s, 10k-change sync p95 1.96s. The listing and sync budgets pass. Five area moves observed by 50 clients take 18.35s p95; all 5 moves and all 50 sync caches converge, the If-Match race still has 1 winner and 7×412, and hostile REPORTs return 400/413 as expected. The DAV name-lock wait is now 0.015ms p95, but `move_journal_put` is 8.74s p95 and App Password verification is 4.61s p95. Load at start/end was 15.13/17.57/17.22 and 13.17/19.40/19.30. I am tracing the Notes write stages next; the profile is saved at `docs/perf/runs/2026-10-01-davscale-573-local-after-move-lock.json`.
Author
Owner

The 1k-entry / 100-change diagnostic after 151c4fd93 confirms that area reads are indexed and do not scan the Home. Under local load 25.76→30.15, the 50-client visibility p95 was 11.876s and writer MOVE p95 was 4.946s. Notes-stage p95 was 1.821s waiting for the per-User write lock, 17ms locating the source, 1.258s for the checked Markdown update, 1.501s for projection indexing, and 381ms for Daily note navigation. App Password verification was 4.622s p95. All 5 moves, 50 sync caches, full listings and Calendar projections converged; the 1-winner/7×412 race passed. These local numbers include severe shared-host load; the perf VM remains unreachable. The trace artifact is docs/perf/runs/2026-10-01-davscale-573-write-stages-probe.json.

The 1k-entry / 100-change diagnostic after `151c4fd93` confirms that area reads are indexed and do not scan the Home. Under local load 25.76→30.15, the 50-client visibility p95 was 11.876s and writer MOVE p95 was 4.946s. Notes-stage p95 was 1.821s waiting for the per-User write lock, 17ms locating the source, 1.258s for the checked Markdown update, 1.501s for projection indexing, and 381ms for Daily note navigation. App Password verification was 4.622s p95. All 5 moves, 50 sync caches, full listings and Calendar projections converged; the 1-winner/7×412 race passed. These local numbers include severe shared-host load; the perf VM remains unreachable. The trace artifact is `docs/perf/runs/2026-10-01-davscale-573-write-stages-probe.json`.
Author
Owner

A corrected 1k-entry/100-change diagnostic confirms the previous result and records the actual scenario sizes. At local start/end load 45.36/42.40/34.95 → 42.40/42.60/35.48, MOVE visibility p95 was 19.099s. App Password verification p95/p99 was 5.283/12.974s. Notes per-User lock wait p95 was 3.529s; source locate was 15ms; checked Markdown update was 140ms; projection index was 1.839s; navigation was 425ms. The DAV lock wait was 0.002ms. All 5 moves, 50 sync caches, full listings, Calendar projection, 1-winner/7×412 If-Match race, and five hostile REPORT cases passed. This is a loaded local diagnostic, not an acceptance measurement. The perf VM remains unreachable. The corrected artifact is docs/perf/runs/2026-10-01-davscale-573-write-stages-final.json.

A corrected 1k-entry/100-change diagnostic confirms the previous result and records the actual scenario sizes. At local start/end load 45.36/42.40/34.95 → 42.40/42.60/35.48, MOVE visibility p95 was 19.099s. App Password verification p95/p99 was 5.283/12.974s. Notes per-User lock wait p95 was 3.529s; source locate was 15ms; checked Markdown update was 140ms; projection index was 1.839s; navigation was 425ms. The DAV lock wait was 0.002ms. All 5 moves, 50 sync caches, full listings, Calendar projection, 1-winner/7×412 If-Match race, and five hostile REPORT cases passed. This is a loaded local diagnostic, not an acceptance measurement. The perf VM remains unreachable. The corrected artifact is `docs/perf/runs/2026-10-01-davscale-573-write-stages-final.json`.
Author
Owner

Finished

Implemented Forgejo #573 on branch job/davscale-573. Head: f943e0dab500f69835462dd3e46b91fa117b9e9f.

The Notes Journal now reads sync tokens from the per-User cursor and serves area listings and copy lookups from indexed projections. Healthy DAV reads skip the Notes writer lock check when no repair is pending. MOVE no longer holds the coarse DAV area-name lock. Notes still checks the source ETag while it writes under its per-User lock. The Apple replay suite now tests two same-source MOVE requests and expects one 201 and one 412.

The benchmark records 50k PROPFIND and calendar-query, 10k sync changes, 50-client move convergence, bounded 429 retries, trace stages, response sizes, CPU and RSS. Results and the baseline entry are in docs/perf/caldav-scale-573.md and docs/perf/baseline.json.

Measurements

The full local profile after MOVE lock removal measured PROPFIND 50k p95 5.969s, calendar-query 50k p95 4.806s, sync after 10k changes p95 1.965s, and move visibility to 50 clients p95 18.349s. Listing and sync met their budgets. The move result improved from 34.075s p95 with the DAV lock, but it did not meet the 5s budget. The perf VM returned No route to host; the local host was busy (load average 15.13/17.57/17.22 at start).

All five moves and all 50 sync caches converged. The Calendar projection matched. The If-Match race had one winner and seven 412 responses. Five hostile REPORT cases returned 400 or 413. The corrected 1k diagnostic shows high local App Password verification and Notes writer/index time; see the write-stage artifact. Do not use these loaded local numbers as a perf VM acceptance result.

Gates

cargo fmt --check
(no output; exit 0)

cargo clippy -p calternal-dav --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 9.02s

cargo test -p calternal-dav
 test result: ok. 41 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.23s
 test result: ok. 37 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.11s

cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 2m 24s

cargo test -p calternal-plugin-notes
 test result: ok. 163 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 72.53s
 test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.41s

cargo clippy -p calternal-server --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 27s

cargo test -p calternal-server
 test result: ok. 107 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 13.96s

cargo clean
Removed 24347 files, 11.0GiB total

Decisions

  • Keep the Notes per-User lock around the Markdown and projection update. It preserves write and index order. The DAV MOVE name lock can be removed because the Notes source ETag check provides the compare-and-swap.
  • The benchmark retries bounded authentication-capacity 429 responses and counts retry time in visibility latency. This keeps the profile running while retaining the backpressure cost in its result.
  • Do not change App Password verification or its concurrency limit in this job. Its p95 was high on the loaded local host; changing authentication behavior is outside this route and Notes performance fix.

Known gap

The 5-second area-move visibility budget is still open. Re-run the full profile on the perf VM when it is reachable. Notes 0024_dav_resource_projection.sql is still the next free migration after origin/dev through 0023.

## Finished Implemented Forgejo #573 on branch `job/davscale-573`. Head: `f943e0dab500f69835462dd3e46b91fa117b9e9f`. The Notes Journal now reads sync tokens from the per-User cursor and serves area listings and copy lookups from indexed projections. Healthy DAV reads skip the Notes writer lock check when no repair is pending. MOVE no longer holds the coarse DAV area-name lock. Notes still checks the source ETag while it writes under its per-User lock. The Apple replay suite now tests two same-source MOVE requests and expects one 201 and one 412. The benchmark records 50k PROPFIND and calendar-query, 10k sync changes, 50-client move convergence, bounded 429 retries, trace stages, response sizes, CPU and RSS. Results and the baseline entry are in `docs/perf/caldav-scale-573.md` and `docs/perf/baseline.json`. ## Measurements The full local profile after MOVE lock removal measured PROPFIND 50k p95 5.969s, calendar-query 50k p95 4.806s, sync after 10k changes p95 1.965s, and move visibility to 50 clients p95 18.349s. Listing and sync met their budgets. The move result improved from 34.075s p95 with the DAV lock, but it did not meet the 5s budget. The perf VM returned `No route to host`; the local host was busy (load average 15.13/17.57/17.22 at start). All five moves and all 50 sync caches converged. The Calendar projection matched. The If-Match race had one winner and seven 412 responses. Five hostile REPORT cases returned 400 or 413. The corrected 1k diagnostic shows high local App Password verification and Notes writer/index time; see the write-stage artifact. Do not use these loaded local numbers as a perf VM acceptance result. ## Gates ```text cargo fmt --check (no output; exit 0) cargo clippy -p calternal-dav --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 9.02s cargo test -p calternal-dav test result: ok. 41 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.23s test result: ok. 37 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.11s cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 2m 24s cargo test -p calternal-plugin-notes test result: ok. 163 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 72.53s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.41s cargo clippy -p calternal-server --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 27s cargo test -p calternal-server test result: ok. 107 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 13.96s cargo clean Removed 24347 files, 11.0GiB total ``` ## Decisions - Keep the Notes per-User lock around the Markdown and projection update. It preserves write and index order. The DAV MOVE name lock can be removed because the Notes source ETag check provides the compare-and-swap. - The benchmark retries bounded authentication-capacity 429 responses and counts retry time in visibility latency. This keeps the profile running while retaining the backpressure cost in its result. - Do not change App Password verification or its concurrency limit in this job. Its p95 was high on the loaded local host; changing authentication behavior is outside this route and Notes performance fix. ## Known gap The 5-second area-move visibility budget is still open. Re-run the full profile on the perf VM when it is reachable. Notes `0024_dav_resource_projection.sql` is still the next free migration after `origin/dev` through `0023`.
Author
Owner

Correction to the gate block in my previous comment: it added one leading space before test result lines. The emitted summary lines are:

cargo fmt --check
(no output; exit 0)

cargo clippy -p calternal-dav --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 9.02s
cargo test -p calternal-dav
test result: ok. 41 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.23s
test result: ok. 37 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.11s

cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 2m 24s
cargo test -p calternal-plugin-notes
test result: ok. 163 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 72.53s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.41s

cargo clippy -p calternal-server --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 27s
cargo test -p calternal-server
test result: ok. 107 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 13.96s
Correction to the gate block in my previous comment: it added one leading space before test result lines. The emitted summary lines are: ```text cargo fmt --check (no output; exit 0) cargo clippy -p calternal-dav --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 9.02s cargo test -p calternal-dav test result: ok. 41 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.23s test result: ok. 37 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.11s cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 2m 24s cargo test -p calternal-plugin-notes test result: ok. 163 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 72.53s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.41s cargo clippy -p calternal-server --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 27s cargo test -p calternal-server test result: ok. 107 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 13.96s ```
Author
Owner

Orchestrator: queued for merge round 7 as an improvement (move visibility 34 s → 18 s, listing/sync within budget). Remaining: move visibility to 50 clients p95 18.3 s vs 5 s budget, and 50k PROPFIND p95 6 s — both measured on the loaded build host; re-measure on the perf VM when it is back online, then a follow-up round. Note for the merge: notes migration 0024 collides with #549's notes migration 24 → renumber at merge.

Orchestrator: queued for merge round 7 as an improvement (move visibility 34 s → 18 s, listing/sync within budget). Remaining: move visibility to 50 clients p95 18.3 s vs 5 s budget, and 50k PROPFIND p95 6 s — both measured on the loaded build host; re-measure on the perf VM when it is back online, then a follow-up round. Note for the merge: notes migration 0024 collides with #549's notes migration 24 → renumber at merge.
Author
Owner

Perf VM baseline A1 (partial)

Pinned origin/dev c4a61e8cf, optimized release server, HDD emulation, 50,000 Calendar entries, 10,000 sync changes, 3 samples per listing/sync scenario. The 50-client move-visibility phase ended with TimeoutError at the profile's 120 s request timeout, so this run is incomplete.

Scenario p50 ms p95 ms Samples
PROPFIND, 50k 10,566.304 13,953.381 3
calendar-query, 50k 6,169.164 6,463.966 3
sync after 10k changes 11,047.479 12,928.446 3
five moves visible to 50 clients timeout timeout incomplete

Fixture seed took 691.451 s; projection settle took 16.400 s. Server peak RSS was 621,699,072 bytes across the completed scenarios. Load in the lock: start 0.04 0.03 0.15, end 1.61 2.07 1.86.

Command: flock -w 14400 /root/perf.lock /root/hdd-emu.sh run-limited env TMPDIR=/srv/hdd-emu/tmp python3 bench/caldav-scale-573.py --server calternal-server --data-root /srv/hdd-emu/perf-rerun/caldav --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50.

This is baseline evidence only; the feature comparison is still running. No regression issue is filed from this partial A run.

### Perf VM baseline A1 (partial) Pinned `origin/dev` c4a61e8cf, optimized release server, HDD emulation, 50,000 Calendar entries, 10,000 sync changes, 3 samples per listing/sync scenario. The 50-client move-visibility phase ended with `TimeoutError` at the profile's 120 s request timeout, so this run is incomplete. | Scenario | p50 ms | p95 ms | Samples | | --- | ---: | ---: | ---: | | PROPFIND, 50k | 10,566.304 | 13,953.381 | 3 | | calendar-query, 50k | 6,169.164 | 6,463.966 | 3 | | sync after 10k changes | 11,047.479 | 12,928.446 | 3 | | five moves visible to 50 clients | timeout | timeout | incomplete | Fixture seed took 691.451 s; projection settle took 16.400 s. Server peak RSS was 621,699,072 bytes across the completed scenarios. Load in the lock: start `0.04 0.03 0.15`, end `1.61 2.07 1.86`. Command: `flock -w 14400 /root/perf.lock /root/hdd-emu.sh run-limited env TMPDIR=/srv/hdd-emu/tmp python3 bench/caldav-scale-573.py --server calternal-server --data-root /srv/hdd-emu/perf-rerun/caldav --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50`. This is baseline evidence only; the feature comparison is still running. No regression issue is filed from this partial A run.
Author
Owner

CalDAV A1/B1, perf VM, pair 1

Pinned dev baseline: c4a61e8cf; feature: f943e0dab. Both used optimized release servers, the same 50k Calendar entries and 10k change workload, 3 request samples, and the emulated HDD. Initial/end load for A1 was 0.04 0.03 0.15 / 1.61 2.07 1.86; B1 was 0.16 1.09 1.51 / 2.52 2.33 2.15.

Scenario Dev p95 ms Feature p95 ms Change Result
PROPFIND, 50k 13,953.381 2,195.127 -84.3% Feature run passed
calendar-query, 50k 6,463.966 2,152.995 -66.7% Feature run passed
sync after 10k changes 12,928.446 2,554.542 -80.2% Feature run passed
five moves visible to 50 clients timeout at 120 s request cap 60,103.550 — 5 moves/50 caches converged; still above the 5 s budget

The feature run returned exit 2 because its 2 s sync and 5 s move budget checks failed; correctness checks passed. Move writer p95 was 39,230.388 ms. Feature server peak RSS across the move phase was 795,365,376 bytes. This is one interleaved pair only; four pairs remain. The p95 improvements are strong, but move visibility still misses the requested budget. No >10% regression is identified from this pair.

Command shape: flock -w 14400 /root/perf.lock /root/hdd-emu.sh run-limited env TMPDIR=/srv/hdd-emu/tmp python3 bench/caldav-scale-573.py --server <release-binary> --data-root /srv/hdd-emu/perf-rerun/caldav --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50.

### CalDAV A1/B1, perf VM, pair 1 Pinned dev baseline: `c4a61e8cf`; feature: `f943e0dab`. Both used optimized release servers, the same 50k Calendar entries and 10k change workload, 3 request samples, and the emulated HDD. Initial/end load for A1 was `0.04 0.03 0.15` / `1.61 2.07 1.86`; B1 was `0.16 1.09 1.51` / `2.52 2.33 2.15`. | Scenario | Dev p95 ms | Feature p95 ms | Change | Result | | --- | ---: | ---: | ---: | --- | | PROPFIND, 50k | 13,953.381 | 2,195.127 | -84.3% | Feature run passed | | calendar-query, 50k | 6,463.966 | 2,152.995 | -66.7% | Feature run passed | | sync after 10k changes | 12,928.446 | 2,554.542 | -80.2% | Feature run passed | | five moves visible to 50 clients | timeout at 120 s request cap | 60,103.550 | — | 5 moves/50 caches converged; still above the 5 s budget | The feature run returned exit 2 because its 2 s sync and 5 s move budget checks failed; correctness checks passed. Move writer p95 was 39,230.388 ms. Feature server peak RSS across the move phase was 795,365,376 bytes. This is one interleaved pair only; four pairs remain. The p95 improvements are strong, but move visibility still misses the requested budget. No >10% regression is identified from this pair. Command shape: `flock -w 14400 /root/perf.lock /root/hdd-emu.sh run-limited env TMPDIR=/srv/hdd-emu/tmp python3 bench/caldav-scale-573.py --server <release-binary> --data-root /srv/hdd-emu/perf-rerun/caldav --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50`.
Author
Owner

Correction to my A1/B1 comment: the same-phase server RSS peak was also higher on the feature build: PROPFIND 524,451,840 → 623,144,960 bytes (+18.8%); calendar-query 576,417,792 → 687,452,160 bytes (+19.3%). I treated the p95 latency gains as the whole comparison and omitted this >10% resource increase. I will use the remaining interleaved pairs to confirm whether this is repeatable, then file one issue if it persists.

Correction to my A1/B1 comment: the same-phase server RSS peak was also higher on the feature build: PROPFIND 524,451,840 → 623,144,960 bytes (+18.8%); calendar-query 576,417,792 → 687,452,160 bytes (+19.3%). I treated the p95 latency gains as the whole comparison and omitted this >10% resource increase. I will use the remaining interleaved pairs to confirm whether this is repeatable, then file one issue if it persists.
Author
Owner

One more resource detail from A1/B1: peak server CPU utilization also rose in the same completed phases: PROPFIND 111.87% → 195.73%, calendar-query 107.80% → 199.70%, sync 103.88% → 192.99%. The sampler reports peak concurrent CPU use (values over 100% span cores), not total CPU seconds. Latency improved sharply in these phases, so I will include both sides in the repeated comparison before deciding whether to file the resource increase.

One more resource detail from A1/B1: peak server CPU utilization also rose in the same completed phases: PROPFIND 111.87% → 195.73%, calendar-query 107.80% → 199.70%, sync 103.88% → 192.99%. The sampler reports peak concurrent CPU use (values over 100% span cores), not total CPU seconds. Latency improved sharply in these phases, so I will include both sides in the repeated comparison before deciding whether to file the resource increase.
Author
Owner

CalDAV baseline A2 repetition

A2 again completed the three 50k/10k scenarios, then ended with TimeoutError during the 50-client move phase at the 120 s request cap. Fixture seed took 683.295 s and projection settle took 14.553 s. Load in the lock was 0.08 1.16 1.71 at start and 1.98 2.25 2.25 at end.

Scenario p50 ms p95 ms Peak RSS bytes Peak CPU %
PROPFIND, 50k 10,150.646 11,926.581 536,039,424 103.88
calendar-query, 50k 4,454.171 6,646.389 601,407,488 103.87
sync after 10k changes 10,738.069 15,031.242 548,708,352 111.86
five moves visible to 50 clients timeout timeout incomplete incomplete

Command: flock -w 14400 /root/perf.lock /root/hdd-emu.sh run-limited env TMPDIR=/srv/hdd-emu/tmp python3 bench/caldav-scale-573.py --server <dev release binary> --data-root /srv/hdd-emu/perf-rerun/caldav --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50.

### CalDAV baseline A2 repetition A2 again completed the three 50k/10k scenarios, then ended with `TimeoutError` during the 50-client move phase at the 120 s request cap. Fixture seed took 683.295 s and projection settle took 14.553 s. Load in the lock was `0.08 1.16 1.71` at start and `1.98 2.25 2.25` at end. | Scenario | p50 ms | p95 ms | Peak RSS bytes | Peak CPU % | | --- | ---: | ---: | ---: | ---: | | PROPFIND, 50k | 10,150.646 | 11,926.581 | 536,039,424 | 103.88 | | calendar-query, 50k | 4,454.171 | 6,646.389 | 601,407,488 | 103.87 | | sync after 10k changes | 10,738.069 | 15,031.242 | 548,708,352 | 111.86 | | five moves visible to 50 clients | timeout | timeout | incomplete | incomplete | Command: `flock -w 14400 /root/perf.lock /root/hdd-emu.sh run-limited env TMPDIR=/srv/hdd-emu/tmp python3 bench/caldav-scale-573.py --server <dev release binary> --data-root /srv/hdd-emu/perf-rerun/caldav --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50`.
Author
Owner

CalDAV feature B2 setup timeout

B2 did not reach a measured scenario. The release feature server created 25,000 of the requested 50,000 Calendar entries, then an API batch request ended with TimeoutError at the 120 s client cap. The profile wrote status=incomplete, no scenarios, and no fixture_seed_seconds. Load in the lock: start 0.93 1.94 2.14, end 1.97 2.19 2.24.

Phase Result
50k fixture creation stopped at 25k / 50k with TimeoutError
PROPFIND / calendar-query / sync / 50-client moves not measured

Command: flock -w 14400 /root/perf.lock /root/hdd-emu.sh run-limited env TMPDIR=/srv/hdd-emu/tmp python3 bench/caldav-scale-573.py --server <feature release binary> --data-root /srv/hdd-emu/perf-rerun/caldav --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50.

This is a setup timeout under load, not a p95 sample; it is excluded from the comparison table. B1 completed the same fixture and all scenarios, so this is a run-to-run stability finding to track with B3.

### CalDAV feature B2 setup timeout B2 did not reach a measured scenario. The release feature server created 25,000 of the requested 50,000 Calendar entries, then an API batch request ended with `TimeoutError` at the 120 s client cap. The profile wrote `status=incomplete`, no scenarios, and no `fixture_seed_seconds`. Load in the lock: start `0.93 1.94 2.14`, end `1.97 2.19 2.24`. | Phase | Result | | --- | --- | | 50k fixture creation | stopped at 25k / 50k with `TimeoutError` | | PROPFIND / calendar-query / sync / 50-client moves | not measured | Command: `flock -w 14400 /root/perf.lock /root/hdd-emu.sh run-limited env TMPDIR=/srv/hdd-emu/tmp python3 bench/caldav-scale-573.py --server <feature release binary> --data-root /srv/hdd-emu/perf-rerun/caldav --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50`. This is a setup timeout under load, not a p95 sample; it is excluded from the comparison table. B1 completed the same fixture and all scenarios, so this is a run-to-run stability finding to track with B3.
Author
Owner

Correction: I described B2 as “under load.” The recorded load averages were 0.93/1.94/2.14 at start and 1.97/2.19/2.24 at end on a 4-vCPU VM; that does not establish external contention. The evidence is only that the 25k fixture batch timed out on the emulated HDD. The cause is unassigned pending the repeated run.

Correction: I described B2 as “under load.” The recorded load averages were 0.93/1.94/2.14 at start and 1.97/2.19/2.24 at end on a 4-vCPU VM; that does not establish external contention. The evidence is only that the 25k fixture batch timed out on the emulated HDD. The cause is unassigned pending the repeated run.
Author
Owner

A3 completed the 50k listing/query and 10k sync samples; the requested 50-client move scenario again timed out at the 120 s request limit. These read/sync results are valid samples from the baseline build even though the run is incomplete.

Scenario p50 / p95 ms Peak RSS Peak CPU
PROPFIND 50k 13,438.980 / 14,051.047 567,435,264 B 103.89%
calendar-query 50k 6,962.884 / 8,068.695 573,546,496 B 103.87%
sync after 10k changes 11,558.647 / 12,547.555 646,651,904 B 107.88%

Fixture seed was 688.055 s. Load average was 0.80/1.81/2.10 at start and 1.62/2.10/2.31 at end. Command: python3 bench/caldav-scale-573.py --server calternal-server --data-root /srv/hdd-emu/perf-rerun/caldav --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50 (optimized release binary, HDD emulation, under flock -w 14400 /root/perf.lock). JSON: /root/perf-rerun/output/caldav-A3.json.

A3 completed the 50k listing/query and 10k sync samples; the requested 50-client move scenario again timed out at the 120 s request limit. These read/sync results are valid samples from the baseline build even though the run is incomplete. | Scenario | p50 / p95 ms | Peak RSS | Peak CPU | | --- | ---: | ---: | ---: | | PROPFIND 50k | 13,438.980 / 14,051.047 | 567,435,264 B | 103.89% | | calendar-query 50k | 6,962.884 / 8,068.695 | 573,546,496 B | 103.87% | | sync after 10k changes | 11,558.647 / 12,547.555 | 646,651,904 B | 107.88% | Fixture seed was 688.055 s. Load average was 0.80/1.81/2.10 at start and 1.62/2.10/2.31 at end. Command: `python3 bench/caldav-scale-573.py --server calternal-server --data-root /srv/hdd-emu/perf-rerun/caldav --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50` (optimized release binary, HDD emulation, under `flock -w 14400 /root/perf.lock`). JSON: `/root/perf-rerun/output/caldav-A3.json`.
Author
Owner

Results

Quiet perf VM; optimized release builds; HDD emulation qualified before the runs. The table shows the median run-level p50/p95 for each read/sync scenario. Each completed run used three request samples. Baseline had three attempts; feature had two completed attempts because B2 timed out during fixture creation at 25k/50k entries.

Build PROPFIND 50k p50 / p95 calendar-query 50k p50 / p95 sync after 10k changes p50 / p95 five moves, visible to 50 clients p50 / p95
dev c4a61e8 (A1–A3) 10,566 / 13,953 ms 6,169 / 6,646 ms 11,047 / 12,928 ms no valid sample; all three timed out at 120 s
feature f943e0d (B1, B3) 2,045 / 2,591 ms 1,853 / 2,265 ms 2,231 / 2,673 ms 52,087 / 58,736 ms

Feature B1/B3 passed the requested 5 s p95 budget for listing, query and sync. Move visibility missed it by about 11.7×; both completed feature runs verified all five moves and 50 sync caches. B3 exited 2 because the benchmark's tighter internal budgets failed, not because the scenarios or correctness checks failed.

Mean RSS rose 27.6% for PROPFIND, 16.6% for calendar-query and 15.9% for sync against the three-run baseline medians. Peak sampled CPU rose 99.6%, 76.9% and 83.8%, respectively. I filed #711 for the RSS regression and #712 for the peak CPU signal. The harness records peak CPU in 250 ms windows, not average CPU-seconds; #712 tracks that measurement gap. Start/end load averages are in the run JSON; starts ranged 0.04–0.80 for dev and 0.16–0.93 for feature attempts.

Commands

Both variants used python3 bench/caldav-scale-573.py --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50, TMPDIR=/srv/hdd-emu/tmp, /root/hdd-emu.sh run-limited, and flock -w 14400 /root/perf.lock. Dev used /root/perf-rerun/bin/dev-calternal-server at c4a61e8cf090170f35b1bed3350d9de20c83ecd5; feature used /root/perf-rerun/bin/davscale-573-calternal-server at f943e0dab500f69835462dd3e46b91fa117b9e9f. JSON: /root/perf-rerun/output/caldav-A1.json through caldav-A3.json, caldav-B1.json, caldav-B2.json, caldav-B3.json.

## Results Quiet perf VM; optimized release builds; HDD emulation qualified before the runs. The table shows the median run-level p50/p95 for each read/sync scenario. Each completed run used three request samples. Baseline had three attempts; feature had two completed attempts because B2 timed out during fixture creation at 25k/50k entries. | Build | PROPFIND 50k p50 / p95 | calendar-query 50k p50 / p95 | sync after 10k changes p50 / p95 | five moves, visible to 50 clients p50 / p95 | | --- | ---: | ---: | ---: | ---: | | dev `c4a61e8` (A1–A3) | 10,566 / 13,953 ms | 6,169 / 6,646 ms | 11,047 / 12,928 ms | no valid sample; all three timed out at 120 s | | feature `f943e0d` (B1, B3) | 2,045 / 2,591 ms | 1,853 / 2,265 ms | 2,231 / 2,673 ms | 52,087 / 58,736 ms | Feature B1/B3 passed the requested 5 s p95 budget for listing, query and sync. Move visibility missed it by about 11.7×; both completed feature runs verified all five moves and 50 sync caches. B3 exited 2 because the benchmark's tighter internal budgets failed, not because the scenarios or correctness checks failed. Mean RSS rose 27.6% for PROPFIND, 16.6% for calendar-query and 15.9% for sync against the three-run baseline medians. Peak sampled CPU rose 99.6%, 76.9% and 83.8%, respectively. I filed [#711](#711) for the RSS regression and [#712](#712) for the peak CPU signal. The harness records peak CPU in 250 ms windows, not average CPU-seconds; #712 tracks that measurement gap. Start/end load averages are in the run JSON; starts ranged 0.04–0.80 for dev and 0.16–0.93 for feature attempts. ## Commands Both variants used `python3 bench/caldav-scale-573.py --entries 50000 --samples 3 --changes 10000 --batch-size 1000 --clients 50`, `TMPDIR=/srv/hdd-emu/tmp`, `/root/hdd-emu.sh run-limited`, and `flock -w 14400 /root/perf.lock`. Dev used `/root/perf-rerun/bin/dev-calternal-server` at `c4a61e8cf090170f35b1bed3350d9de20c83ecd5`; feature used `/root/perf-rerun/bin/davscale-573-calternal-server` at `f943e0dab500f69835462dd3e46b91fa117b9e9f`. JSON: `/root/perf-rerun/output/caldav-A1.json` through `caldav-A3.json`, `caldav-B1.json`, `caldav-B2.json`, `caldav-B3.json`.
Author
Owner

Sync architecture audit for #663, base c4a61e8cf; round-7a 2f4482ded checked.

Additional source evidence for #573:

  • round-7a crates/plugins/notes/src/tasks_dav.rs:156 fetches every retained reminder_changes row newer than the token, then folds by UID. The table retains 10,000 changes (tasks_store.rs, delete at base :683). Scratch SQLite, using the actual migration and query: 10,000 changes of one UID fetch 10,000 rows before one UID remains. EXPLAIN: SEARCH reminder_changes USING INDEX sqlite_autoindex_reminder_changes_1 (user_id=? AND seq>?). The issue is work/response bounds, not a missing seek index.
  • Journal changes at round-7a notes/src/lib.rs:1353 (queued #653 :1058) use an indexed seq range but fetch_all with no page cap. Initial Reminder sync calls list_inner and buffers all Tasks.
  • round-7a calternal-dav/src/protocol.rs:1185 area_report and the Reminder REPORT path (:2551 onward) build whole XML responses. The request body is capped at 512 KiB; that is not a response/work cap. I found no nresults handling in these server REPORT paths.

Keep #573 as the single DAV scale owner. Add RFC-compatible limited sync pages, a response-byte bound and a continuation token that advances only through emitted records. Combine duplicates without fetching the full retained history when possible. Verify tombstones, scope/epoch changes, range/multiget behavior, initial sync, and concurrent writes. #653 publication-before-ACK must remain intact. A 100-row app rule is not a reason to silently truncate DAV: its client protocol needs an explicit continuation.

This is source and query-plan evidence, not a timing measurement. No latency-budget or regression ratio is claimed. No hostile traffic was run by this audit.

Sync architecture audit for #663, base c4a61e8cf; round-7a 2f4482ded checked. Additional source evidence for #573: - round-7a `crates/plugins/notes/src/tasks_dav.rs:156` fetches every retained `reminder_changes` row newer than the token, then folds by UID. The table retains 10,000 changes (`tasks_store.rs`, delete at base :683). Scratch SQLite, using the actual migration and query: 10,000 changes of one UID fetch 10,000 rows before one UID remains. EXPLAIN: `SEARCH reminder_changes USING INDEX sqlite_autoindex_reminder_changes_1 (user_id=? AND seq>?)`. The issue is work/response bounds, not a missing seek index. - Journal changes at round-7a `notes/src/lib.rs:1353` (queued #653 :1058) use an indexed seq range but fetch_all with no page cap. Initial Reminder sync calls list_inner and buffers all Tasks. - round-7a `calternal-dav/src/protocol.rs:1185` area_report and the Reminder REPORT path (:2551 onward) build whole XML responses. The request body is capped at 512 KiB; that is not a response/work cap. I found no nresults handling in these server REPORT paths. Keep #573 as the single DAV scale owner. Add RFC-compatible limited sync pages, a response-byte bound and a continuation token that advances only through emitted records. Combine duplicates without fetching the full retained history when possible. Verify tombstones, scope/epoch changes, range/multiget behavior, initial sync, and concurrent writes. #653 publication-before-ACK must remain intact. A 100-row app rule is not a reason to silently truncate DAV: its client protocol needs an explicit continuation. This is source and query-plan evidence, not a timing measurement. No latency-budget or regression ratio is claimed. No hostile traffic was run by this audit.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#573
No description provided.