PERF: use one resettable flush task per collaboration room (#663) #778

Open
opened 2026-10-02 13:10:32 +00:00 by kayg · 2 comments
Owner

Context: #663 coalescing/backpressure audit (rules 5 and 8). Source evidence at c4a61e8cf and job/merge-round-7a 2f4482ded.

Evidence at round-7a crates/calternal-collab/src/session.rs:

  • :1176 schedule_flush starts a new Tokio task for each dirty update (:1184). Generation invalidates old tasks after their sleep, but does not cancel them before they run. At the three-second cap, overdue tasks can all attempt the same flush lock.
  • :2881 calls schedule_flush after each changed WebSocket frame. An authenticated client has a message byte cap; that cap does not bound the count of pending watcher hints or timer tasks.
  • The last-client retry path already uses one retrying flag per room (:2890 onward). Reuse that ownership discipline.

Reasoned impact: high edit rates allocate many timers holding Hub/Room references. Slow filesystem or projection publication increases pending work. Generation checks skip stale work only after the timers wake. This is unbounded per-room scheduling; this audit did not run a hostile load or demonstrate process exhaustion, so this is not labelled BLOCKER.

Concrete fix: one resettable flush task per room with a first-dirty deadline and bounded retries. Preserve the 750 ms quiet/3 s max-save guarantees and unload/shutdown invariants.

Tests: use a controlled slow flush, then send a bounded burst. Assert a fixed task count per room and convergence to the last source revision. Continuous edits must still save by the maximum wait; shutdown and last-client leave must flush; duplicate events and rename/delete must converge. Measure pending tasks high-water and CPU/RSS with the existing collaboration fixture.

Duplicate search: #265 large-Note sync, #270 public frame limit, #668 shared app delta. None replaces per-frame flush timers. Concurrent audit #761 owns the watcher queue; reuse its ingress and overflow contract, but keep this issue scoped to room persistence scheduling.

Context: #663 coalescing/backpressure audit (rules 5 and 8). Source evidence at c4a61e8cf and job/merge-round-7a 2f4482ded. Evidence at round-7a crates/calternal-collab/src/session.rs: - :1176 schedule_flush starts a new Tokio task for each dirty update (:1184). Generation invalidates old tasks after their sleep, but does not cancel them before they run. At the three-second cap, overdue tasks can all attempt the same flush lock. - :2881 calls schedule_flush after each changed WebSocket frame. An authenticated client has a message byte cap; that cap does not bound the count of pending watcher hints or timer tasks. - The last-client retry path already uses one retrying flag per room (:2890 onward). Reuse that ownership discipline. Reasoned impact: high edit rates allocate many timers holding Hub/Room references. Slow filesystem or projection publication increases pending work. Generation checks skip stale work only after the timers wake. This is unbounded per-room scheduling; this audit did not run a hostile load or demonstrate process exhaustion, so this is not labelled BLOCKER. Concrete fix: one resettable flush task per room with a first-dirty deadline and bounded retries. Preserve the 750 ms quiet/3 s max-save guarantees and unload/shutdown invariants. Tests: use a controlled slow flush, then send a bounded burst. Assert a fixed task count per room and convergence to the last source revision. Continuous edits must still save by the maximum wait; shutdown and last-client leave must flush; duplicate events and rename/delete must converge. Measure pending tasks high-water and CPU/RSS with the existing collaboration fixture. Duplicate search: #265 large-Note sync, #270 public frame limit, #668 shared app delta. None replaces per-frame flush timers. Concurrent audit #761 owns the watcher queue; reuse its ingress and overflow contract, but keep this issue scoped to room persistence scheduling.
kayg changed title from PERF: bound collaboration watcher hints and flush scheduling (#663) to PERF: use one resettable flush task per collaboration room (#663) 2026-10-02 13:11:31 +00:00
Author
Owner

Starting work on job/notesperf, based on job/merge-round-7a at 2f4482ded066d9c5d9c59130377907f7fd2916c9. I will address the issue with focused changes and regression coverage, then report the final head SHA and verbatim gate output here.

Starting work on `job/notesperf`, based on `job/merge-round-7a` at `2f4482ded066d9c5d9c59130377907f7fd2916c9`. I will address the issue with focused changes and regression coverage, then report the final head SHA and verbatim gate output here.
Author
Owner

Each dirty room now owns one resettable debounce task with a three-second maximum dirty deadline. A 100-call burst regression test passed: cargo test -p calternal-collab --lib flush_burst_starts_one_room_timer_task (1 passed). Commit: 750b36129. The full crate gate and performance profile are pending.

Each dirty room now owns one resettable debounce task with a three-second maximum dirty deadline. A 100-call burst regression test passed: `cargo test -p calternal-collab --lib flush_burst_starts_one_room_timer_task` (1 passed). Commit: 750b36129. The full crate gate and performance profile are pending.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#778
No description provided.