PERF: move WAL checkpoint work off interactive commits (#663) #823

Open
opened 2026-10-02 13:18:04 +00:00 by kayg · 3 comments
Owner

Context: SQLite architecture audit #663, DESIGN §58 rules 4 and 8; related HDD work #549. Base c4a61e8cf0 and queued merge-round-7a 2f4482ded0.

Evidence: crates/calternal-db/src/db.rs:48–69 creates the sole writer with WAL and synchronous=NORMAL. There is no wal_autocheckpoint setting, explicit wal_checkpoint, checkpoint worker, or checkpoint timing in the inspected production source. Round 7a db.rs:144–171 retains that policy; its new page and statement cache limits do not change checkpoint scheduling. The SQLx 0.9.0 local options source has no automatic checkpoint override.

Reasoned impact: SQLite's default threshold is 1,000 pages. The commit that crosses it runs a passive checkpoint on the committing connection. This can add sync and scattered Index writes to an interactive mutation and keeps the sole writer unavailable to all other Users. A long WAL reader can stop checkpoint progress. This is a source-level latency risk on HDD, not a measured regression or evidence that every commit is slow. WAL readers do not normally wait for a write transaction's lock.

Primary source: https://www.sqlite.org/wal.html sections 2.2–2.3 and 3.1. Read/write concurrency, commit-triggered checkpoints, and the NORMAL/FULL distinction are separate concerns.

Concrete fix: give calternal-db one checkpoint policy and its own writable maintenance connection. Schedule bounded PASSIVE work outside interactive commits, record remaining/processed frames and duration without User content, and retry later when readers pin frames. Set a tested WAL growth limit and a shutdown policy. Do not simply turn off auto-checkpointing and leave an unbounded WAL. Keep stronger Security state durability as a separate concern; off-thread checkpoints do not make NORMAL commits power-loss durable.

Tests: force a small checkpoint threshold on a scratch Index; hold a reader snapshot, issue bounded writes, and prove the policy reports incomplete progress without blocking the interactive writer. Release the reader and prove checkpoint completion. Check restart/shutdown, failed checkpoints, and a bound on WAL growth. Run at least five accepted/durable mutation samples with large ingest on the locked perf VM with HDD emulation; report median/p95/max, CPU/RSS and load. No production data or destructive power-loss test is needed for this scheduling regression.

Duplicate check: searched all issue titles through #800 for SQLite, WAL, checkpoint, fsync and durability; inspected #549 and #748. #549 owns Tab paint, #748 owns unnecessary startup snapshots; neither specifies WAL checkpoint ownership. Non-blocking performance finding.

Context: SQLite architecture audit #663, DESIGN §58 rules 4 and 8; related HDD work #549. Base c4a61e8cf090170f35b1bed3350d9de20c83ecd5 and queued merge-round-7a 2f4482ded066d9c5d9c59130377907f7fd2916c9. Evidence: crates/calternal-db/src/db.rs:48–69 creates the sole writer with WAL and synchronous=NORMAL. There is no wal_autocheckpoint setting, explicit wal_checkpoint, checkpoint worker, or checkpoint timing in the inspected production source. Round 7a db.rs:144–171 retains that policy; its new page and statement cache limits do not change checkpoint scheduling. The SQLx 0.9.0 local options source has no automatic checkpoint override. Reasoned impact: SQLite's default threshold is 1,000 pages. The commit that crosses it runs a passive checkpoint on the committing connection. This can add sync and scattered Index writes to an interactive mutation and keeps the sole writer unavailable to all other Users. A long WAL reader can stop checkpoint progress. This is a source-level latency risk on HDD, not a measured regression or evidence that every commit is slow. WAL readers do not normally wait for a write transaction's lock. Primary source: https://www.sqlite.org/wal.html sections 2.2–2.3 and 3.1. Read/write concurrency, commit-triggered checkpoints, and the NORMAL/FULL distinction are separate concerns. Concrete fix: give calternal-db one checkpoint policy and its own writable maintenance connection. Schedule bounded PASSIVE work outside interactive commits, record remaining/processed frames and duration without User content, and retry later when readers pin frames. Set a tested WAL growth limit and a shutdown policy. Do not simply turn off auto-checkpointing and leave an unbounded WAL. Keep stronger Security state durability as a separate concern; off-thread checkpoints do not make NORMAL commits power-loss durable. Tests: force a small checkpoint threshold on a scratch Index; hold a reader snapshot, issue bounded writes, and prove the policy reports incomplete progress without blocking the interactive writer. Release the reader and prove checkpoint completion. Check restart/shutdown, failed checkpoints, and a bound on WAL growth. Run at least five accepted/durable mutation samples with large ingest on the locked perf VM with HDD emulation; report median/p95/max, CPU/RSS and load. No production data or destructive power-loss test is needed for this scheduling regression. Duplicate check: searched all issue titles through #800 for SQLite, WAL, checkpoint, fsync and durability; inspected #549 and #748. #549 owns Tab paint, #748 owns unnecessary startup snapshots; neither specifies WAL checkpoint ownership. Non-blocking performance finding.
Author
Owner

Finding: production wire.rs:1233 builds SqliteAuthStore from Db::writer_pool and Db::reader_pool. The shared writer currently sets NORMAL; its connection options do not override SQLite auto-checkpoint (default 1,000 pages). cargo search sqlx --limit 1 confirms 0.9.0; dependencies are unchanged.

Decision for #824: FULL on the shared writer, including replacement connections, because Security state and Derived data share the Index and raw writer pool. Per-transaction switching would require a broader authority-write refactor to be cancellation safe. macOS fullfsync is enabled; sync-honoring storage remains required.

Decision for #823: independent writable maintenance connection, PASSIVE each second, no commit hook, 16 MiB retained-WAL target, progress/error counters and a best-effort shutdown pass. A pinned read snapshot can exceed the retention target: journal_size_limit is not a hard active-WAL cap. A hard bound needs reader expiry or writer admission policy; this limitation will remain explicit.

Finding: production `wire.rs:1233` builds `SqliteAuthStore` from `Db::writer_pool` and `Db::reader_pool`. The shared writer currently sets NORMAL; its connection options do not override SQLite auto-checkpoint (default 1,000 pages). `cargo search sqlx --limit 1` confirms 0.9.0; dependencies are unchanged. Decision for #824: FULL on the shared writer, including replacement connections, because Security state and Derived data share the Index and raw writer pool. Per-transaction switching would require a broader authority-write refactor to be cancellation safe. macOS fullfsync is enabled; sync-honoring storage remains required. Decision for #823: independent writable maintenance connection, PASSIVE each second, no commit hook, 16 MiB retained-WAL target, progress/error counters and a best-effort shutdown pass. A pinned read snapshot can exceed the retention target: `journal_size_limit` is not a hard active-WAL cap. A hard bound needs reader expiry or writer admission policy; this limitation will remain explicit.
Author
Owner

wal-824 final report

Head: fa68c3d41d3fa064b41156baa28408cba8466c9e on job/wal-824. No push, deploy or outgoing merge. The required one fetch and merge of origin/dev completed as 466d3baa8; it changed only media sandbox scripts. The worktree is clean. No issue is closed.

Built

  • #824: the shared Index writer uses FULL and macOS fullfsync. All core and Plugin repositories using Db inherit the policy, including replacement connections. Production SqliteAuthStore receives this same writer pool. No temporary synchronous downgrade is introduced. The stronger NORMAL→FULL test expectation is required by this issue.
  • #823: automatic writer checkpoints are off. One maintenance connection runs PASSIVE each second, records content-free frame/duration/error diagnostics, retries incomplete work and makes a final best-effort pass. Db clones share its lifetime. Close remains joinable after cancellation.
  • Regression coverage: pinned readers, >16 MiB active WAL, writer liveness, completion after release, retention after restart, rollback/error/cancelled transactions, replacement writer, failed maintenance, retry, close cancellation and reopen.
  • Linux sync-boundary harness: successful fsync/fdatasync images only, SIGKILL immediately after ACK, main/WAL recovery from a durable initial WAL baseline. The NORMAL negative control loses the acknowledged change, FULL preserves all four synthetic authority fields, and failed sync receives no ACK.

Files

  • crates/calternal-db/src/db.rs
  • crates/calternal-db/src/checkpoint.rs
  • crates/calternal-db/src/lib.rs
  • crates/calternal-db/Cargo.toml
  • Cargo.lock
  • crates/calternal-db/tests/queue.rs
  • crates/calternal-db/tests/checkpoint.rs
  • crates/calternal-db/examples/wal_probe.rs
  • tests/adversarial/sqlite/durability.py
  • tests/adversarial/sqlite/sync_image.c
  • bench/sqlite-wal.sh
  • docs/DESIGN.md
  • docs/perf/2026-10-02-wal-824.md

Measurements

Locked HDD VM, release builds, SQLite 3.51.3, 100k synthetic 1 KiB rows, 1,000 accepted changed-row commits and a 32-request burst per phase. Lock released between phases. HDD emulation used direct I/O, ext4, 8 ms delays and 200 IOPS. Load before phases: 0.08/0.73/0.66 and 0.00/0.03/0.24.

Metric Before After
Mutation p50 ms 0.179 54.897
Mutation p95 ms 0.363 103.992
Mutation max ms 391.877 578.404
Ingest p95 ms 0.870 110.972
Ingest max ms 1077.680 636.519
Burst p95 ms 4.151 1894.848
Probe seconds 20.178 141.508
CPU user + system seconds 0.88 1.62
Average CPU % (time) 4 1
Peak RSS KiB 8132 8700

The old mutation commits were not power-loss durable. docs/perf/baseline.json has no corresponding SQLite profile; the recorded old-opener run is the comparison. The latency increase is filed as #853. Peak RSS increased about 7%, below the 15% threshold. Full results and repeat commands: docs/perf/2026-10-02-wal-824.md.

Gate output, verbatim

cargo fmt --check: passed, no output (including a final check after documentation edits).

cargo clippy -p calternal-db --all-targets -- -D warnings:

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 27.08s

cargo test -p calternal-db:

    Finished `test` profile [unoptimized + debuginfo] target(s) in 20.35s
     Running unittests src/lib.rs (/mnt/hdd/targets/jobs/wal-824/debug/deps/calternal_db-fc6d78b9df826065)

running 12 tests
test cron::tests::zoned_cron_keeps_its_wall_clock_time_across_daylight_saving ... ok
test sqlite::tests::saturated_or_closed_sqlite_pools_are_transient_service_errors ... ok
test checkpoint::tests::failed_pass_counts_error_and_a_later_pass_recovers ... ok
test sqlite::tests::wrapped_sqlite_contention_retries_the_whole_operation ... ok
test sqlite::tests::persistent_wrapped_sqlite_contention_stops_at_the_attempt_limit ... ok
test checkpoint::tests::cancelled_close_keeps_the_final_pass_joinable ... ok
test secrets::tests::rejects_invalid_names_and_empty_candidates ... ok
test secrets::tests::named_secret_is_stable_and_first_candidate_wins ... ok
test secrets::tests::named_secret_can_be_replaced_and_cleared_without_reading_it_for_status ... ok
test secrets::tests::secret_survives_database_reopen ... ok
test worker::tests::disabled_handler_keeps_jobs_pending_and_finishes_active_work ... ok
test jobs::contention_tests::background_queue_write_waits_past_five_seconds_for_the_writer ... ok

test result: ok. 12 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 12.44s

     Running tests/checkpoint.rs (/mnt/hdd/targets/jobs/wal-824/debug/deps/checkpoint-0fd30139b93d0bfc)

running 3 tests
test durability_policy_survives_errors_and_replacement_connections ... ok
test wal_retention_target_and_reuse_are_configured ... ok
test pinned_reader_reports_remaining_frames_and_retries_after_release ... ok

test result: ok. 3 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 18.63s

     Running tests/queue.rs (/mnt/hdd/targets/jobs/wal-824/debug/deps/queue-37225b1abc3ce8b0)

running 17 tests
test enqueue_lease_throughput_microbenchmark ... ignored, manual enqueue plus lease throughput measurement
test opens_wal_database_with_required_pragmas ... ok
test queue_change_subscribers_receive_a_hint_after_state_changes ... ok
test job_list_summaries_do_not_read_handler_payloads ... ok
test deduplicates_pending_and_leased_jobs ... ok
test plugin_migrations_share_namespaces_and_check_applied_sql ... ok
test cron_enqueues_each_occurrence_once_and_keeps_one_active_job ... ok
test expired_lease_is_recovered_and_old_owner_loses_lease ... ok
test controlled_worker_observes_stop_at_handler_checkpoint ... ok
test failure_backoff_uses_fake_clock_and_dead_letters_at_limit ... ok
test heartbeat_does_not_deadlock_a_handler_inside_a_write_transaction ... ok
test owned_job_progress_is_private_and_cancellation_finishes_at_checkpoint ... ok
test queue_pause_is_idempotent_persistent_and_blocks_new_leases ... ok
test worker_publishes_registered_kind_metadata_to_the_shared_queue ... ok
test queue_retry_run_now_and_clear_only_change_failed_jobs_of_one_kind ... ok
test snapshot_restores_as_a_readable_database ... ok
test workers_do_not_execute_a_job_twice ... ok

test result: ok. 16 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 19.46s

   Doc-tests calternal_db

running 0 tests

test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s

Final sync-image adversarial round:

NORMAL negative control: acknowledged state lost
FULL simulated power loss: recovered: revoked session, revoked App Password, changed password, enabled 2FA
sync failure: no ACK

C shim compilation with -Wall -Wextra -Werror, Python byte compilation, bash syntax and git diff checks passed with no output. No route or external API contract changed, so no server or web gate was run. No real-server hostile request round was run. The adversarial evidence here is the real SQLite fault harness and maintenance regressions. No migration was added.

Cleanup:

     Removed 3171 files, 1.2GiB total

No web build output was created. Review artifacts and saved probe binaries remain ignored under artifacts/. No screenshots or other review files were committed.

Decisions

Use FULL for the shared Index because Security state and Derived data share an exposed writer pool. A narrower boundary needs a cancellation-safe authority lease or separate Security state Index; Auth-only coverage would miss Shares, Roles and Plugin state. The measured HDD increase is recorded for that follow-up rather than weakening durability.

Use one-second PASSIVE maintenance, a 16 MiB retained-WAL target and a best-effort final pass. A commit does not invoke or await a checkpoint.

Known gaps

  • #823 remains partial: there is no hard active-WAL growth bound while a reader pins frames or a transaction grows. Enforcing one needs transaction limits and reader lifetime or writer admission policy outside the current raw-pool API. The >16 MiB test explicitly proves the distinction from retained-WAL trimming. Keep #823 open.
  • The fault harness uses synthetic authority rows through the real shared opener/VFS. It does not exercise actual Auth HTTP mutations or queued receipt recovery. Physical power loss and torn sectors were not tested. Durability assumes storage honors sync, as required by SQLite.
  • Standalone stores outside Db are not changed. The production Security state Index and Db-opened Plugin Index files receive the shared policy.

UX gaps closed / UX gaps left

Not applicable: no UI changed. No screenshot or device-width evidence is required for this backend change.

# wal-824 final report Head: `fa68c3d41d3fa064b41156baa28408cba8466c9e` on `job/wal-824`. No push, deploy or outgoing merge. The required one fetch and merge of origin/dev completed as `466d3baa8`; it changed only media sandbox scripts. The worktree is clean. No issue is closed. ## Built - #824: the shared Index writer uses FULL and macOS fullfsync. All core and Plugin repositories using Db inherit the policy, including replacement connections. Production SqliteAuthStore receives this same writer pool. No temporary synchronous downgrade is introduced. The stronger NORMAL→FULL test expectation is required by this issue. - #823: automatic writer checkpoints are off. One maintenance connection runs PASSIVE each second, records content-free frame/duration/error diagnostics, retries incomplete work and makes a final best-effort pass. Db clones share its lifetime. Close remains joinable after cancellation. - Regression coverage: pinned readers, >16 MiB active WAL, writer liveness, completion after release, retention after restart, rollback/error/cancelled transactions, replacement writer, failed maintenance, retry, close cancellation and reopen. - Linux sync-boundary harness: successful fsync/fdatasync images only, SIGKILL immediately after ACK, main/WAL recovery from a durable initial WAL baseline. The NORMAL negative control loses the acknowledged change, FULL preserves all four synthetic authority fields, and failed sync receives no ACK. ## Files - `crates/calternal-db/src/db.rs` - `crates/calternal-db/src/checkpoint.rs` - `crates/calternal-db/src/lib.rs` - `crates/calternal-db/Cargo.toml` - `Cargo.lock` - `crates/calternal-db/tests/queue.rs` - `crates/calternal-db/tests/checkpoint.rs` - `crates/calternal-db/examples/wal_probe.rs` - `tests/adversarial/sqlite/durability.py` - `tests/adversarial/sqlite/sync_image.c` - `bench/sqlite-wal.sh` - `docs/DESIGN.md` - `docs/perf/2026-10-02-wal-824.md` ## Measurements Locked HDD VM, release builds, SQLite 3.51.3, 100k synthetic 1 KiB rows, 1,000 accepted changed-row commits and a 32-request burst per phase. Lock released between phases. HDD emulation used direct I/O, ext4, 8 ms delays and 200 IOPS. Load before phases: 0.08/0.73/0.66 and 0.00/0.03/0.24. | Metric | Before | After | | --- | ---: | ---: | | Mutation p50 ms | 0.179 | 54.897 | | Mutation p95 ms | 0.363 | 103.992 | | Mutation max ms | 391.877 | 578.404 | | Ingest p95 ms | 0.870 | 110.972 | | Ingest max ms | 1077.680 | 636.519 | | Burst p95 ms | 4.151 | 1894.848 | | Probe seconds | 20.178 | 141.508 | | CPU user + system seconds | 0.88 | 1.62 | | Average CPU % (time) | 4 | 1 | | Peak RSS KiB | 8132 | 8700 | The old mutation commits were not power-loss durable. docs/perf/baseline.json has no corresponding SQLite profile; the recorded old-opener run is the comparison. The latency increase is filed as [#853](https://git.kayg.org/kayg/calternal/issues/853). Peak RSS increased about 7%, below the 15% threshold. Full results and repeat commands: `docs/perf/2026-10-02-wal-824.md`. ## Gate output, verbatim `cargo fmt --check`: passed, no output (including a final check after documentation edits). `cargo clippy -p calternal-db --all-targets -- -D warnings`: ``` Finished `dev` profile [unoptimized + debuginfo] target(s) in 27.08s ``` `cargo test -p calternal-db`: ``` Finished `test` profile [unoptimized + debuginfo] target(s) in 20.35s Running unittests src/lib.rs (/mnt/hdd/targets/jobs/wal-824/debug/deps/calternal_db-fc6d78b9df826065) running 12 tests test cron::tests::zoned_cron_keeps_its_wall_clock_time_across_daylight_saving ... ok test sqlite::tests::saturated_or_closed_sqlite_pools_are_transient_service_errors ... ok test checkpoint::tests::failed_pass_counts_error_and_a_later_pass_recovers ... ok test sqlite::tests::wrapped_sqlite_contention_retries_the_whole_operation ... ok test sqlite::tests::persistent_wrapped_sqlite_contention_stops_at_the_attempt_limit ... ok test checkpoint::tests::cancelled_close_keeps_the_final_pass_joinable ... ok test secrets::tests::rejects_invalid_names_and_empty_candidates ... ok test secrets::tests::named_secret_is_stable_and_first_candidate_wins ... ok test secrets::tests::named_secret_can_be_replaced_and_cleared_without_reading_it_for_status ... ok test secrets::tests::secret_survives_database_reopen ... ok test worker::tests::disabled_handler_keeps_jobs_pending_and_finishes_active_work ... ok test jobs::contention_tests::background_queue_write_waits_past_five_seconds_for_the_writer ... ok test result: ok. 12 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 12.44s Running tests/checkpoint.rs (/mnt/hdd/targets/jobs/wal-824/debug/deps/checkpoint-0fd30139b93d0bfc) running 3 tests test durability_policy_survives_errors_and_replacement_connections ... ok test wal_retention_target_and_reuse_are_configured ... ok test pinned_reader_reports_remaining_frames_and_retries_after_release ... ok test result: ok. 3 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 18.63s Running tests/queue.rs (/mnt/hdd/targets/jobs/wal-824/debug/deps/queue-37225b1abc3ce8b0) running 17 tests test enqueue_lease_throughput_microbenchmark ... ignored, manual enqueue plus lease throughput measurement test opens_wal_database_with_required_pragmas ... ok test queue_change_subscribers_receive_a_hint_after_state_changes ... ok test job_list_summaries_do_not_read_handler_payloads ... ok test deduplicates_pending_and_leased_jobs ... ok test plugin_migrations_share_namespaces_and_check_applied_sql ... ok test cron_enqueues_each_occurrence_once_and_keeps_one_active_job ... ok test expired_lease_is_recovered_and_old_owner_loses_lease ... ok test controlled_worker_observes_stop_at_handler_checkpoint ... ok test failure_backoff_uses_fake_clock_and_dead_letters_at_limit ... ok test heartbeat_does_not_deadlock_a_handler_inside_a_write_transaction ... ok test owned_job_progress_is_private_and_cancellation_finishes_at_checkpoint ... ok test queue_pause_is_idempotent_persistent_and_blocks_new_leases ... ok test worker_publishes_registered_kind_metadata_to_the_shared_queue ... ok test queue_retry_run_now_and_clear_only_change_failed_jobs_of_one_kind ... ok test snapshot_restores_as_a_readable_database ... ok test workers_do_not_execute_a_job_twice ... ok test result: ok. 16 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 19.46s Doc-tests calternal_db running 0 tests test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s ``` Final sync-image adversarial round: ``` NORMAL negative control: acknowledged state lost FULL simulated power loss: recovered: revoked session, revoked App Password, changed password, enabled 2FA sync failure: no ACK ``` C shim compilation with -Wall -Wextra -Werror, Python byte compilation, bash syntax and git diff checks passed with no output. No route or external API contract changed, so no server or web gate was run. No real-server hostile request round was run. The adversarial evidence here is the real SQLite fault harness and maintenance regressions. No migration was added. Cleanup: ``` Removed 3171 files, 1.2GiB total ``` No web build output was created. Review artifacts and saved probe binaries remain ignored under artifacts/. No screenshots or other review files were committed. ## Decisions Use FULL for the shared Index because Security state and Derived data share an exposed writer pool. A narrower boundary needs a cancellation-safe authority lease or separate Security state Index; Auth-only coverage would miss Shares, Roles and Plugin state. The measured HDD increase is recorded for that follow-up rather than weakening durability. Use one-second PASSIVE maintenance, a 16 MiB retained-WAL target and a best-effort final pass. A commit does not invoke or await a checkpoint. ## Known gaps - #823 remains partial: there is no hard active-WAL growth bound while a reader pins frames or a transaction grows. Enforcing one needs transaction limits and reader lifetime or writer admission policy outside the current raw-pool API. The >16 MiB test explicitly proves the distinction from retained-WAL trimming. Keep #823 open. - The fault harness uses synthetic authority rows through the real shared opener/VFS. It does not exercise actual Auth HTTP mutations or queued receipt recovery. Physical power loss and torn sectors were not tested. Durability assumes storage honors sync, as required by SQLite. - Standalone stores outside Db are not changed. The production Security state Index and Db-opened Plugin Index files receive the shared policy. ## UX gaps closed / UX gaps left Not applicable: no UI changed. No screenshot or device-width evidence is required for this backend change.
Author
Owner

Implemented on job/wal-824 (current HEAD 8e8718003; not merged). Both write connections set journal_size_limit=16777216. The independent maintenance connection uses zero lock wait. After PASSIVE copies all active frames above 16 MiB, it attempts TRUNCATE; a new reader or writer causes a later retry rather than a wait on an interactive commit.

Regression pinned_reader_reports_remaining_frames_and_retries_after_release now verifies that releasing a pinned reader reclaims the oversized WAL to at most 16 MiB without another write. cargo test -p calternal-db passed, including the rollback, cancellation, replacement-connection and mixed-policy tests.

This bounds reclaimable active WAL. A reader or one transaction can still pin more than 16 MiB until it ends; it is not an absolute byte cap. Ordinary writes remain WAL/NORMAL, Security state uses a dedicated FULL connection. The locked HDD profile measured ordinary p95 0.401852 ms and authority p95 117.234803 ms. Full evidence: docs/perf/2026-10-02-wal-824.md.

Implemented on `job/wal-824` (current HEAD `8e8718003`; not merged). Both write connections set `journal_size_limit=16777216`. The independent maintenance connection uses zero lock wait. After PASSIVE copies all active frames above 16 MiB, it attempts TRUNCATE; a new reader or writer causes a later retry rather than a wait on an interactive commit. Regression `pinned_reader_reports_remaining_frames_and_retries_after_release` now verifies that releasing a pinned reader reclaims the oversized WAL to at most 16 MiB without another write. `cargo test -p calternal-db` passed, including the rollback, cancellation, replacement-connection and mixed-policy tests. This bounds reclaimable active WAL. A reader or one transaction can still pin more than 16 MiB until it ends; it is not an absolute byte cap. Ordinary writes remain WAL/NORMAL, Security state uses a dedicated FULL connection. The locked HDD profile measured ordinary p95 0.401852 ms and authority p95 117.234803 ms. Full evidence: `docs/perf/2026-10-02-wal-824.md`.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#823
No description provided.