Jobs: leases of in-flight jobs are never recovered after a restart (silent stall) #1042

Closed
opened 2026-10-04 08:09:28 +00:00 by kayg · 12 comments
Owner

Finding (2026-10-04, #1038 stress round, branch job/mailstress-1038 at cae707d27)

Worker::run in calternal-db calls recover_expired_leases only at startup. A job lease lasts 120 s. When the server restarts before the leases of in-flight jobs expire (every deploy, every crash-restart), those jobs stay leased for good: no worker claims them, no error appears, status looks healthy. Only a second restart after expiry frees them.

Observed: with 3 Mail accounts, one account's Inbox stayed at 0 messages for 180+ s with accounts_with_errors=0; three expired mail.sync leases live at 180 s; a manual restart recovered them. A deterministic regression exists:

thread 'restart_before_mail_lease_expiry_does_not_strand_an_account' panicked at crates/calternal-db/tests/mailstress_restart.rs:96:5:
Mail sync lease expired after startup but was never recovered
test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 2.10s

Source: artifacts/mailstress-1038/mailstress_restart.rs and a proposed lease-recovery.patch in that worktree.

Scope

This is the shared Jobs queue, so it affects every job kind (Mail sync, thumbnails, search indexing, projection rebuilds, imports), not only Mail. Production restarts on each deploy.

Wanted

  • Expired leases become claimable again while the server runs: either the claim query treats lease_expires_at < now as claimable, or the worker recovers expired leases periodically (bounded cost, no busy loop). Choose the simplest correct option and explain it in the module doc.
  • A job whose lease expired but whose worker is still running (slow job) must not run twice at the same time: heartbeat/renewal or an attempt token so the stale worker's completion is rejected. State the invariant.
  • Regression tests: the restart case above; a slow job that renews its lease is not stolen; a stolen job's late completion is ignored.
  • Check the production log/DB after deploy (read-only aggregate) for jobs leased past expiry.
## Finding (2026-10-04, #1038 stress round, branch job/mailstress-1038 at cae707d27) `Worker::run` in `calternal-db` calls `recover_expired_leases` **only at startup**. A job lease lasts 120 s. When the server restarts before the leases of in-flight jobs expire (every deploy, every crash-restart), those jobs stay leased for good: no worker claims them, no error appears, status looks healthy. Only a second restart after expiry frees them. Observed: with 3 Mail accounts, one account's Inbox stayed at 0 messages for 180+ s with `accounts_with_errors=0`; three expired `mail.sync` leases live at 180 s; a manual restart recovered them. A deterministic regression exists: ```text thread 'restart_before_mail_lease_expiry_does_not_strand_an_account' panicked at crates/calternal-db/tests/mailstress_restart.rs:96:5: Mail sync lease expired after startup but was never recovered test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 2.10s ``` Source: `artifacts/mailstress-1038/mailstress_restart.rs` and a proposed `lease-recovery.patch` in that worktree. ## Scope This is the shared Jobs queue, so it affects every job kind (Mail sync, thumbnails, search indexing, projection rebuilds, imports), not only Mail. Production restarts on each deploy. ## Wanted - Expired leases become claimable again while the server runs: either the claim query treats `lease_expires_at < now` as claimable, or the worker recovers expired leases periodically (bounded cost, no busy loop). Choose the simplest correct option and explain it in the module doc. - A job whose lease expired but whose worker is still running (slow job) must not run twice at the same time: heartbeat/renewal or an attempt token so the stale worker's completion is rejected. State the invariant. - Regression tests: the restart case above; a slow job that renews its lease is not stolen; a stolen job's late completion is ignored. - Check the production log/DB after deploy (read-only aggregate) for jobs leased past expiry.
Author
Owner

Starting #1042 on job/lease-1042, base 6074f71d18abe73b2b4255acb564f275a9acc851. Read CLAUDE.md, CONTEXT.md and DESIGN Jobs decisions. Evaluating the #1038 restart regression and proposed recovery patch; auditing all Worker registrations for long-running lease renewal. No pushes or deploys.

Starting #1042 on `job/lease-1042`, base `6074f71d18abe73b2b4255acb564f275a9acc851`. Read CLAUDE.md, CONTEXT.md and DESIGN Jobs decisions. Evaluating the #1038 restart regression and proposed recovery patch; auditing all Worker registrations for long-running lease renewal. No pushes or deploys.
Author
Owner

Production evidence (2026-10-04, 6074f71d1, read-only aggregates)

  • After the 08:56 IST deploy, a mail.sync job stayed leased with its lease expired 72 min earlier: the owner's Mail sync was silently stalled on production. 66 mail.sync, 16 notes.reconcile and 6 notes.daily-log-projection-rebuild jobs are dead.
  • Mitigation: restarted production at ~10:15 IST (healthy in 9 s); startup recovery freed the lease, and mail.sync is running and renewing again.
  • After the restart, notes.log-batch-index was leased 2.7 min earlier and its lease had expired 0.7 min earlier while still running. So at least one long job kind does not renew its lease. Include it in the renewal audit; with the claim-side fix it could otherwise run twice.
## Production evidence (2026-10-04, 6074f71d1, read-only aggregates) - After the 08:56 IST deploy, a `mail.sync` job stayed `leased` with its lease **expired 72 min** earlier: the owner's Mail sync was silently stalled on production. 66 `mail.sync`, 16 `notes.reconcile` and 6 `notes.daily-log-projection-rebuild` jobs are `dead`. - Mitigation: restarted production at ~10:15 IST (healthy in 9 s); startup recovery freed the lease, and `mail.sync` is running and renewing again. - After the restart, `notes.log-batch-index` was leased 2.7 min earlier and its lease had **expired 0.7 min** earlier while still running. So at least one long job kind does not renew its lease. Include it in the renewal audit; with the claim-side fix it could otherwise run twice.
Author
Owner

Confirmed the supplied restart regression fails before the fix:


test restart_before_mail_lease_expiry_does_not_strand_an_account ... FAILED

failures:

---- restart_before_mail_lease_expiry_does_not_strand_an_account stdout ----

thread 'restart_before_mail_lease_expiry_does_not_strand_an_account' (923768) panicked at crates/calternal-db/tests/mailstress_restart.rs:96:5:
Mail sync lease expired after startup but was never recovered
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace


failures:
    restart_before_mail_lease_expiry_does_not_strand_an_account

test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 2.07s

error: test failed, to rerun pass `-p calternal-db --test mailstress_restart`

The proposed periodic-recovery patch is suitable: the existing recovery query uses the expired-lease index and handles cancellation and attempt limits. I use Tokio monotonic time and check at the loop head, once per max(lease/3, poll interval), so continuous dispatch cannot starve upkeep.

The shared Worker already heartbeats every job kind at lease/3, including plugin handlers. The server config uses the same server owner for every attempt. I will use a unique claim token in the existing leased_by field to fence old progress, renewal and settlement when that configured name is reused. No schema or public route change is needed.

Confirmed the supplied restart regression fails before the fix: ```text test restart_before_mail_lease_expiry_does_not_strand_an_account ... FAILED failures: ---- restart_before_mail_lease_expiry_does_not_strand_an_account stdout ---- thread 'restart_before_mail_lease_expiry_does_not_strand_an_account' (923768) panicked at crates/calternal-db/tests/mailstress_restart.rs:96:5: Mail sync lease expired after startup but was never recovered note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace failures: restart_before_mail_lease_expiry_does_not_strand_an_account test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 2.07s error: test failed, to rerun pass `-p calternal-db --test mailstress_restart` ``` The proposed periodic-recovery patch is suitable: the existing recovery query uses the expired-lease index and handles cancellation and attempt limits. I use Tokio monotonic time and check at the loop head, once per max(lease/3, poll interval), so continuous dispatch cannot starve upkeep. The shared Worker already heartbeats every job kind at lease/3, including plugin handlers. The server config uses the same `server` owner for every attempt. I will use a unique claim token in the existing `leased_by` field to fence old progress, renewal and settlement when that configured name is reused. No schema or public route change is needed.
Author
Owner

Lease audit (#1042): all 15 production job kinds enter the one Worker::run_handler heartbeat path in calternal-db. Server handlers in wire.rs: system.user.deletion, system.search.rebuild, system.search.user-rebuild, system.backup, system.cas.scrub, system.user.archive-purge, files.uploads.cleanup, files.thumbnail. Plugin handlers: mail.sync, notes.reconcile, notes.log-batch-index, notes.daily-log-projection-rebuild, photos.clip.index, video.transcode, notifications.dispatch. No direct production queue claimant exists outside Worker (repository-wide lease_next search).

Examples longer than the 120-second server lease: thumbnail full-job cap is 180 seconds (DESIGN §39); video transcode permits six hours; Mail traverses account folders; Search, Notes projections, User Home deletion/transfer, integrity checks and backup scale with the data set. All receive shared renewal every 40 seconds. CLIP inference, media snapshots, Home archive purge, integrity checks and deletion use blocking tasks or dedicated workers; Search rebuild waits on its indexer actor. These waits leave the shared renewal future pollable. No per-plugin heartbeat additions are needed.

Invariants and limits: renewal and settlement require an unexpired per-claim owner token. Lease loss drops the handler future and late settlement is ignored. This fences Jobs queue state, not already committed filesystem/provider writes or detached blocking work. Existing handlers must remain idempotent, as required by DESIGN §3. A process-wide scheduler stall or prolonged writer hold can still lose a lease; fencing prevents stale queue mutation.

git fetch origin and git merge origin/dev completed once before final gates: Already up to date. Both feature slices are committed (694e198ec, fc0425400). The web production build passed; final database and server gates are running with four build jobs, no incremental build, line-tables-only debug and OPENSSL_NO_VENDOR=1.

Lease audit (#1042): all 15 production job kinds enter the one `Worker::run_handler` heartbeat path in `calternal-db`. Server handlers in `wire.rs`: `system.user.deletion`, `system.search.rebuild`, `system.search.user-rebuild`, `system.backup`, `system.cas.scrub`, `system.user.archive-purge`, `files.uploads.cleanup`, `files.thumbnail`. Plugin handlers: `mail.sync`, `notes.reconcile`, `notes.log-batch-index`, `notes.daily-log-projection-rebuild`, `photos.clip.index`, `video.transcode`, `notifications.dispatch`. No direct production queue claimant exists outside Worker (repository-wide `lease_next` search). Examples longer than the 120-second server lease: thumbnail full-job cap is 180 seconds (DESIGN §39); video transcode permits six hours; Mail traverses account folders; Search, Notes projections, User Home deletion/transfer, integrity checks and backup scale with the data set. All receive shared renewal every 40 seconds. CLIP inference, media snapshots, Home archive purge, integrity checks and deletion use blocking tasks or dedicated workers; Search rebuild waits on its indexer actor. These waits leave the shared renewal future pollable. No per-plugin heartbeat additions are needed. Invariants and limits: renewal and settlement require an unexpired per-claim owner token. Lease loss drops the handler future and late settlement is ignored. This fences Jobs queue state, not already committed filesystem/provider writes or detached blocking work. Existing handlers must remain idempotent, as required by DESIGN §3. A process-wide scheduler stall or prolonged writer hold can still lose a lease; fencing prevents stale queue mutation. `git fetch origin` and `git merge origin/dev` completed once before final gates: `Already up to date.` Both feature slices are committed (`694e198ec`, `fc0425400`). The web production build passed; final database and server gates are running with four build jobs, no incremental build, line-tables-only debug and OPENSSL_NO_VENDOR=1.
Author
Owner

READY FOR MERGE: yes

Built #1042 on job/lease-1042. Head: fc0425400227259bf74f33d4887352a7d6cceead.

  • Recover expired leases during service, including leases that expire after replacement startup. The existing indexed query runs at the loop head at max(lease/3, poll interval), which is 40 seconds with the production configuration. Continuous ready work cannot starve recovery.
  • Give each Worker claim a fresh UUID owner token in the existing leased_by field. Renewal, progress and settlement cannot adopt a replacement attempt when the configured server name is reused. Late success or failure is ignored by Worker.
  • Added the supplied restart regression, a slow-handler renewal regression, a reused-worker-name stale completion regression, and a direct Worker settlement regression. Existing test expectations remain unchanged.
  • Audited all 15 production kinds. All run through shared renewal; no plugin call-site changes are needed. Detailed audit is in the preceding comment.

Files:

  • crates/calternal-db/src/worker.rs
  • crates/calternal-db/src/jobs.rs
  • crates/calternal-db/tests/mailstress_restart.rs
  • crates/calternal-db/tests/queue.rs

Atomic commits: 694e198ec (recovery and restart regression), fc0425400 (attempt fencing and renewal/stale settlement regressions). Fetched origin and merged origin/dev once before final gates: Already up to date.

Validation: CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 OPENSSL_NO_VENDOR=1, with the supplied target directory and worktree target/tmp. Built apps/web before server gates. cargo fmt --check exited 0 with no output. Verbatim gate excerpts follow; complete logs are under artifacts/lease-1042/.

cargo clippy -p calternal-db --all-targets -- -D warnings

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.25s

cargo test -p calternal-db -- --test-threads=4

    Finished `test` profile [unoptimized + debuginfo] target(s) in 0.24s
test result: ok. 22 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 5.83s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.11s
test result: ok. 18 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 0.79s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.03s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

cargo clippy -p calternal-server --all-targets -- -D warnings

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 3m 31s

cargo test -p calternal-server -- --test-threads=4

    Finished `test` profile [unoptimized + debuginfo] target(s) in 6m 38s
test result: ok. 163 passed; 0 failed; 6 ignored; 0 measured; 0 filtered out; finished in 48.05s

Web production build (bun run --cwd apps/web build) exited 0. Verbatim build receipts:

✓ built in 10.44s
✓ built in 14ms
✓ built in 25.05s

Known gaps:

  • Post-deploy production expired-lease aggregate and logs remain for the orchestrator; this job does not deploy. No known blocking implementation gap remains.
  • Fencing protects Jobs queue state. It cannot undo external writes or stop detached blocking work after lease loss; handlers still require the existing idempotency invariant. Scheduler stalls or long writer holds can lose leases.
  • No new dependencies, migrations, UI changes or performance measurements. The latest verification policy reserves full adversarial matrices, live e2e and performance measurements for the merge round.

Decisions not specified by DESIGN:

  • Periodic indexed recovery, rather than expanding the claim query: keeps cancellation and attempt-limit handling in the existing recovery operation, with bounded frequency and no busy loop.
  • Reuse leased_by as a per-claim identity (worker-name:UUID), rather than adding a migration or changing every queue operation's signature. Direct queue callers must also use fresh attempt identities, as documented.

UX gaps closed / UX gaps left: not applicable; no UI changed.

For the merge round:

  • bash tests/adversarial/run.sh: run one time-boxed real-server round for queue controls, concurrency and cross-plugin consistency.
  • bun run --cwd apps/web test:e2e:mail-sync-613: confirm Mail sync works through the production Worker.
  • After deployment and restart, check the following read-only aggregate after interrupted leases have expired and one 40-second recovery period has passed. It must not show persistently expired leases. Review the production log's aggregate count of job worker failed; restarting as well. Do not publish credentials, message bodies or job payloads.
sqlite3 -readonly "$INDEX_DB" "SELECT kind, count(*) AS expired_leases FROM jobs WHERE state='leased' AND lease_until <= CAST(strftime('%s','now') AS INTEGER)*1000 GROUP BY kind;"

Cleanup: cargo clean run; web build/ and .svelte-kit/output/ removed. Working tree is clean. No push or deploy. Issue remains open.

READY FOR MERGE: yes Built #1042 on `job/lease-1042`. Head: `fc0425400227259bf74f33d4887352a7d6cceead`. - Recover expired leases during service, including leases that expire after replacement startup. The existing indexed query runs at the loop head at max(lease/3, poll interval), which is 40 seconds with the production configuration. Continuous ready work cannot starve recovery. - Give each Worker claim a fresh UUID owner token in the existing `leased_by` field. Renewal, progress and settlement cannot adopt a replacement attempt when the configured server name is reused. Late success or failure is ignored by Worker. - Added the supplied restart regression, a slow-handler renewal regression, a reused-worker-name stale completion regression, and a direct Worker settlement regression. Existing test expectations remain unchanged. - Audited all 15 production kinds. All run through shared renewal; no plugin call-site changes are needed. Detailed audit is in the preceding comment. Files: - `crates/calternal-db/src/worker.rs` - `crates/calternal-db/src/jobs.rs` - `crates/calternal-db/tests/mailstress_restart.rs` - `crates/calternal-db/tests/queue.rs` Atomic commits: `694e198ec` (recovery and restart regression), `fc0425400` (attempt fencing and renewal/stale settlement regressions). Fetched origin and merged `origin/dev` once before final gates: `Already up to date.` Validation: `CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 OPENSSL_NO_VENDOR=1`, with the supplied target directory and worktree `target/tmp`. Built `apps/web` before server gates. `cargo fmt --check` exited 0 with no output. Verbatim gate excerpts follow; complete logs are under `artifacts/lease-1042/`. `cargo clippy -p calternal-db --all-targets -- -D warnings` ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.25s ``` `cargo test -p calternal-db -- --test-threads=4` ```text Finished `test` profile [unoptimized + debuginfo] target(s) in 0.24s test result: ok. 22 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 5.83s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.11s test result: ok. 18 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 0.79s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.03s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` `cargo clippy -p calternal-server --all-targets -- -D warnings` ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 3m 31s ``` `cargo test -p calternal-server -- --test-threads=4` ```text Finished `test` profile [unoptimized + debuginfo] target(s) in 6m 38s test result: ok. 163 passed; 0 failed; 6 ignored; 0 measured; 0 filtered out; finished in 48.05s ``` Web production build (`bun run --cwd apps/web build`) exited 0. Verbatim build receipts: ```text ✓ built in 10.44s ✓ built in 14ms ✓ built in 25.05s ``` Known gaps: - Post-deploy production expired-lease aggregate and logs remain for the orchestrator; this job does not deploy. No known blocking implementation gap remains. - Fencing protects Jobs queue state. It cannot undo external writes or stop detached blocking work after lease loss; handlers still require the existing idempotency invariant. Scheduler stalls or long writer holds can lose leases. - No new dependencies, migrations, UI changes or performance measurements. The latest verification policy reserves full adversarial matrices, live e2e and performance measurements for the merge round. Decisions not specified by DESIGN: - Periodic indexed recovery, rather than expanding the claim query: keeps cancellation and attempt-limit handling in the existing recovery operation, with bounded frequency and no busy loop. - Reuse `leased_by` as a per-claim identity (`worker-name:UUID`), rather than adding a migration or changing every queue operation's signature. Direct queue callers must also use fresh attempt identities, as documented. UX gaps closed / UX gaps left: not applicable; no UI changed. For the merge round: - `bash tests/adversarial/run.sh`: run one time-boxed real-server round for queue controls, concurrency and cross-plugin consistency. - `bun run --cwd apps/web test:e2e:mail-sync-613`: confirm Mail sync works through the production Worker. - After deployment and restart, check the following read-only aggregate after interrupted leases have expired and one 40-second recovery period has passed. It must not show persistently expired leases. Review the production log's aggregate count of `job worker failed; restarting` as well. Do not publish credentials, message bodies or job payloads. ```sh sqlite3 -readonly "$INDEX_DB" "SELECT kind, count(*) AS expired_leases FROM jobs WHERE state='leased' AND lease_until <= CAST(strftime('%s','now') AS INTEGER)*1000 GROUP BY kind;" ``` Cleanup: `cargo clean` run; web `build/` and `.svelte-kit/output/` removed. Working tree is clean. No push or deploy. Issue remains open.
Author
Owner

Send-back started on job/lease-1042, base/head fc0425400. Reproduce the single-writer heartbeat delay, protect live attempts during recovery, and audit long write transactions. No push or deploy. Gates: calternal-db and calternal-server, web build first, OPENSSL_NO_VENDOR=1.

Send-back started on job/lease-1042, base/head fc0425400. Reproduce the single-writer heartbeat delay, protect live attempts during recovery, and audit long write transactions. No push or deploy. Gates: calternal-db and calternal-server, web build first, OPENSSL_NO_VENDOR=1.
Author
Owner

Confirmed: a 30 ms lease with the single writer held across expiry was recovered while its handler was still alive. Deterministic fake-clock regression before fix:



thread 'live_writer_transaction_outlasts_lease_and_completes_once' (1115163) panicked at crates/calternal-db/tests/queue.rs:973:5:
assertion `left == right` failed: recovery must skip the live attempt
  left: 1
 right: 0
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
test live_writer_transaction_outlasts_lease_and_completes_once ... FAILED

failures:

failures:
    live_writer_transaction_outlasts_lease_and_completes_once

test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 19 filtered out; finished in 0.97s

error: test failed, to rerun pass `-p calternal-db --test queue`
 v2.0.119
   Compiling synstructure v0.14.0
   Compiling zerovec-derive v0.11.6
   Compiling displaydoc v0.2.7
   Compiling zerofrom-derive v0.1.8
   Compiling yoke-derive v0.8.3
   Compiling scopeguard v1.2.0
   Compiling lock_api v0.4.14
   Compiling zerofrom v0.1.8
   Compiling yoke v0.8.3
   Compiling once_cell v1.21.4
   Compiling parking_lot_core v0.9.12
   Compiling serde_core v1.0.229
   Compiling zerovec v0.11.8
   Compiling zerotrie v0.2.5
   Compiling futures-sink v0.3.34
   Compiling crypto-common v0.2.2
   Compiling equivalent v1.0.2
   Compiling tinystr v0.8.4
   Compiling potential_utf v0.1.6
   Compiling memchr v2.8.3
   Compiling crossbeam-utils v0.8.23
   Compiling icu_collections v2.3.0
   Compiling icu_locale_core v2.3.0
   Compiling crypto-common v0.1.6
   Compiling block-buffer v0.10.4
   Compiling tokio-macros v2.7.2
   Compiling icu_provider v2.3.1
   Compiling socket2 v0.6.5
   Compiling mio v1.2.3
   Compiling icu_normalizer v2.3.0
   Compiling icu_properties v2.3.0
   Compiling bytes v1.12.1
   Compiling vcpkg v0.2.15
   Compiling futures-io v0.3.34
   Compiling thiserror v2.0.21
   Compiling allocator-api2 v0.2.21
   Compiling idna_adapter v1.2.2
   Compiling percent-encoding v2.3.2
   Compiling slab v0.4.12
   Compiling rand_core v0.6.4
   Compiling pkg-config v0.3.34
   Compiling futures-task v0.3.34
   Compiling serde v1.0.229
   Compiling foldhash v0.2.0
   Compiling libsqlite3-sys v0.37.0
   Compiling futures-util v0.3.34
   Compiling hashbrown v0.16.1
   Compiling rand v0.8.8
   Compiling form_urlencoded v1.2.2
   Compiling idna v1.1.0
   Compiling tokio v1.53.1
   Compiling digest v0.10.7
   Compiling parking_lot v0.12.5
   Compiling tracing-core v0.1.36
   Compiling tracing-attributes v0.1.31
   Compiling thiserror-impl v2.0.21
   Compiling serde_derive v1.0.229
   Compiling phf_shared v0.11.3
   Compiling inout v0.2.2
   Compiling cmov v0.5.4
   Compiling cpufeatures v0.2.17
   Compiling parking v2.2.1
   Compiling hashbrown v0.17.1
   Compiling crc-catalog v2.5.0
   Compiling cpufeatures v0.3.1
   Compiling log v0.4.34
   Compiling indexmap v2.14.2
   Compiling event-listener v5.4.2
   Compiling crc v3.4.0
   Compiling tracing v0.1.44
   Compiling sha2 v0.10.9
   Compiling ctutils v0.4.2
   Compiling phf_generator v0.11.3
   Compiling tokio-stream v0.1.19
   Compiling crossbeam-queue v0.3.14
   Compiling futures-intrusive v0.5.0
   Compiling url v2.5.8
   Compiling hashlink v0.11.1
   Compiling spin v0.9.9
   Compiling block-buffer v0.12.1
   Compiling iana-time-zone v0.1.65
   Compiling either v1.18.0
   Compiling base64 v0.22.1
   Compiling zmij v1.0.23
   Compiling chrono v0.4.45
   Compiling cipher v0.5.2
   Compiling sqlx-core v0.9.0
   Compiling flume v0.12.0
   Compiling phf_macros v0.11.3
   Compiling universal-hash v0.6.1
   Compiling futures-executor v0.3.34
   Compiling atoi v2.0.0
   Compiling futures-channel v0.3.34
   Compiling blake3 v1.8.7
   Compiling phf_shared v0.12.1
   Compiling serde_json v1.0.151
   Compiling rustix v1.1.5
   Compiling chrono-tz v0.10.4
   Compiling phf v0.11.3
   Compiling phf v0.12.1
   Compiling sqlx-sqlite v0.9.0
   Compiling poly1305 v0.9.1
   Compiling chacha20 v0.10.2
   Compiling aead v0.6.1
   Compiling arrayvec v0.7.8
   Compiling linux-raw-sys v0.12.1
   Compiling winnow v0.7.15
   Compiling constant_time_eq v0.4.2
   Compiling bitflags v2.13.2
   Compiling itoa v1.0.18
   Compiling cron v0.17.0
   Compiling chacha20poly1305 v0.11.0
   Compiling sqlx v0.9.0
   Compiling uuid v1.26.1
   Compiling fastrand v2.5.0
   Compiling tempfile v3.27.0
   Compiling calternal-db v0.1.0 (/home/kayg/Developer/calternal-wt/lease-1042/crates/calternal-db)
    Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 10s
     Running tests/queue.rs (/home/kayg/build/targets/lease-1042/debug/deps/queue-bc7da7cdac84b88b)


Fix combines (a) and (b): recovery skips live attempt tokens shared by every JobQueue over the same Db; renewal, checkpoints, progress and settlement require token ownership, without an expiry predicate. A drop guard removes tokens on task completion, panic or abort. It is installed before the claim waits for the writer and held through settlement. Recovery takes its token snapshot after acquiring the writer, so it cannot miss an already claimed live attempt. The focused writer-hold regression passes, with one attempt and one committed row. No handler chunking changed; the transaction audit is in progress.

Confirmed: a 30 ms lease with the single writer held across expiry was recovered while its handler was still alive. Deterministic fake-clock regression before fix: ```text thread 'live_writer_transaction_outlasts_lease_and_completes_once' (1115163) panicked at crates/calternal-db/tests/queue.rs:973:5: assertion `left == right` failed: recovery must skip the live attempt left: 1 right: 0 note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace test live_writer_transaction_outlasts_lease_and_completes_once ... FAILED failures: failures: live_writer_transaction_outlasts_lease_and_completes_once test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 19 filtered out; finished in 0.97s error: test failed, to rerun pass `-p calternal-db --test queue` v2.0.119 Compiling synstructure v0.14.0 Compiling zerovec-derive v0.11.6 Compiling displaydoc v0.2.7 Compiling zerofrom-derive v0.1.8 Compiling yoke-derive v0.8.3 Compiling scopeguard v1.2.0 Compiling lock_api v0.4.14 Compiling zerofrom v0.1.8 Compiling yoke v0.8.3 Compiling once_cell v1.21.4 Compiling parking_lot_core v0.9.12 Compiling serde_core v1.0.229 Compiling zerovec v0.11.8 Compiling zerotrie v0.2.5 Compiling futures-sink v0.3.34 Compiling crypto-common v0.2.2 Compiling equivalent v1.0.2 Compiling tinystr v0.8.4 Compiling potential_utf v0.1.6 Compiling memchr v2.8.3 Compiling crossbeam-utils v0.8.23 Compiling icu_collections v2.3.0 Compiling icu_locale_core v2.3.0 Compiling crypto-common v0.1.6 Compiling block-buffer v0.10.4 Compiling tokio-macros v2.7.2 Compiling icu_provider v2.3.1 Compiling socket2 v0.6.5 Compiling mio v1.2.3 Compiling icu_normalizer v2.3.0 Compiling icu_properties v2.3.0 Compiling bytes v1.12.1 Compiling vcpkg v0.2.15 Compiling futures-io v0.3.34 Compiling thiserror v2.0.21 Compiling allocator-api2 v0.2.21 Compiling idna_adapter v1.2.2 Compiling percent-encoding v2.3.2 Compiling slab v0.4.12 Compiling rand_core v0.6.4 Compiling pkg-config v0.3.34 Compiling futures-task v0.3.34 Compiling serde v1.0.229 Compiling foldhash v0.2.0 Compiling libsqlite3-sys v0.37.0 Compiling futures-util v0.3.34 Compiling hashbrown v0.16.1 Compiling rand v0.8.8 Compiling form_urlencoded v1.2.2 Compiling idna v1.1.0 Compiling tokio v1.53.1 Compiling digest v0.10.7 Compiling parking_lot v0.12.5 Compiling tracing-core v0.1.36 Compiling tracing-attributes v0.1.31 Compiling thiserror-impl v2.0.21 Compiling serde_derive v1.0.229 Compiling phf_shared v0.11.3 Compiling inout v0.2.2 Compiling cmov v0.5.4 Compiling cpufeatures v0.2.17 Compiling parking v2.2.1 Compiling hashbrown v0.17.1 Compiling crc-catalog v2.5.0 Compiling cpufeatures v0.3.1 Compiling log v0.4.34 Compiling indexmap v2.14.2 Compiling event-listener v5.4.2 Compiling crc v3.4.0 Compiling tracing v0.1.44 Compiling sha2 v0.10.9 Compiling ctutils v0.4.2 Compiling phf_generator v0.11.3 Compiling tokio-stream v0.1.19 Compiling crossbeam-queue v0.3.14 Compiling futures-intrusive v0.5.0 Compiling url v2.5.8 Compiling hashlink v0.11.1 Compiling spin v0.9.9 Compiling block-buffer v0.12.1 Compiling iana-time-zone v0.1.65 Compiling either v1.18.0 Compiling base64 v0.22.1 Compiling zmij v1.0.23 Compiling chrono v0.4.45 Compiling cipher v0.5.2 Compiling sqlx-core v0.9.0 Compiling flume v0.12.0 Compiling phf_macros v0.11.3 Compiling universal-hash v0.6.1 Compiling futures-executor v0.3.34 Compiling atoi v2.0.0 Compiling futures-channel v0.3.34 Compiling blake3 v1.8.7 Compiling phf_shared v0.12.1 Compiling serde_json v1.0.151 Compiling rustix v1.1.5 Compiling chrono-tz v0.10.4 Compiling phf v0.11.3 Compiling phf v0.12.1 Compiling sqlx-sqlite v0.9.0 Compiling poly1305 v0.9.1 Compiling chacha20 v0.10.2 Compiling aead v0.6.1 Compiling arrayvec v0.7.8 Compiling linux-raw-sys v0.12.1 Compiling winnow v0.7.15 Compiling constant_time_eq v0.4.2 Compiling bitflags v2.13.2 Compiling itoa v1.0.18 Compiling cron v0.17.0 Compiling chacha20poly1305 v0.11.0 Compiling sqlx v0.9.0 Compiling uuid v1.26.1 Compiling fastrand v2.5.0 Compiling tempfile v3.27.0 Compiling calternal-db v0.1.0 (/home/kayg/Developer/calternal-wt/lease-1042/crates/calternal-db) Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 10s Running tests/queue.rs (/home/kayg/build/targets/lease-1042/debug/deps/queue-bc7da7cdac84b88b) ``` Fix combines (a) and (b): recovery skips live attempt tokens shared by every JobQueue over the same Db; renewal, checkpoints, progress and settlement require token ownership, without an expiry predicate. A drop guard removes tokens on task completion, panic or abort. It is installed before the claim waits for the writer and held through settlement. Recovery takes its token snapshot after acquiring the writer, so it cannot miss an already claimed live attempt. The focused writer-hold regression passes, with one attempt and one committed row. No handler chunking changed; the transaction audit is in progress.
Author
Owner

Job writer leases (#1042)

This audit supports DESIGN §§2 and 3. It covers the 15 production Job kinds.
The server lease is 120 seconds. Renewal is due every 40 seconds.
The owner reported about 7,800 Log entries and 14,800 Jobs in the Index.

Evidence and limits

The production report on #1042 shows notes.log-batch-index still leased
2.7 minutes after its last update, with expiry 0.7 minutes earlier.
This proves a renewal delay. It does not measure a particular transaction.
This audit checks source code and transaction boundaries. It does not claim
measured production transaction durations. Job count alone cannot give those
durations. Entry distribution per Daily note and Index size also matter.

The deterministic regression uses a 30 ms lease. The handler holds the only
writer across expiry. A separate queue wrapper requests recovery while the
writer is held. Before the fix, recovery resets the live attempt. After the
fix, the same attempt completes with one call and one committed row.

Handlers that can delay renewal

These paths have no 40-second limit on writer occupancy. Inspect these first
if production traces show long writer holds. Every live handler is protected
by the shared token set, including handlers with otherwise small writes.

Job kind Writer boundary and work that can exceed the renewal interval
notes.log-batch-index notes/src/lib.rs: at most 3,000 paths, with a commit per Note. store::index_note_projection holds a transaction through Journal, Calendar, Search, backlinks and reminder updates. tasks_store::index_source then has a separate per-source transaction for Tasks, Tags and attachments. Large individual Notes or many affected backlinks can hold the writer past 40 seconds. The batch does not hold one transaction across all paths.
notes.daily-log-projection-rebuild Uses the same projection functions. Commits and saves a cursor per Daily note, then waits 10 ms between Notes. Large individual Notes have the same risk; the whole Home is not one transaction.
notes.reconcile store::reconcile_user calls index per Note and removes stale projections per path. It has the same per-Note and backlink risk.
system.user.deletion Calls Notes reconciliation, so it inherits the projection risk. Files cleanup already commits batches of 500 rows. Home transfer and archive work do not hold the Index writer. Final security-state cleanup can also cascade through a User's Index rows.
system.backup calternal-db/src/snapshot.rs: VACUUM INTO holds the only writer for the complete Index copy. Its duration grows with Index size and disk latency.
system.search.rebuild calternal-search/src/indexer.rs: publish_staged_manifest copies and deletes complete manifest tables in one transaction. Rollback copies the previous manifest in one transaction. File scans and Tantivy work are outside these Index transactions.
mail.sync cache/store.rs: store_window commits windows of 80 UIDs. Folder-list replacement loops over the complete folder list in one transaction. Generation activation deletes old memberships and unused account messages inside the window transaction. This cleanup is not bounded by the 80-UID window. Provider fetches occur outside the transaction.

Other handlers

These paths have small or bounded writer units, or do their long work outside
the shared Index writer. This does not guarantee a duration on a loaded host.

Job kind Writer use
system.search.user-rebuild Private Search generation is built outside the shared Index writer. Shared invalidation cleanup is one statement.
system.cas.scrub Hashing and repair run on a dedicated thread. Progress uses separate queue writes; no Index transaction encloses the scan.
system.user.archive-purge Filesystem work on a blocking task; no Index transaction encloses the purge.
files.uploads.cleanup Deletes one upload row, releases the writer, then removes staging files.
files.thumbnail Media processing runs outside the Index writer.
photos.clip.index Batches of eight identities; inference uses a separate derived-data Index. Shared queue writes are separate.
video.transcode Media work runs outside the writer. State updates are separate statements.
notifications.dispatch Reminder and push loops commit per notification or delivery. Network delivery is outside the notification transaction. Inbox retention is bounded.

Decision

Use options (a) and (b) from the send-back. Expiry permits recovery of an
orphan; it does not revoke a live attempt. Recovery skips tokens registered
in the shared Db. Each Worker registers its unique token before claim and
keeps a drop guard through settlement. Recovery obtains the writer before
it reads the token set. Task completion, panic and cancellation remove the
token. A new process starts with an empty set.

Renewal, settlement, progress and stop checkpoints use the owner token even
after expiry. Recovery clears that token before it permits a new claim.
An old token cannot change its successor's queue state.

Do not split the per-Note projection transaction in this fix. Journal
snapshots and their resource revisions must publish together (#549).
Several handlers already commit per Note, per window or per batch. Chunking
other paths can improve latency, but lease safety must not depend on a
fixed transaction duration. The single server writer is required. This is
not a coordination protocol for multiple server processes.

Checks for the merge round

Deploy the combined branch through the normal merge-round process. Check
read-only aggregate Job state after restart and after one lease interval.
An expired timestamp on a live write attempt can occur; it must not cause
another claim. Check that orphaned attempts become pending, completed or
dead, and that affected projection Jobs complete without repeated attempts.
Use production write-stage traces to measure transaction durations before
any later change to chunk sizes. Do not publish Home content or credentials.

# Job writer leases (#1042) This audit supports DESIGN §§2 and 3. It covers the 15 production Job kinds. The server lease is 120 seconds. Renewal is due every 40 seconds. The owner reported about 7,800 Log entries and 14,800 Jobs in the Index. ## Evidence and limits The production report on #1042 shows `notes.log-batch-index` still leased 2.7 minutes after its last update, with expiry 0.7 minutes earlier. This proves a renewal delay. It does not measure a particular transaction. This audit checks source code and transaction boundaries. It does not claim measured production transaction durations. Job count alone cannot give those durations. Entry distribution per Daily note and Index size also matter. The deterministic regression uses a 30 ms lease. The handler holds the only writer across expiry. A separate queue wrapper requests recovery while the writer is held. Before the fix, recovery resets the live attempt. After the fix, the same attempt completes with one call and one committed row. ## Handlers that can delay renewal These paths have no 40-second limit on writer occupancy. Inspect these first if production traces show long writer holds. Every live handler is protected by the shared token set, including handlers with otherwise small writes. | Job kind | Writer boundary and work that can exceed the renewal interval | | --- | --- | | `notes.log-batch-index` | `notes/src/lib.rs`: at most 3,000 paths, with a commit per Note. `store::index_note_projection` holds a transaction through Journal, Calendar, Search, backlinks and reminder updates. `tasks_store::index_source` then has a separate per-source transaction for Tasks, Tags and attachments. Large individual Notes or many affected backlinks can hold the writer past 40 seconds. The batch does not hold one transaction across all paths. | | `notes.daily-log-projection-rebuild` | Uses the same projection functions. Commits and saves a cursor per Daily note, then waits 10 ms between Notes. Large individual Notes have the same risk; the whole Home is not one transaction. | | `notes.reconcile` | `store::reconcile_user` calls `index` per Note and removes stale projections per path. It has the same per-Note and backlink risk. | | `system.user.deletion` | Calls Notes reconciliation, so it inherits the projection risk. Files cleanup already commits batches of 500 rows. Home transfer and archive work do not hold the Index writer. Final security-state cleanup can also cascade through a User's Index rows. | | `system.backup` | `calternal-db/src/snapshot.rs`: `VACUUM INTO` holds the only writer for the complete Index copy. Its duration grows with Index size and disk latency. | | `system.search.rebuild` | `calternal-search/src/indexer.rs`: `publish_staged_manifest` copies and deletes complete manifest tables in one transaction. Rollback copies the previous manifest in one transaction. File scans and Tantivy work are outside these Index transactions. | | `mail.sync` | `cache/store.rs`: `store_window` commits windows of 80 UIDs. Folder-list replacement loops over the complete folder list in one transaction. Generation activation deletes old memberships and unused account messages inside the window transaction. This cleanup is not bounded by the 80-UID window. Provider fetches occur outside the transaction. | ## Other handlers These paths have small or bounded writer units, or do their long work outside the shared Index writer. This does not guarantee a duration on a loaded host. | Job kind | Writer use | | --- | --- | | `system.search.user-rebuild` | Private Search generation is built outside the shared Index writer. Shared invalidation cleanup is one statement. | | `system.cas.scrub` | Hashing and repair run on a dedicated thread. Progress uses separate queue writes; no Index transaction encloses the scan. | | `system.user.archive-purge` | Filesystem work on a blocking task; no Index transaction encloses the purge. | | `files.uploads.cleanup` | Deletes one upload row, releases the writer, then removes staging files. | | `files.thumbnail` | Media processing runs outside the Index writer. | | `photos.clip.index` | Batches of eight identities; inference uses a separate derived-data Index. Shared queue writes are separate. | | `video.transcode` | Media work runs outside the writer. State updates are separate statements. | | `notifications.dispatch` | Reminder and push loops commit per notification or delivery. Network delivery is outside the notification transaction. Inbox retention is bounded. | ## Decision Use options (a) and (b) from the send-back. Expiry permits recovery of an orphan; it does not revoke a live attempt. Recovery skips tokens registered in the shared `Db`. Each Worker registers its unique token before claim and keeps a drop guard through settlement. Recovery obtains the writer before it reads the token set. Task completion, panic and cancellation remove the token. A new process starts with an empty set. Renewal, settlement, progress and stop checkpoints use the owner token even after expiry. Recovery clears that token before it permits a new claim. An old token cannot change its successor's queue state. Do not split the per-Note projection transaction in this fix. Journal snapshots and their resource revisions must publish together (#549). Several handlers already commit per Note, per window or per batch. Chunking other paths can improve latency, but lease safety must not depend on a fixed transaction duration. The single server writer is required. This is not a coordination protocol for multiple server processes. ## Checks for the merge round Deploy the combined branch through the normal merge-round process. Check read-only aggregate Job state after restart and after one lease interval. An expired timestamp on a live write attempt can occur; it must not cause another claim. Check that orphaned attempts become pending, completed or dead, and that affected projection Jobs complete without repeated attempts. Use production write-stage traces to measure transaction durations before any later change to chunk sizes. Do not publish Home content or credentials.
Author
Owner

Additional writer-hold finding: the production writer pool has a 30-second checkout timeout. A healthy handler holding its transaction for minutes would also be dropped when heartbeat or periodic recovery returns PoolTimedOut, even with live-token exclusions. Worker-only queue operations now retry this pre-SQL timeout; handler polling continues. Other Index errors still stop the run, and HTTP checkout deadlines are unchanged. A focused regression with a 30 ms lease, a 100 ms pool checkout timeout and a 350 ms transaction hold passes with one attempt. The full calternal-db tests and clippy pass; server gates run next.

Additional writer-hold finding: the production writer pool has a 30-second checkout timeout. A healthy handler holding its transaction for minutes would also be dropped when heartbeat or periodic recovery returns PoolTimedOut, even with live-token exclusions. Worker-only queue operations now retry this pre-SQL timeout; handler polling continues. Other Index errors still stop the run, and HTTP checkout deadlines are unchanged. A focused regression with a 30 ms lease, a 100 ms pool checkout timeout and a 350 ms transaction hold passes with one attempt. The full calternal-db tests and clippy pass; server gates run next.
Author
Owner

READY FOR MERGE: yes

Built the #1042 send-back on job/lease-1042. Head: 5822d8c24c3193483938c435fb823dbf5d78eca5.

  • Recovery now skips live attempt tokens shared by every JobQueue over the same Db. A drop guard registers before claim and stays through settlement; finish, panic and abort remove it.
  • Renewal, completion, failure, progress and cancellation checkpoints accept the owned token after expiry. Recovery or settlement must remove ownership before another attempt can own the row. Existing stale-token rejection tests still pass.
  • Worker heartbeat, recovery, claim and settlement retry writer checkout timeouts. The 30-second production pool timeout can no longer abort a healthy minutes-long write. Other Index errors still propagate. HTTP deadlines are unchanged.
  • Added a deterministic single-writer regression: 30 ms lease, fake-clock expiry, recovery queued during the transaction, one call, one claim and one committed row. It failed before the fix (recovery reset 1 live attempt) and passes now. A second regression holds the writer for 350 ms with a 100 ms checkout timeout and finishes once. Added token-ownership coverage for delayed progress, checkpoints, renewal, success, failure and cancellation.

Files: crates/calternal-db/src/db.rs, crates/calternal-db/src/jobs.rs, crates/calternal-db/src/worker.rs, crates/calternal-db/tests/queue.rs, docs/audits/job-writer-leases-1042.md.

Decisions: combine options (a) and (b). Expiry is an orphan-recovery clock; it is not proof that a live attempt is dead. Keep existing per-Note atomic projections (#549) and existing per-window/per-batch boundaries. No Plugin behavior changed. Also retry pre-SQL checkout timeouts inside Worker, because expiry protection alone leaves the 30-second timeout failure path. No dependencies changed.

Handler audit: all 15 production kinds were inspected. Paths without a 40-second writer limit are notes.log-batch-index, notes.daily-log-projection-rebuild, notes.reconcile, system.user.deletion (reconciliation and final cleanup), system.backup (whole-Index VACUUM), system.search.rebuild (whole-manifest publication/rollback), and mail.sync (folder replacement and generation cleanup). Notes commit per Note; Mail commits windows of 80 UIDs. Source evidence and the remaining eight kinds are listed in the audit document and preceding issue comment. The production observation proves delayed renewal, not the duration of a particular transaction. Exact production transaction times are not measured in this job.

Atomic commits: 68c0cd02b (live ownership), 131dc3a7b (handler audit), bebdc7044 (checkout timeout), 5822d8c24 (comment consistency). Fetched origin and merged origin/dev once before final gates: Already up to date. Re-read changed-file doc comments. The working tree is clean.

Validation environment: CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 OPENSSL_NO_VENDOR=1 TMPDIR=<worktree>/target/tmp, supplied target directory retained. Built apps/web first. bun run --cwd apps/web build exited 0; verbatim receipt:

  Wrote site to "build"

cargo fmt --check exited 0 with no output. Verbatim gate excerpts:

cargo clippy -p calternal-db --all-targets -- -D warnings

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 5.47s

cargo test -p calternal-db -- --test-threads=4

    Finished `test` profile [unoptimized + debuginfo] target(s) in 6m 26s
test result: ok. 23 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 6.66s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.33s
test result: ok. 20 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 1.66s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.07s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

cargo clippy -p calternal-server --all-targets -- -D warnings

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 2m 01s

cargo test -p calternal-server -- --test-threads=4

    Finished `test` profile [unoptimized + debuginfo] target(s) in 10m 30s
test result: ok. 163 passed; 0 failed; 6 ignored; 0 measured; 0 filtered out; finished in 27.05s

Complete receipts and reproduction logs remain in artifacts/lease-1042/. Cargo and web build output are removed at job end.

Known gaps: no deploy or production mutation was performed. Read-only production aggregates after deployment and actual transaction timing remain for the merge round. The live set follows the single-server writer rule; it is not multi-process coordination. Detached blocking work and already committed external writes still require idempotent handlers. Existing ignored tests remain ignored. No UI changed; UX gaps closed/left: not applicable. No performance run under the current verification policy.

For the merge round: run bash tests/adversarial/run.sh once on the combined branch for the full real-server robustness, authorization and restart checks. Run bun run --cwd apps/web test and bun run --cwd apps/web test:e2e for combined web behavior. After the authorized staging deployment, check read-only Job aggregates immediately after restart and after 120 seconds: orphaned attempts must become claimable; live projection attempts must complete without recovery/reclaim loops. Inspect write-stage traces to measure the unbounded transaction paths. Do not infer an orphan solely from an expired timestamp.

READY FOR MERGE: yes Built the #1042 send-back on `job/lease-1042`. Head: `5822d8c24c3193483938c435fb823dbf5d78eca5`. - Recovery now skips live attempt tokens shared by every JobQueue over the same Db. A drop guard registers before claim and stays through settlement; finish, panic and abort remove it. - Renewal, completion, failure, progress and cancellation checkpoints accept the owned token after expiry. Recovery or settlement must remove ownership before another attempt can own the row. Existing stale-token rejection tests still pass. - Worker heartbeat, recovery, claim and settlement retry writer checkout timeouts. The 30-second production pool timeout can no longer abort a healthy minutes-long write. Other Index errors still propagate. HTTP deadlines are unchanged. - Added a deterministic single-writer regression: 30 ms lease, fake-clock expiry, recovery queued during the transaction, one call, one claim and one committed row. It failed before the fix (recovery reset 1 live attempt) and passes now. A second regression holds the writer for 350 ms with a 100 ms checkout timeout and finishes once. Added token-ownership coverage for delayed progress, checkpoints, renewal, success, failure and cancellation. Files: `crates/calternal-db/src/db.rs`, `crates/calternal-db/src/jobs.rs`, `crates/calternal-db/src/worker.rs`, `crates/calternal-db/tests/queue.rs`, `docs/audits/job-writer-leases-1042.md`. Decisions: combine options (a) and (b). Expiry is an orphan-recovery clock; it is not proof that a live attempt is dead. Keep existing per-Note atomic projections (#549) and existing per-window/per-batch boundaries. No Plugin behavior changed. Also retry pre-SQL checkout timeouts inside Worker, because expiry protection alone leaves the 30-second timeout failure path. No dependencies changed. Handler audit: all 15 production kinds were inspected. Paths without a 40-second writer limit are `notes.log-batch-index`, `notes.daily-log-projection-rebuild`, `notes.reconcile`, `system.user.deletion` (reconciliation and final cleanup), `system.backup` (whole-Index VACUUM), `system.search.rebuild` (whole-manifest publication/rollback), and `mail.sync` (folder replacement and generation cleanup). Notes commit per Note; Mail commits windows of 80 UIDs. Source evidence and the remaining eight kinds are listed in the audit document and preceding issue comment. The production observation proves delayed renewal, not the duration of a particular transaction. Exact production transaction times are not measured in this job. Atomic commits: `68c0cd02b` (live ownership), `131dc3a7b` (handler audit), `bebdc7044` (checkout timeout), `5822d8c24` (comment consistency). Fetched origin and merged `origin/dev` once before final gates: `Already up to date.` Re-read changed-file doc comments. The working tree is clean. Validation environment: `CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 OPENSSL_NO_VENDOR=1 TMPDIR=<worktree>/target/tmp`, supplied target directory retained. Built apps/web first. `bun run --cwd apps/web build` exited 0; verbatim receipt: ```text Wrote site to "build" ``` `cargo fmt --check` exited 0 with no output. Verbatim gate excerpts: `cargo clippy -p calternal-db --all-targets -- -D warnings` ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 5.47s ``` `cargo test -p calternal-db -- --test-threads=4` ```text Finished `test` profile [unoptimized + debuginfo] target(s) in 6m 26s test result: ok. 23 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 6.66s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.33s test result: ok. 20 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 1.66s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.07s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` `cargo clippy -p calternal-server --all-targets -- -D warnings` ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 2m 01s ``` `cargo test -p calternal-server -- --test-threads=4` ```text Finished `test` profile [unoptimized + debuginfo] target(s) in 10m 30s test result: ok. 163 passed; 0 failed; 6 ignored; 0 measured; 0 filtered out; finished in 27.05s ``` Complete receipts and reproduction logs remain in `artifacts/lease-1042/`. Cargo and web build output are removed at job end. Known gaps: no deploy or production mutation was performed. Read-only production aggregates after deployment and actual transaction timing remain for the merge round. The live set follows the single-server writer rule; it is not multi-process coordination. Detached blocking work and already committed external writes still require idempotent handlers. Existing ignored tests remain ignored. No UI changed; UX gaps closed/left: not applicable. No performance run under the current verification policy. For the merge round: run `bash tests/adversarial/run.sh` once on the combined branch for the full real-server robustness, authorization and restart checks. Run `bun run --cwd apps/web test` and `bun run --cwd apps/web test:e2e` for combined web behavior. After the authorized staging deployment, check read-only Job aggregates immediately after restart and after 120 seconds: orphaned attempts must become claimable; live projection attempts must complete without recovery/reclaim loops. Inspect write-stage traces to measure the unbounded transaction paths. Do not infer an orphan solely from an expired timestamp.
Author
Owner

Deployed to production (2026-10-04 ~11:33 IST, small round 5 = d0fc1f463)

Round: dev 6074f71d1 + job/apprevoke-1041 + job/passkeybind-1043 + job/lease-1042. Gates on the round: cargo fmt --check clean; clippy clean and tests green for calternal-auth (92 passed), calternal-db (45), calternal-plugin-notes (193), calternal-plugin-mail (46), calternal-server (163), 0 failed. Staging healthy first; production healthy in 9 s, no panics or errors.

## Deployed to production (2026-10-04 ~11:33 IST, small round 5 = d0fc1f463) Round: dev 6074f71d1 + job/apprevoke-1041 + job/passkeybind-1043 + job/lease-1042. Gates on the round: `cargo fmt --check` clean; clippy clean and tests green for calternal-auth (92 passed), calternal-db (45), calternal-plugin-notes (193), calternal-plugin-mail (46), calternal-server (163), 0 failed. Staging healthy first; production healthy in 9 s, no panics or errors.
Author
Owner

Post-deploy check on production (5 min after the round 5 deploy, read-only aggregate): the only leased job is mail.sync, with its lease renewed and ending 1.5 min in the future. No job is leased past expiry. Closing.

Post-deploy check on production (5 min after the round 5 deploy, read-only aggregate): the only leased job is `mail.sync`, with its lease renewed and ending 1.5 min in the future. No job is leased past expiry. Closing.
kayg closed this issue 2026-10-04 09:38:27 +00:00
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#1042
No description provided.