BLOCKER: cap PDF search extractor memory before parsing User content #781

Open
opened 2026-10-02 13:10:35 +00:00 by kayg · 1 comment
Owner

Found during the read-only sec-fs audit requested on #663. Source evidence; no decompression payload or OOM run was used.

Context and evidence

Source: c4a61e8cf0 (origin/dev). The same implementation remains in round 7a (2f4482ded0).

  • crates/calternal-search/src/pdf.rs:69–75: SubprocessPdfTextExtractor starts the server's hidden extraction command without a memory limit, a separate memory cgroup or the native media launcher.
  • crates/calternal-server/src/main.rs:257–269: the hidden command reads a capped source and calls the parser. It adds no memory limit.
  • crates/calternal-search/src/pdf.rs:129–151: the parser loads the document and builds each whole page as a String before append_capped applies the 256 KiB text limit. The 128-page cap is applied after loading the document.
  • The parent kills the child after two seconds. This limits elapsed time, not bytes allocated within that time.
  • Cargo.lock pins pdf-extract 0.12.1 and lopdf 0.42.0. The exact lopdf 0.42.0 source confirms that zlib decompression reads into a growing Vec without an output-byte cap. This version was checked against primary source on the web.

Reasoned impact

The 16 MiB source cap and final text cap do not bound parser memory, decompressed streams or a page String. An ordinary uploaded PDF automatically enters this background path (crates/calternal-search/src/indexer.rs:2703–2708). The child shares the Instance's resource domain. Memory pressure can affect the server before its deadline fires. This is a missing hard resource boundary; no OOM or maximum RSS is claimed.

Concrete fix

Set a hard child memory budget before parsing. Use a separate capped resource domain when available, or a conservative address-space limit applied in the child before loading content. Clear unnecessary inherited environment and descriptors. Bound decoded streams and page output at their sinks, and keep the deadline, input cap and final text cap. Reuse the existing decoder isolation conventions in DESIGN §39 rather than adding another unbounded parser path.

Defensive tests

Use an inert child that reports its inherited limits, plus small ordinary PDF fixtures. Assert that limits apply before parser startup, termination is reaped, the output sink stops at its cap, and indexing continues with a subsequent valid file after an extraction failure. Test the boundary with small allocations in a separate test resource domain; do not stress the shared build host.

Duplicate search: all-state PDF, extraction and cancellation searches. #58 adds PDF search; #410/#547 cover thumbnail rendering; #501 covers image/video input isolation. None covers the PDF search child's memory boundary.

Found during the read-only sec-fs audit requested on #663. Source evidence; no decompression payload or OOM run was used. ## Context and evidence Source: c4a61e8cf090170f35b1bed3350d9de20c83ecd5 (`origin/dev`). The same implementation remains in round 7a (2f4482ded066d9c5d9c59130377907f7fd2916c9). - `crates/calternal-search/src/pdf.rs:69–75`: `SubprocessPdfTextExtractor` starts the server's hidden extraction command without a memory limit, a separate memory cgroup or the native media launcher. - `crates/calternal-server/src/main.rs:257–269`: the hidden command reads a capped source and calls the parser. It adds no memory limit. - `crates/calternal-search/src/pdf.rs:129–151`: the parser loads the document and builds each whole page as a String before `append_capped` applies the 256 KiB text limit. The 128-page cap is applied after loading the document. - The parent kills the child after two seconds. This limits elapsed time, not bytes allocated within that time. - Cargo.lock pins pdf-extract 0.12.1 and lopdf 0.42.0. The exact [lopdf 0.42.0 source](https://docs.rs/lopdf/0.42.0/src/lopdf/object.rs.html#850-872) confirms that zlib decompression reads into a growing Vec without an output-byte cap. This version was checked against primary source on the web. ## Reasoned impact The 16 MiB source cap and final text cap do not bound parser memory, decompressed streams or a page String. An ordinary uploaded PDF automatically enters this background path (`crates/calternal-search/src/indexer.rs:2703–2708`). The child shares the Instance's resource domain. Memory pressure can affect the server before its deadline fires. This is a missing hard resource boundary; no OOM or maximum RSS is claimed. ## Concrete fix Set a hard child memory budget before parsing. Use a separate capped resource domain when available, or a conservative address-space limit applied in the child before loading content. Clear unnecessary inherited environment and descriptors. Bound decoded streams and page output at their sinks, and keep the deadline, input cap and final text cap. Reuse the existing decoder isolation conventions in DESIGN §39 rather than adding another unbounded parser path. ## Defensive tests Use an inert child that reports its inherited limits, plus small ordinary PDF fixtures. Assert that limits apply before parser startup, termination is reaped, the output sink stops at its cap, and indexing continues with a subsequent valid file after an extraction failure. Test the boundary with small allocations in a separate test resource domain; do not stress the shared build host. Duplicate search: all-state PDF, extraction and cancellation searches. #58 adds PDF search; #410/#547 cover thumbnail rendering; #501 covers image/video input isolation. None covers the PDF search child's memory boundary.
Author
Owner

The hard PDF child limits, pre-runtime dispatch and capped text sink are committed. All six focused PDF tests passed. Full Search/Server gates and the real server command remain unverified. Head: bd6e97e09b. Full handoff and verbatim output are on #779. No push or deploy.

The hard PDF child limits, pre-runtime dispatch and capped text sink are committed. All six focused PDF tests passed. Full Search/Server gates and the real server command remain unverified. Head: bd6e97e09b56c8df6ae770f5bb440aaf2d9d8f36. Full handoff and verbatim output are on #779. No push or deploy.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#781
No description provided.