Repository navigation
feat: read a scan's invisible OCR text layer when asked - #628
logancyang wants to merge 3 commits into
Conversation
Add an opt-in readInvisibleTextLayer option to processPdf and processPdfAsync (DetectionConfig::read_invisible_text_layer, pdf2md --read-invisible-text-layer). With it, the detector does not flag a page whose only text is an invisible layer under a covering image, and the extractor reads a page again with its invisible text when that adds at least 40 letters or digits and the visible text is at most a quarter of that gain, such as a download footer over an OCR scan. Off by default. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
a655b94 to
afd3a9c
Compare
Detection now clears a scan's invisible_text_layer flag only when extraction would adopt that layer (OCR_LAYER_MIN_ALNUM, not garbage), so a short layer keeps its page on the OCR path. The extraction retry reads the layer only on a page with a covering image, and shares the page's form-walk budget with the visible pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A page whose invisible layer is read as text still counts as a scan page, so an OCR'd book is Mixed and the per-page rule decides every page: one with no layer, or a layer extraction would not keep, goes to OCR even when the sample left it out. Detection asks extraction's own per-page path (the visible pass, then the layer retry, on one form budget) whether it keeps a page's layer, with the document's fonts read once. The reason for an unsampled page of a scanned document goes back to the base rule. Tests now write the render mode inside each text object, as OCR tools do, so they exercise the retry and its 4x guard. pdf2md --help lists --read-invisible-text-layer, and the changelog notes the gate and the new DetectionConfig field. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| if (analysis.has_template_image | ||
| && (analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_ok)) | ||
| || analysis.has_invisible_text_layer | ||
| || analysis.has_readable_invisible_text_layer |
There was a problem hiding this comment.
🟡 Readable scans falsely recommend OCR
When sampled scan pages have readable layers and sufficient text operators, has_readable_invisible_text_layer forces Mixed classification. The Mixed branch sets ocr_recommended even when pages_needing_ocr is empty, prompting unnecessary OCR.
Learn more
A readable OCR layer is accepted as text by the extraction probe. Counting that page as a template forces the Mixed branch when pages_with_text > 0, and classification sets ocr_recommended unconditionally. The per-page routing then leaves a readable scan out of pages_needing_ocr, so detection recommends OCR for pages that need none.
Example: A one-page scan has 120 individually positioned invisible glyphs and a covering raster. With the option enabled, its layer is extracted and pages_needing_ocr is empty, but ocr_recommended is true.
Recommended fix: Preserve the per-page Mixed routing for unread scans, but derive the OCR recommendation from the resulting OCR list or otherwise exclude accepted layers from the recommendation calculation.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Valid, but narrow. Every Mixed document already sets ocr_recommended, and processPdf does not return it; pages_needing_ocr is the per-page answer. No change to the field. The docs for this option will say to route on pages_needing_ocr.
| pub(crate) fn keeps_layer(&mut self, doc: &Document, page_id: ObjectId, page_num: u32) -> bool { | ||
| let (font_cmaps, style_cache) = self | ||
| .fonts | ||
| .get_or_insert_with(|| (FontCMaps::from_doc(doc), FontStyleCache::new())); |
There was a problem hiding this comment.
🟡 Sampled detection parses every page's fonts
On a sampled page with a hidden layer, keeps_layer builds font maps for the entire document. FontCMaps::from_doc traverses unsampled pages and parses their embedded fonts, making sampled detection pay full-document font costs.
Learn more
Detection normally scans only the pages chosen by ScanStrategy, but the first candidate hidden layer now initializes a document-wide font map. FontCMaps::from_doc_pages walks every page and processes its page and form fonts when no filter is supplied. Thus even ScanStrategy::Pages(vec![1]) triggers font processing for every unselected page when page 1 has a hidden layer.
Example: Page 1 of a 1,000-page PDF has an OCR layer; the other 999 pages each embed a distinct TrueType font. Detecting only page 1 now parses the other pages' fonts before answering.
Recommended fix: Scope CMap construction to pages actually probed. Cache page-scoped maps or lazily extend a font map as new pages are analyzed, preserving the sampled strategy's resource bound.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Valid, and part of a larger cost. With the option on, detection reads every hidden-layer page twice, whatever pages were asked, and extraction then reads it again. We will fix that cost in this PR.
There was a problem hiding this comment.
3 issues found across 6 files (changes from recent commits).
Confidence score: 3/5
src/detector.rscan classify an all-readable scan as Mixed even when no pages need OCR. Deriveocr_recommendedfrom the pages that still need OCR.src/extractor/mod.rsbuilds CMaps for unsampled pages, defeating the resource bound inScanStrategy::Pages. Keep CMap construction scoped to the pages detection probes.CHANGELOG.mdomits the four-to-one gate, so readers may miss that a 40+ character text layer is still rejected unless added alphanumeric text is at least four times the visible text. Document both thresholds.
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="CHANGELOG.md">
<violation number="1" location="CHANGELOG.md:18">
P2: This omits the four-to-one gate: a 40+ character layer is still rejected unless its added alphanumeric text is at least four times the visible text. Document the added-text threshold and ratio so the changelog does not promise adoption for these pages.</violation>
</file>
<file name="src/detector.rs">
<violation number="1" location="src/detector.rs:326">
P2: Derive `ocr_recommended` from the pages that still need OCR; an all-readable scan can now be classified as Mixed while `pages_needing_ocr` is empty.</violation>
</file>
<file name="src/extractor/mod.rs">
<violation number="1" location="src/extractor/mod.rs:757">
P2: Keep CMap construction scoped to the pages detection probes; this document-wide initialization parses fonts from unsampled pages and defeats `ScanStrategy::Pages`' resource bound.</violation>
</file>
Shadow auto-approve: would not auto-approve because issues were found.
View guided diff | Re-trigger cubic
| option on `processPdfAsync`, `DetectionConfig::read_invisible_text_layer` | ||
| and `pdf2md --read-invisible-text-layer`, read the invisible OCR text layer | ||
| (render mode 3) of a scanned page as its text. A layer of at least 40 | ||
| letters or digits that is not garbage is read, and its page is not |
There was a problem hiding this comment.
P2: This omits the four-to-one gate: a 40+ character layer is still rejected unless its added alphanumeric text is at least four times the visible text. Document the added-text threshold and ratio so the changelog does not promise adoption for these pages.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At CHANGELOG.md, line 18:
<comment>This omits the four-to-one gate: a 40+ character layer is still rejected unless its added alphanumeric text is at least four times the visible text. Document the added-text threshold and ratio so the changelog does not promise adoption for these pages.</comment>
<file context>
@@ -14,10 +14,14 @@ version and date. Earlier releases are described in their
- stamp over such a layer, like an archive's download footer, returns the
- layer too. Off by default; output without the option is unchanged.
+ (render mode 3) of a scanned page as its text. A layer of at least 40
+ letters or digits that is not garbage is read, and its page is not
+ flagged `invisible_text_layer`; a shorter layer keeps the page on the OCR
+ path. A page whose only visible text is a small stamp over such a layer,
</file context>
There was a problem hiding this comment.
Valid. We will add the 4:1 ratio against visible text to the changelog in this PR.
| if (analysis.has_template_image | ||
| && (analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_ok)) | ||
| || analysis.has_invisible_text_layer | ||
| || analysis.has_readable_invisible_text_layer |
There was a problem hiding this comment.
P2: Derive ocr_recommended from the pages that still need OCR; an all-readable scan can now be classified as Mixed while pages_needing_ocr is empty.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/detector.rs, line 326:
<comment>Derive `ocr_recommended` from the pages that still need OCR; an all-readable scan can now be classified as Mixed while `pages_needing_ocr` is empty.</comment>
<file context>
@@ -314,9 +317,13 @@ pub(crate) fn detect_from_document(
if (analysis.has_template_image
&& (analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_ok))
|| analysis.has_invisible_text_layer
+ || analysis.has_readable_invisible_text_layer
{
pages_with_template_images += 1;
</file context>
There was a problem hiding this comment.
Same point as the Devin thread on this line. No change to the field; the docs for this option will say to route on pages_needing_ocr.
| pub(crate) fn keeps_layer(&mut self, doc: &Document, page_id: ObjectId, page_num: u32) -> bool { | ||
| let (font_cmaps, style_cache) = self | ||
| .fonts | ||
| .get_or_insert_with(|| (FontCMaps::from_doc(doc), FontStyleCache::new())); |
There was a problem hiding this comment.
P2: Keep CMap construction scoped to the pages detection probes; this document-wide initialization parses fonts from unsampled pages and defeats ScanStrategy::Pages' resource bound.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/extractor/mod.rs, line 757:
<comment>Keep CMap construction scoped to the pages detection probes; this document-wide initialization parses fonts from unsampled pages and defeats `ScanStrategy::Pages`' resource bound.</comment>
<file context>
@@ -702,66 +730,42 @@ fn read_invisible_layer(
+ pub(crate) fn keeps_layer(&mut self, doc: &Document, page_id: ObjectId, page_num: u32) -> bool {
+ let (font_cmaps, style_cache) = self
+ .fonts
+ .get_or_insert_with(|| (FontCMaps::from_doc(doc), FontStyleCache::new()));
+ extract_page(
+ doc,
</file context>
There was a problem hiding this comment.
Same point as the Devin thread on this line. We will fix it in this PR, together with the cost of the probe itself.
Summary
Closes #627.
Adds an opt-in
readInvisibleTextLayeroption, so a caller can read the invisible OCR text layer (render mode 3) of a scanned page as its text instead of having the page flagged for OCR. It is off by default, and output without it is unchanged.processPdf(buffer, pages, { readInvisibleTextLayer: true }), the same third argument onprocessPdfAsync,DetectionConfig::read_invisible_text_layerin Rust, andpdf2md --read-invisible-text-layer.invisible_text_layer, and it classifies as a text page.OCR_LAYER_MIN_ALNUM(40) letters or digits, the visible text is at most a quarter of that gain, and the text is not garbage. This reads a scan whose only visible text is a download footer, as JSTOR stamps its pages. A page with a real visible body keeps its visible reading, so an OCR layer never doubles its words.Tests
test_read_invisible_text_layer_reads_the_layer_of_a_scancovers a layer alone, a layer under a visible footer drawn from page content and from a Form XObject, and a page with a visible body, which extracts as it does without the option. It fails with either the detection or the extraction change removed.On the public samples in #627, with the option:
--read-invisible-text-layertext_vektor.pdfscan_ocr_sandwich.pdfscan_ocr_fpdf2.pdfscan_ocr_sandwich.pdf+ visible footerscan_ohne_text.pdfscannedcargo test(1,952 tests),cargo fmt --check,cargo clippy -D warningswith and withoutocr, and the napi crate's clippy pass.Risk
If an OCR layer repeats a small amount of visible text, such as a heading that is both painted and in the layer, that text can show twice with the option on.
Summary by cubic
Adds an opt-in
readInvisibleTextLayeroption to read a scanned page’s invisible OCR text instead of sending it to OCR. It is off by default, so existing output is unchanged.processPdf,processPdfAsync,DetectionConfig, andpdf2md --read-invisible-text-layer.DetectionConfigstruct literals must set the new field or use..Default::default().Written for commit 289504c. Summary will update on new commits.