Skip to content

feat: read a scan's invisible OCR text layer when asked - #628

Open
logancyang wants to merge 3 commits into
firecrawl:mainfrom
logancyang:read-invisible-text-layer
Open

logancyang wants to merge 3 commits into
firecrawl:mainfrom
logancyang:read-invisible-text-layer

Conversation

@logancyang

@logancyang logancyang commented Oct 10, 2026 •

Copy link
Copy Markdown

Summary

Closes #627.

Adds an opt-in readInvisibleTextLayer option, so a caller can read the invisible OCR text layer (render mode 3) of a scanned page as its text instead of having the page flagged for OCR. It is off by default, and output without it is unchanged.

  • API. processPdf(buffer, pages, { readInvisibleTextLayer: true }), the same third argument on processPdfAsync, DetectionConfig::read_invisible_text_layer in Rust, and pdf2md --read-invisible-text-layer.
  • Detection. With the option, a page whose only text is an invisible layer under a covering image is not flagged invisible_text_layer, and it classifies as a text page.
  • Extraction. With the option, each page whose visible pass skipped invisible text is read again with it. The reading is kept when the layer adds at least OCR_LAYER_MIN_ALNUM (40) letters or digits, the visible text is at most a quarter of that gain, and the text is not garbage. This reads a scan whose only visible text is a download footer, as JSTOR stamps its pages. A page with a real visible body keeps its visible reading, so an OCR layer never doubles its words.

Tests

test_read_invisible_text_layer_reads_the_layer_of_a_scan covers a layer alone, a layer under a visible footer drawn from page content and from a Form XObject, and a page with a visible body, which extracts as it does without the option. It fails with either the detection or the extraction change removed.

On the public samples in #627, with the option:

File Default --read-invisible-text-layer
text_vektor.pdf marker found unchanged
scan_ocr_sandwich.pdf 0 chars, flagged layer text, marker found, not flagged
scan_ocr_fpdf2.pdf 0 chars, flagged layer text, marker found, not flagged
scan_ocr_sandwich.pdf + visible footer footer only footer and layer text
scan_ohne_text.pdf flagged scanned unchanged

cargo test (1,952 tests), cargo fmt --check, cargo clippy -D warnings with and without ocr, and the napi crate's clippy pass.

Risk

If an OCR layer repeats a small amount of visible text, such as a heading that is both painted and in the layer, that text can show twice with the option on.


Devin Review


Summary by cubic

Adds an opt-in readInvisibleTextLayer option to read a scanned page’s invisible OCR text instead of sending it to OCR. It is off by default, so existing output is unchanged.

  • Available through processPdf, processPdfAsync, DetectionConfig, and pdf2md --read-invisible-text-layer.
  • Reads a layer only when the page has a covering image and the layer adds at least 40 letters or digits, is not garbage, and dominates the visible text; the retry shares the page’s form budget.
  • Detection still sends pages without a readable layer to OCR, including pages missed by the scan sample.
  • Rust DetectionConfig struct literals must set the new field or use ..Default::default().
  • A small amount of visible text repeated in the OCR layer may appear twice.

Written for commit 289504c. Summary will update on new commits.

View guided diff

devin-ai-integration[bot]

This comment was marked as resolved.

Add an opt-in readInvisibleTextLayer option to processPdf and
processPdfAsync (DetectionConfig::read_invisible_text_layer, pdf2md
--read-invisible-text-layer). With it, the detector does not flag a page
whose only text is an invisible layer under a covering image, and the
extractor reads a page again with its invisible text when that adds at
least 40 letters or digits and the visible text is at most a quarter of
that gain, such as a download footer over an OCR scan. Off by default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@logancyang
logancyang force-pushed the read-invisible-text-layer branch from a655b94 to afd3a9c Compare October 10, 2026 23:55
cubic-dev-ai[bot]

This comment was marked as resolved.

Detection now clears a scan's invisible_text_layer flag only when
extraction would adopt that layer (OCR_LAYER_MIN_ALNUM, not garbage), so
a short layer keeps its page on the OCR path. The extraction retry reads
the layer only on a page with a covering image, and shares the page's
form-walk budget with the visible pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
devin-ai-integration[bot]

This comment was marked as resolved.

cubic-dev-ai[bot]

This comment was marked as resolved.

A page whose invisible layer is read as text still counts as a scan
page, so an OCR'd book is Mixed and the per-page rule decides every
page: one with no layer, or a layer extraction would not keep, goes to
OCR even when the sample left it out.

Detection asks extraction's own per-page path (the visible pass, then
the layer retry, on one form budget) whether it keeps a page's layer,
with the document's fonts read once. The reason for an unsampled page
of a scanned document goes back to the base rule.

Tests now write the render mode inside each text object, as OCR tools
do, so they exercise the retry and its 4x guard. pdf2md --help lists
--read-invisible-text-layer, and the changelog notes the gate and the
new DetectionConfig field.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 3 new potential issues.

Devin Review

Comment thread src/detector.rs
Comment thread src/detector.rs
if (analysis.has_template_image
&& (analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_ok))
|| analysis.has_invisible_text_layer
|| analysis.has_readable_invisible_text_layer

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Readable scans falsely recommend OCR

When sampled scan pages have readable layers and sufficient text operators, has_readable_invisible_text_layer forces Mixed classification. The Mixed branch sets ocr_recommended even when pages_needing_ocr is empty, prompting unnecessary OCR.

Learn more

A readable OCR layer is accepted as text by the extraction probe. Counting that page as a template forces the Mixed branch when pages_with_text > 0, and classification sets ocr_recommended unconditionally. The per-page routing then leaves a readable scan out of pages_needing_ocr, so detection recommends OCR for pages that need none.

Example: A one-page scan has 120 individually positioned invisible glyphs and a covering raster. With the option enabled, its layer is extracted and pages_needing_ocr is empty, but ocr_recommended is true.

Recommended fix: Preserve the per-page Mixed routing for unread scans, but derive the OCR recommendation from the resulting OCR list or otherwise exclude accepted layers from the recommendation calculation.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid, but narrow. Every Mixed document already sets ocr_recommended, and processPdf does not return it; pages_needing_ocr is the per-page answer. No change to the field. The docs for this option will say to route on pages_needing_ocr.

Comment thread src/extractor/mod.rs
pub(crate) fn keeps_layer(&mut self, doc: &Document, page_id: ObjectId, page_num: u32) -> bool {
let (font_cmaps, style_cache) = self
.fonts
.get_or_insert_with(|| (FontCMaps::from_doc(doc), FontStyleCache::new()));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Sampled detection parses every page's fonts

On a sampled page with a hidden layer, keeps_layer builds font maps for the entire document. FontCMaps::from_doc traverses unsampled pages and parses their embedded fonts, making sampled detection pay full-document font costs.

Learn more

Detection normally scans only the pages chosen by ScanStrategy, but the first candidate hidden layer now initializes a document-wide font map. FontCMaps::from_doc_pages walks every page and processes its page and form fonts when no filter is supplied. Thus even ScanStrategy::Pages(vec![1]) triggers font processing for every unselected page when page 1 has a hidden layer.

Example: Page 1 of a 1,000-page PDF has an OCR layer; the other 999 pages each embed a distinct TrueType font. Detecting only page 1 now parses the other pages' fonts before answering.

Recommended fix: Scope CMap construction to pages actually probed. Cache page-scoped maps or lazily extend a font map as new pages are analyzed, preserving the sampled strategy's resource bound.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid, and part of a larger cost. With the option on, detection reads every hidden-layer page twice, whatever pages were asked, and extraction then reads it again. We will fix that cost in this PR.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

3 issues found across 6 files (changes from recent commits).

Confidence score: 3/5

  • src/detector.rs can classify an all-readable scan as Mixed even when no pages need OCR. Derive ocr_recommended from the pages that still need OCR.
  • src/extractor/mod.rs builds CMaps for unsampled pages, defeating the resource bound in ScanStrategy::Pages. Keep CMap construction scoped to the pages detection probes.
  • CHANGELOG.md omits the four-to-one gate, so readers may miss that a 40+ character text layer is still rejected unless added alphanumeric text is at least four times the visible text. Document both thresholds.
Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="CHANGELOG.md">

<violation number="1" location="CHANGELOG.md:18">
P2: This omits the four-to-one gate: a 40+ character layer is still rejected unless its added alphanumeric text is at least four times the visible text. Document the added-text threshold and ratio so the changelog does not promise adoption for these pages.</violation>
</file>

<file name="src/detector.rs">

<violation number="1" location="src/detector.rs:326">
P2: Derive `ocr_recommended` from the pages that still need OCR; an all-readable scan can now be classified as Mixed while `pages_needing_ocr` is empty.</violation>
</file>

<file name="src/extractor/mod.rs">

<violation number="1" location="src/extractor/mod.rs:757">
P2: Keep CMap construction scoped to the pages detection probes; this document-wide initialization parses fonts from unsampled pages and defeats `ScanStrategy::Pages`' resource bound.</violation>
</file>

Shadow auto-approve: would not auto-approve because issues were found.

View guided diff | Re-trigger cubic

Comment thread CHANGELOG.md
option on `processPdfAsync`, `DetectionConfig::read_invisible_text_layer`
and `pdf2md --read-invisible-text-layer`, read the invisible OCR text layer
(render mode 3) of a scanned page as its text. A layer of at least 40
letters or digits that is not garbage is read, and its page is not

@cubic-dev-ai cubic-dev-ai Bot Oct 11, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This omits the four-to-one gate: a 40+ character layer is still rejected unless its added alphanumeric text is at least four times the visible text. Document the added-text threshold and ratio so the changelog does not promise adoption for these pages.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At CHANGELOG.md, line 18:

<comment>This omits the four-to-one gate: a 40+ character layer is still rejected unless its added alphanumeric text is at least four times the visible text. Document the added-text threshold and ratio so the changelog does not promise adoption for these pages.</comment>

<file context>
@@ -14,10 +14,14 @@ version and date. Earlier releases are described in their
-  stamp over such a layer, like an archive's download footer, returns the
-  layer too. Off by default; output without the option is unchanged.
+  (render mode 3) of a scanned page as its text. A layer of at least 40
+  letters or digits that is not garbage is read, and its page is not
+  flagged `invisible_text_layer`; a shorter layer keeps the page on the OCR
+  path. A page whose only visible text is a small stamp over such a layer,
</file context>
Fix with cubic

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid. We will add the 4:1 ratio against visible text to the changelog in this PR.

Comment thread src/extractor/mod.rs
Comment thread src/detector.rs
if (analysis.has_template_image
&& (analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_ok))
|| analysis.has_invisible_text_layer
|| analysis.has_readable_invisible_text_layer

@cubic-dev-ai cubic-dev-ai Bot Oct 11, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Derive ocr_recommended from the pages that still need OCR; an all-readable scan can now be classified as Mixed while pages_needing_ocr is empty.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/detector.rs, line 326:

<comment>Derive `ocr_recommended` from the pages that still need OCR; an all-readable scan can now be classified as Mixed while `pages_needing_ocr` is empty.</comment>

<file context>
@@ -314,9 +317,13 @@ pub(crate) fn detect_from_document(
             if (analysis.has_template_image
                 && (analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_ok))
                 || analysis.has_invisible_text_layer
+                || analysis.has_readable_invisible_text_layer
             {
                 pages_with_template_images += 1;
</file context>
Fix with cubic

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same point as the Devin thread on this line. No change to the field; the docs for this option will say to route on pages_needing_ocr.

Comment thread src/extractor/mod.rs
pub(crate) fn keeps_layer(&mut self, doc: &Document, page_id: ObjectId, page_num: u32) -> bool {
let (font_cmaps, style_cache) = self
.fonts
.get_or_insert_with(|| (FontCMaps::from_doc(doc), FontStyleCache::new()));

@cubic-dev-ai cubic-dev-ai Bot Oct 11, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Keep CMap construction scoped to the pages detection probes; this document-wide initialization parses fonts from unsampled pages and defeats ScanStrategy::Pages' resource bound.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/extractor/mod.rs, line 757:

<comment>Keep CMap construction scoped to the pages detection probes; this document-wide initialization parses fonts from unsampled pages and defeats `ScanStrategy::Pages`' resource bound.</comment>

<file context>
@@ -702,66 +730,42 @@ fn read_invisible_layer(
+    pub(crate) fn keeps_layer(&mut self, doc: &Document, page_id: ObjectId, page_num: u32) -> bool {
+        let (font_cmaps, style_cache) = self
+            .fonts
+            .get_or_insert_with(|| (FontCMaps::from_doc(doc), FontStyleCache::new()));
+        extract_page(
+            doc,
</file context>
Fix with cubic

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same point as the Devin thread on this line. We will fix it in this PR, together with the cost of the probe itself.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Option to read a scan's invisible OCR text layer instead of flagging it for OCR

1 participant