Skip to content

fix(extractor): a simple font's sparse ToUnicode CMap reads the codes its program does not - #621

Open
tlq5l wants to merge 2 commits into
firecrawl:mainfrom
tlq5l:fix/sparse-tounicode-simple-font
Open

tlq5l wants to merge 2 commits into
firecrawl:mainfrom
tlq5l:fix/sparse-tounicode-simple-font

Conversation

@tlq5l

@tlq5l tlq5l commented Oct 6, 2026 •

Copy link
Copy Markdown

Fixes #620.

Summary

A ToUnicode CMap of fewer than ten entries yields the primary role to the
embedded program's reading and is kept as the alternative (cmap_entry).
The byte-level reading of a simple font never looks at that alternative.
So a code the CMap maps, but whose glyph the program cannot name, falls
through to the printable-byte guess and reads as its byte value.

macOS Quartz writes such fonts: the letters of a heading go into a small
TrueType subset with codes from 0x21 and a short ToUnicode CMap. The
reproduction in #620 (Times New Roman Bold):

text
source BIÊN BẢN
before BIÊ! BẢ!
after BIÊN BẢN

pdftotext reads the same file right.

Fix

In the byte-level reading, a simple font's alternative CMap is read just
before the printable-byte guess: after the program's reading, the
/Differences and the base encoding. Only text that was a guess (or
dropped) can change. It is limited to simple fonts because their only
alternative is the demoted sparse CMap; the subset remap is made for
Identity-H/V composite fonts alone.

A code the sparse CMap maps to a control character is marked as any
other CMap's control destination is: it reads as U+FFFD and is not
guessed from its byte value (second commit, from cubic's review).

The entry could arguably go earlier (ahead of /Differences and the base
encoding, as a ToUnicode CMap normally does). I kept it last to keep the
change small; say if you prefer the other order.

Tests

  • Unit: a_simple_fonts_sparse_cmap_reads_the_codes_the_program_does_not
    (synthetic font, no fixture file). Reads !o$$ before, coee after;
    a code both read still takes the program's reading.
  • Unit: a_simple_fonts_sparse_cmap_marks_its_control_destination.
    A code the sparse CMap maps to U+0003 reads # before the second
    commit, U+FFFD after.
  • cargo fmt --all -- --check, cargo clippy -- -D warnings,
    cargo clippy --features ocr -- -D warnings, cargo test (1620 unit,
    292 integration, all pass), scripts/version.py --check, script tests.
  • pdf2md output of all 44 PDFs in tests/fixtures/ is byte-identical
    before and after.
  • The PDF from A simple font's short ToUnicode CMap is ignored for codes the embedded program does not name: BIÊN BẢN reads as BIÊ! BẢ! #620 and two more Quartz-made PDFs with the fault now read
    right; four from the same generator that read right before are unchanged.

Not verified


Summary by cubic

Fixes sparse ToUnicode CMap entries being ignored for simple fonts, so codes the embedded program can't name no longer misread as their byte values.

  • Reads the sparse CMap as the last resort before the printable-byte guess, so only text that was previously guessed (or dropped) can change.
  • The program's own reading still takes precedence where both map the same code.
  • A code the CMap maps to a control character now reads as U+FFFD instead of its byte value.

Written for commit ebf28a2. Summary will update on new commits.

Review in cubic


Devin Review

… its program does not

A ToUnicode CMap of fewer than ten entries yields the primary role to the
embedded program's own reading and is kept as the alternative. The
byte-level reading of a simple font never looked at that alternative, so
a code the CMap maps but the program does not read (its glyph has no
usable name) fell through to the printable-byte guess: macOS Quartz
writes such fonts for the few accented capitals of a heading, and
`CÔNG TY` read as `CÔ!G TY`.

The alternative is now read before that guess, after the program's
reading, the Differences and the base encoding, so only text that was a
guess changes. A simple font has no other alternative: the subset remap
is made for Identity-H/V composite fonts alone.
cubic-dev-ai[bot]

This comment was marked as resolved.

…reads as U+FFFD

The sparse ToUnicode CMap of a simple font was read for text alone. A
code it maps to a control character, which nothing repaired because the
program does not name the glyph, fell through to the printable-byte
guess and read as its byte value.

The code is now marked as any other CMap's control destination is, so
it reads as U+FFFD.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 potential issue.

Devin Review

Comment thread src/extractor/fonts.rs
Comment on lines +2666 to +2673
if is_simple_font {
match entry.remapped.as_ref().map(|c| c.lookup_code(code)) {
Some(CodeMapping::Text(text)) if !text.contains('\u{FFFD}') => {
return Some(text);
}
Some(CodeMapping::ControlDestination) => control_destination = true,
_ => {}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Sparse mappings lose to font encodings

With a declared base encoding, entry.remapped never reads codes that entry.primary misses; the base returns first. A sparse ToUnicode mapping from 0x21 to c still reads ! under WinAnsi, and unreadable /Differences names drop mapped codes.

Learn more

A sparse ToUnicode CMap is moved into remapped when a font-program reading takes the primary role in cmap_entry. The new lookup occurs after differences_reading and the base-encoding return. Thus an explicit /BaseEncoding supplies a character even for a code the sparse CMap maps differently; an unreadable Differences name exits before the CMap can read the code at all. Control destinations in the sparse map likewise cannot be marked when the base encoding has already returned.

Example: A simple font with a sparse <21> <0063> ToUnicode entry, an embedded program that reads only code 0x22, and /BaseEncoding /WinAnsiEncoding reads byte 0x21 as ! rather than c. With an unresolved /Differences name at 0x21 and another readable name keeping the encoding active, it omits c altogether.

Recommended fix: After the primary program reading, consult the demoted ToUnicode map before returning a base encoding or discarding an unreadable Differences code. Preserve the existing priority for a Differences name that actually identifies a readable glyph, and apply control-destination handling consistently.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

0 issues found across 2 files (changes from recent commits).

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Shadow auto-approve: would auto-approve. Fixes a simple font's sparse ToUnicode CMap being ignored by reading it before the printable-byte guess, with tests pinning the corrected codes and control-character handling. The change is focused and clearly corrects misread text.

Turn on auto-fix | Re-trigger cubic

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A simple font's short ToUnicode CMap is ignored for codes the embedded program does not name: BIÊN BẢN reads as BIÊ! BẢ!

1 participant