Repository navigation
Conversation
… its program does not A ToUnicode CMap of fewer than ten entries yields the primary role to the embedded program's own reading and is kept as the alternative. The byte-level reading of a simple font never looked at that alternative, so a code the CMap maps but the program does not read (its glyph has no usable name) fell through to the printable-byte guess: macOS Quartz writes such fonts for the few accented capitals of a heading, and `CÔNG TY` read as `CÔ!G TY`. The alternative is now read before that guess, after the program's reading, the Differences and the base encoding, so only text that was a guess changes. A simple font has no other alternative: the subset remap is made for Identity-H/V composite fonts alone.
…reads as U+FFFD The sparse ToUnicode CMap of a simple font was read for text alone. A code it maps to a control character, which nothing repaired because the program does not name the glyph, fell through to the printable-byte guess and read as its byte value. The code is now marked as any other CMap's control destination is, so it reads as U+FFFD.
| if is_simple_font { | ||
| match entry.remapped.as_ref().map(|c| c.lookup_code(code)) { | ||
| Some(CodeMapping::Text(text)) if !text.contains('\u{FFFD}') => { | ||
| return Some(text); | ||
| } | ||
| Some(CodeMapping::ControlDestination) => control_destination = true, | ||
| _ => {} | ||
| } |
There was a problem hiding this comment.
🔴 Sparse mappings lose to font encodings
With a declared base encoding, entry.remapped never reads codes that entry.primary misses; the base returns first. A sparse ToUnicode mapping from 0x21 to c still reads ! under WinAnsi, and unreadable /Differences names drop mapped codes.
Learn more
A sparse ToUnicode CMap is moved into remapped when a font-program reading takes the primary role in cmap_entry. The new lookup occurs after differences_reading and the base-encoding return. Thus an explicit /BaseEncoding supplies a character even for a code the sparse CMap maps differently; an unreadable Differences name exits before the CMap can read the code at all. Control destinations in the sparse map likewise cannot be marked when the base encoding has already returned.
Example: A simple font with a sparse <21> <0063> ToUnicode entry, an embedded program that reads only code 0x22, and /BaseEncoding /WinAnsiEncoding reads byte 0x21 as ! rather than c. With an unresolved /Differences name at 0x21 and another readable name keeping the encoding active, it omits c altogether.
Recommended fix: After the primary program reading, consult the demoted ToUnicode map before returning a base encoding or discarding an unreadable Differences code. Preserve the existing priority for a Differences name that actually identifies a readable glyph, and apply control-destination handling consistently.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
0 issues found across 2 files (changes from recent commits).
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Shadow auto-approve: would auto-approve. Fixes a simple font's sparse ToUnicode CMap being ignored by reading it before the printable-byte guess, with tests pinning the corrected codes and control-character handling. The change is focused and clearly corrects misread text.
Turn on auto-fix | Re-trigger cubic
Fixes #620.
Summary
A ToUnicode CMap of fewer than ten entries yields the primary role to the
embedded program's reading and is kept as the alternative (
cmap_entry).The byte-level reading of a simple font never looks at that alternative.
So a code the CMap maps, but whose glyph the program cannot name, falls
through to the printable-byte guess and reads as its byte value.
macOS Quartz writes such fonts: the letters of a heading go into a small
TrueType subset with codes from
0x21and a short ToUnicode CMap. Thereproduction in #620 (Times New Roman Bold):
BIÊN BẢNBIÊ! BẢ!BIÊN BẢNpdftotextreads the same file right.Fix
In the byte-level reading, a simple font's alternative CMap is read just
before the printable-byte guess: after the program's reading, the
/Differencesand the base encoding. Only text that was a guess (ordropped) can change. It is limited to simple fonts because their only
alternative is the demoted sparse CMap; the subset remap is made for
Identity-H/V composite fonts alone.
A code the sparse CMap maps to a control character is marked as any
other CMap's control destination is: it reads as U+FFFD and is not
guessed from its byte value (second commit, from cubic's review).
The entry could arguably go earlier (ahead of
/Differencesand the baseencoding, as a ToUnicode CMap normally does). I kept it last to keep the
change small; say if you prefer the other order.
Tests
a_simple_fonts_sparse_cmap_reads_the_codes_the_program_does_not(synthetic font, no fixture file). Reads
!o$$before,coeeafter;a code both read still takes the program's reading.
a_simple_fonts_sparse_cmap_marks_its_control_destination.A code the sparse CMap maps to U+0003 reads
#before the secondcommit, U+FFFD after.
cargo fmt --all -- --check,cargo clippy -- -D warnings,cargo clippy --features ocr -- -D warnings,cargo test(1620 unit,292 integration, all pass),
scripts/version.py --check, script tests.pdf2mdoutput of all 44 PDFs intests/fixtures/is byte-identicalbefore and after.
BIÊN BẢNreads asBIÊ! BẢ!#620 and two more Quartz-made PDFs with the fault now readright; four from the same generator that read right before are unchanged.
Not verified
pdf-evalsis not public, sobench.py test/bench.py scorewerenot run. Please run them before merge.
The script in A simple font's short ToUnicode CMap is ignored for codes the embedded program does not name:
BIÊN BẢNreads asBIÊ! BẢ!#620 makes one on macOS.Summary by cubic
Fixes sparse ToUnicode CMap entries being ignored for simple fonts, so codes the embedded program can't name no longer misread as their byte values.
Written for commit ebf28a2. Summary will update on new commits.