Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,25 @@ version. A separate release pull request bumps the manifests with
version and date. Earlier releases are described in their
[GitHub releases](https://github.com/firecrawl/pdf-inspector/releases).

## [Unreleased]

### Fixed

- A TrueType subset with no `cmap`, no glyph names and no ToUnicode that
comes from a core Windows font (Arial, Times New Roman, Verdana, ...) is
read through the Windows variant of the standard glyph order. That
variant leaves out the Macintosh no-break space (172) and `apple` (210)
and has the euro sign in the place of the currency sign, so the two
orders diverge from glyph 172 on, and the Macintosh order read those
glyphs as their neighbours: `“N/A”` as `—N/A“`, `company’s` as
`company‘s` and en dashes as `œ`. The font's
own advances choose the variant, since an accented letter advances like
its base letter.
- A glyph in such a subset past the standard order that paints nothing but
advances reads as a space. Word justifies lines with Arial's en and em
spaces, which came out as U+FFFD between every word
(`Proposer�Legal�Name`) and sent the pages to OCR as garbled text.

## [1.25.2] - 2026-09-28

Changes since 1.25.1.
Expand Down
143 changes: 131 additions & 12 deletions src/mac_glyph_order.rs
Original file line number Diff line number Diff line change
Expand Up @@ -3,20 +3,24 @@
//! A TrueType subset embedded for an Identity-H CIDFont sometimes keeps
//! neither a `cmap` table nor glyph names, and the PDF carries no
//! ToUnicode. The glyph IDs are then the only handle on the text. The core
//! Windows and Macintosh fonts (Arial, Times New Roman, Helvetica, ...) place
//! the 258 glyphs of the standard Macintosh ordering first, so for those
//! fonts the glyph ID alone still names the character. The order is only
//! assumed when the font's own metrics corroborate it.
//! Macintosh fonts (Helvetica, Times, ...) place the 258 glyphs of the
//! standard Macintosh ordering first, and the core Windows fonts (Arial,
//! Times New Roman, ...) a variant of it, so for those fonts the glyph ID
//! alone still names the character. The order is only assumed when the
//! font's own metrics corroborate it, and the font's advances choose between
//! the two variants.

use std::sync::LazyLock;

use crate::glyph_names::glyph_to_char;
use crate::tounicode::ToUnicodeCMap;
use lopdf::{Document, Object};

/// The 258 glyph names of the standard Macintosh ordering (`post` format 1;
/// TrueType Reference Manual, "The 'post' table"). The core Windows and
/// Macintosh fonts (Arial, Times New Roman, Helvetica, ...) place these
/// glyphs first in this order, so a subset that dropped its `cmap` and its
/// glyph names still exposes them through the glyph ID alone.
/// TrueType Reference Manual, "The 'post' table"). The core Macintosh fonts
/// (Helvetica, Times, ...) place these glyphs first in this order, so a
/// subset that dropped its `cmap` and its glyph names still exposes them
/// through the glyph ID alone.
const MAC_GLYPH_NAMES: [&str; 258] = [
".notdef",
".null",
Expand Down Expand Up @@ -278,6 +282,98 @@ const MAC_GLYPH_NAMES: [&str; 258] = [
"dcroat",
];

/// The first slot at which [`WINDOWS_GLYPH_NAMES`] departs from
/// [`MAC_GLYPH_NAMES`]; every slot before it names the same glyph in both.
const FIRST_WINDOWS_ONLY_SLOT: u16 = 172;

/// The order the core Windows fonts (Arial, Times New Roman, Courier New,
/// Verdana, Georgia, Tahoma) give their first 258 glyphs: the Macintosh
/// order without `nonbreakingspace` (172) and `apple` (210), with `Euro` in
/// the place of `currency`, closed by `overscore` and `middot`. From slot
/// 172 on the two orders name different glyphs, so a Windows subset read
/// through the Macintosh order turns curly quotes into dashes (“N/A” as
/// —N/A“), apostrophes into opening quotes and en dashes into `œ`.
static WINDOWS_GLYPH_NAMES: LazyLock<Vec<&'static str>> = LazyLock::new(|| {
MAC_GLYPH_NAMES
.iter()
.copied()
.filter(|name| !matches!(*name, "nonbreakingspace" | "apple"))
.map(|name| if name == "currency" { "Euro" } else { name })
.chain(["overscore", "middot"])
.collect()
});

/// The glyph a slot named `name` advances like in either order: the base
/// letter of a precomposed accented letter (`Agrave` advances like `A` at
/// 36, `ydieresis` like `y` at 92), and the space at 3 for the no-break
/// space. Both orders keep the space and the letters at the same slots.
fn advance_twin(name: &str) -> Option<u16> {
const ACCENTS: [&str; 10] = [
"grave",
"acute",
"circumflex",
"tilde",
"dieresis",
"caron",
"breve",
"cedilla",
"dotaccent",
"ring",
];
if name == "nonbreakingspace" {
return Some(3);
}
let mut chars = name.chars();
let base = chars.next()?;
if !ACCENTS.contains(&chars.as_str()) {
return None;
}
match base {
'A'..='Z' => Some(36 + (base as u16 - 'A' as u16)),
'a'..='z' => Some(68 + (base as u16 - 'a' as u16)),
_ => None,
}
}

/// The variant of the standard order the font's advances support. Each
/// order is scored over the slots from 172 on that it names after a twin
/// ([`advance_twin`]): how many advance like their twin and how many do
/// not. The advances of glyphs a subset dropped still count, since
/// subsetters that keep glyph IDs usually keep the metrics too. The
/// Windows order is taken only when it fits more slots and misfits fewer;
/// equal evidence, as in a monospaced font or a subset that zeroed its
/// dropped metrics, keeps the Macintosh order.
fn standard_order_names(face: &ttf_parser::Face) -> &'static [&'static str] {
let advance = |gid: u16| {
face.glyph_hor_advance(ttf_parser::GlyphId(gid))
.filter(|&w| w > 0)
};
let score = |names: &[&str]| {
let (mut fits, mut misfits) = (0usize, 0usize);
let last = face.number_of_glyphs().min(names.len() as u16);
for gid in FIRST_WINDOWS_ONLY_SLOT..last {
let Some(twin) = advance_twin(names[usize::from(gid)]) else {
continue;
};
if let (Some(own), Some(theirs)) = (advance(gid), advance(twin)) {
if own == theirs {
fits += 1;
} else {
misfits += 1;
}
}
}
(fits, misfits)
};
let (mac_fits, mac_misfits) = score(&MAC_GLYPH_NAMES);
let (windows_fits, windows_misfits) = score(&WINDOWS_GLYPH_NAMES);
if windows_fits > mac_fits && windows_misfits < mac_misfits {
&WINDOWS_GLYPH_NAMES
} else {
&MAC_GLYPH_NAMES
}
}

/// Whether the CIDFont maps CIDs to glyph IDs without a table, so the CIDs
/// in the content stream are the embedded font's glyph IDs.
pub(crate) fn cid_to_gid_is_identity(cid_font_dict: &lopdf::Dictionary, doc: &Document) -> bool {
Expand All @@ -299,15 +395,21 @@ pub(crate) fn font_file_follows_mac_order(font_data: &[u8]) -> bool {
}

/// Build a GID→Unicode CMap for a TrueType font that has neither a `cmap`
/// table nor glyph names, assuming the standard Macintosh glyph order.
/// table nor glyph names, assuming the standard Macintosh glyph order or
/// its Windows variant ([`standard_order_names`]).
///
/// The assumption is only accepted when the font's own metrics corroborate
/// it: at least three glyphs at the digit positions (19–28) exist and share
/// one advance, as tabular figures do; the glyph at the space position (3)
/// advances without an outline; every present `i` and `l` is narrower than
/// every present `m` and `w`;
/// and capitals average wider than lowercase. Each check only applies to
/// glyphs the subset kept. Everything else keeps today's behaviour.
/// glyphs the subset kept, and all of them to slots both variants share.
/// Everything else keeps today's behaviour.
///
/// A kept glyph past the order that paints nothing but advances reads as a
/// space: the core Windows fonts draw U+2000–U+200A (en space, em space,
/// ...) that way, and Word justifies lines with them.
pub(crate) fn build_cmap_from_mac_glyph_order(font_data: &[u8]) -> Option<ToUnicodeCMap> {
use ttf_parser::GlyphId;

Expand Down Expand Up @@ -370,9 +472,10 @@ pub(crate) fn build_cmap_from_mac_glyph_order(font_data: &[u8]) -> Option<ToUnic

// Only slots the subset kept: an outline, or a blank glyph that still
// advances (space, no-break space).
let names = standard_order_names(&face);
let mut cmap = ToUnicodeCMap::new();
let count = usize::from(face.number_of_glyphs()).min(MAC_GLYPH_NAMES.len());
for (gid, name) in MAC_GLYPH_NAMES.iter().enumerate().take(count) {
let count = usize::from(face.number_of_glyphs()).min(names.len());
for (gid, name) in names.iter().enumerate().take(count) {
let gid = gid as u16;
let kept = present(gid) || face.glyph_hor_advance(GlyphId(gid)).is_some_and(|w| w > 0);
if !kept {
Expand All @@ -382,6 +485,22 @@ pub(crate) fn build_cmap_from_mac_glyph_order(font_data: &[u8]) -> Option<ToUnic
cmap.char_map.insert(gid, ch.to_string());
}
}
// Past the order a glyph has no name to read, but one that paints
// nothing and advances is a space whatever its code point. Only glyphs
// with an outline record count: the core fonts give those spaces a
// one-point outline, while a glyph the subset dropped has no record,
// though it often keeps its advance.
if let Some(glyf) = face.tables().glyf {
for gid in (names.len() as u16)..face.number_of_glyphs() {
let id = GlyphId(gid);
if glyf.bbox(id).is_some()
&& face.glyph_bounding_box(id).is_none()
&& face.glyph_hor_advance(id).is_some_and(|w| w > 0)
{
cmap.char_map.insert(gid, " ".to_string());
}
}
}
if cmap.char_map.is_empty() {
return None;
}
Expand Down
133 changes: 128 additions & 5 deletions src/tounicode_mac_order_tests.rs
Original file line number Diff line number Diff line change
Expand Up @@ -6,12 +6,27 @@ use lopdf::{dictionary, Document, Object, Stream};
/// `glyf` but no `cmap` and no `post` names. Glyph `i` is `(outlined,
/// advance)`; a missing glyph is `(false, 0)`.
fn cmapless_truetype(glyphs: &[(bool, u16)]) -> Vec<u8> {
truetype(glyphs, false, false)
truetype(glyphs, &[], false, false)
}

/// The same font, with the glyphs at `blank_records` drawn as a single
/// point: an outline record that paints nothing, the way the core Windows
/// fonts draw their en and em spaces.
fn cmapless_truetype_with_blank_records(
glyphs: &[(bool, u16)],
blank_records: &[usize],
) -> Vec<u8> {
truetype(glyphs, blank_records, false, false)
}

/// The same font, optionally with a (3,1) `cmap` mapping `A` to glyph 36,
/// or a `post` format 2 table naming glyph 36 `A`.
fn truetype(glyphs: &[(bool, u16)], with_cmap: bool, with_post_names: bool) -> Vec<u8> {
fn truetype(
glyphs: &[(bool, u16)],
blank_records: &[usize],
with_cmap: bool,
with_post_names: bool,
) -> Vec<u8> {
let square: Vec<u8> = {
let mut g = Vec::new();
g.extend(1i16.to_be_bytes());
Expand All @@ -26,12 +41,25 @@ fn truetype(glyphs: &[(bool, u16)], with_cmap: bool, with_post_names: bool) -> V
}
g
};
let point: Vec<u8> = {
let mut g = Vec::new();
g.extend(1i16.to_be_bytes());
g.extend([0u8; 8]); // bbox 0,0,0,0
g.extend(0u16.to_be_bytes()); // endPtsOfContours
g.extend(0u16.to_be_bytes()); // instructionLength
g.push(1); // on curve, word coordinates
g.extend([0u8; 4]); // x, y
g.extend([0u8; 3]); // pad to a 4-byte boundary
g
};
let mut glyf = Vec::new();
let mut loca: Vec<u8> = Vec::new();
loca.extend(0u32.to_be_bytes());
for &(outlined, _) in glyphs {
for (gid, &(outlined, _)) in glyphs.iter().enumerate() {
if outlined {
glyf.extend(&square);
} else if blank_records.contains(&gid) {
glyf.extend(&point);
}
loca.extend((glyf.len() as u32).to_be_bytes());
}
Expand Down Expand Up @@ -188,11 +216,11 @@ fn mac_order_is_not_used_when_the_font_has_a_cmap_or_glyph_names() {
// A cmap or post names are authoritative; those fonts take the regular
// embedded-cmap and glyph-name paths instead.
let glyphs = arial_like(&[556; 10], (false, 278));
let with_cmap = truetype(&glyphs, true, false);
let with_cmap = truetype(&glyphs, &[], true, false);
let face = ttf_parser::Face::parse(&with_cmap, 0).unwrap();
assert_eq!(face.glyph_index('A').map(|g| g.0), Some(36));
assert!(build_cmap_from_mac_glyph_order(&with_cmap).is_none());
let with_names = truetype(&glyphs, false, true);
let with_names = truetype(&glyphs, &[], false, true);
let face = ttf_parser::Face::parse(&with_names, 0).unwrap();
assert_eq!(face.glyph_name(ttf_parser::GlyphId(36)), Some("A"));
assert!(build_cmap_from_mac_glyph_order(&with_names).is_none());
Expand Down Expand Up @@ -346,3 +374,98 @@ fn mac_order_declines_a_unicode_keyed_font() {
}
assert!(build_cmap_from_mac_glyph_order(&cmapless_truetype(&glyphs)).is_none());
}

/// A core font laid out in `names` order, every slot kept: tabular digits,
/// a blank space and no-break space, capitals and lowercase each one width
/// apart, every accented letter as wide as its base letter, and the other
/// glyphs at widths of their own.
fn core_font(names: &[&str]) -> Vec<(bool, u16)> {
let letter_width = |c: char| match c {
'i' | 'l' => 222,
'm' => 833,
'w' => 722,
'A'..='Z' => 600 + 10 * (c as u16 - 'A' as u16),
'a'..='z' => 400 + 5 * (c as u16 - 'a' as u16),
_ => unreachable!(),
};
names
.iter()
.enumerate()
.map(|(gid, &name)| {
if matches!(name, "space" | "nonbreakingspace") {
return (false, 278);
}
if (19..=28).contains(&gid) {
return (true, 556);
}
let mut chars = name.chars();
let first = chars.next().unwrap();
if first.is_ascii_alphabetic()
&& (chars.as_str().is_empty() || advance_twin(name).is_some())
{
return (true, letter_width(first));
}
(true, 300 + (gid as u16 % 50) * 7)
})
.collect()
}

#[test]
fn windows_order_reads_a_core_windows_subset() {
// Arial and Times New Roman leave out the Macintosh no-break space at
// 172, so from there each glyph sits one slot earlier: “N/A” read
// through the Macintosh order would come out as —N/A“.
let font = cmapless_truetype(&core_font(&WINDOWS_GLYPH_NAMES));
let cmap = build_cmap_from_mac_glyph_order(&font).expect("corroborated");
assert_eq!(decode(&cmap, &[179, 49, 18, 36, 180]), "“N/A”");
assert_eq!(decode(&cmap, &[177, 3, 182, 3, 188]), "– ’ €");
}

#[test]
fn mac_order_still_reads_a_core_macintosh_subset() {
let font = cmapless_truetype(&core_font(&MAC_GLYPH_NAMES));
let cmap = build_cmap_from_mac_glyph_order(&font).expect("corroborated");
assert_eq!(decode(&cmap, &[180, 49, 18, 36, 181]), "“N/A”");
assert_eq!(decode(&cmap, &[178, 3, 179, 3, 183]), "– — ’");
}

#[test]
fn mac_order_stands_without_advances_to_tell_the_orders_apart() {
// A subset that zeroed the metrics of the glyphs it dropped leaves no
// accented letter to compare, so the Macintosh reading stands.
let mut glyphs = core_font(&WINDOWS_GLYPH_NAMES);
for (gid, glyph) in glyphs.iter_mut().enumerate().skip(172) {
if !matches!(gid, 179 | 180) {
*glyph = (false, 0);
}
}
let cmap = build_cmap_from_mac_glyph_order(&cmapless_truetype(&glyphs)).unwrap();
assert_eq!(decode(&cmap, &[179, 180]), "—“");
}

#[test]
fn windows_glyph_names_differ_from_the_mac_order_from_slot_172() {
assert_eq!(WINDOWS_GLYPH_NAMES.len(), MAC_GLYPH_NAMES.len());
let first_difference =
(0..MAC_GLYPH_NAMES.len()).find(|&gid| WINDOWS_GLYPH_NAMES[gid] != MAC_GLYPH_NAMES[gid]);
assert_eq!(first_difference, Some(usize::from(FIRST_WINDOWS_ONLY_SLOT)));
assert_eq!(WINDOWS_GLYPH_NAMES[188], "Euro");
assert_eq!(WINDOWS_GLYPH_NAMES[209], "Ograve");
}

#[test]
fn a_blank_glyph_past_the_order_reads_as_a_space() {
// Word justifies lines with Arial's en space (glyph 3024), a one-point
// outline that advances. A dropped glyph that kept its advance has no
// outline record and stays unmapped, as do a zero-width blank and an
// outlined glyph past the order.
let mut glyphs = core_font(&WINDOWS_GLYPH_NAMES);
glyphs.extend([(false, 500), (false, 500), (false, 0), (true, 600)]);
let font = cmapless_truetype_with_blank_records(&glyphs, &[258, 260]);
let cmap = build_cmap_from_mac_glyph_order(&font).unwrap();
assert_eq!(cmap.lookup(258).as_deref(), Some(" "));
for gid in [259, 260, 261] {
assert_eq!(cmap.lookup(gid), None, "glyph {gid}");
}
assert_eq!(decode(&cmap, &[51, 85, 82, 258, 47]), "Pro L");
}