Skip to content

Decode AcroForm field names and values as PDF text strings - #616

Open
litoosh13 wants to merge 1 commit into
firecrawl:mainfrom
litoosh13:fix-form-field-text-strings
Open

litoosh13 wants to merge 1 commit into
firecrawl:mainfrom
litoosh13:fix-form-field-text-strings

Conversation

@litoosh13

@litoosh13 litoosh13 commented Oct 5, 2026 •

Copy link
Copy Markdown

Fixes #504.

Form field names (/T) and text/choice values (/V) were decoded with String::from_utf8_lossy. They are PDF
text strings (PDF 32000-1 §7.9.2.2): UTF-16 with a byte-order mark, or PDFDocEncoding. This change uses the
crate's existing decode_pdf_text_string for the three sites in src/extractor/links.rs (the /URI and the
checkbox name stay as they are, as noted in the issue).

One consequence that #504 does not mention: the U+FFFD characters end up in the page's text, the page is then
flagged needs_ocr with suspected_garbled_text, and extract_pages_markdown returns empty markdown for the
whole page
— the printed text included. A filled German form whose field names contain umlauts
(PDFDocEncoding, no UTF-16) loses all its text this way, so this is not only about Acrobat's UTF-16 values.

The test covers both encodings: a PDFDocEncoding name (Straße und Größe) and a UTF-16BE value
(Jürgen Groß). cargo test --lib: 1619 passed.

We have been carrying this fix in a patched copy of pdf-inspector in leafmind
(https://github.com/litoosh13/leafmind, third_party/pdf-inspector/leafmind-form-field-strings.patch).


Summary by cubic

Fixes #504 by decoding AcroForm field names (/T) and text/choice values (/V) as PDF text strings instead of reading them as UTF-8 with lossy replacement. This prevents non-ASCII names and values (PDFDocEncoding or UTF-16 with BOM) from becoming U+FFFD, which previously caused the page to be flagged as garbled and extract_pages_markdown to return empty markdown for the whole page. The test covers both encodings.

Written for commit d30f20b. Summary will update on new commits.

Review in cubic

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 1 file

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Shadow auto-approve: would auto-approve. Replaces lossy UTF-8 decoding of AcroForm field names and values with the existing PDF text string decoder, fixing garbled non-ASCII form data; a new test covers PDFDocEncoding and UTF-16 values.

Re-trigger cubic

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

acroform field names and values are decoded as utf-8, mangling utf-16be text strings

1 participant