Repository navigation
Conversation
There was a problem hiding this comment.
No issues found across 1 file
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Shadow auto-approve: would auto-approve. Replaces lossy UTF-8 decoding of AcroForm field names and values with the existing PDF text string decoder, fixing garbled non-ASCII form data; a new test covers PDFDocEncoding and UTF-16 values.
Re-trigger cubic
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #504.
Form field names (
/T) and text/choice values (/V) were decoded withString::from_utf8_lossy. They are PDFtext strings (PDF 32000-1 §7.9.2.2): UTF-16 with a byte-order mark, or PDFDocEncoding. This change uses the
crate's existing
decode_pdf_text_stringfor the three sites insrc/extractor/links.rs(the/URIand thecheckbox name stay as they are, as noted in the issue).
One consequence that #504 does not mention: the U+FFFD characters end up in the page's text, the page is then
flagged
needs_ocrwithsuspected_garbled_text, andextract_pages_markdownreturns empty markdown for thewhole page — the printed text included. A filled German form whose field names contain umlauts
(PDFDocEncoding, no UTF-16) loses all its text this way, so this is not only about Acrobat's UTF-16 values.
The test covers both encodings: a PDFDocEncoding name (
Straße und Größe) and a UTF-16BE value(
Jürgen Groß).cargo test --lib: 1619 passed.We have been carrying this fix in a patched copy of pdf-inspector in leafmind
(https://github.com/litoosh13/leafmind,
third_party/pdf-inspector/leafmind-form-field-strings.patch).Summary by cubic
Fixes #504 by decoding AcroForm field names (
/T) and text/choice values (/V) as PDF text strings instead of reading them as UTF-8 with lossy replacement. This prevents non-ASCII names and values (PDFDocEncoding or UTF-16 with BOM) from becoming U+FFFD, which previously caused the page to be flagged as garbled andextract_pages_markdownto return empty markdown for the whole page. The test covers both encodings.Written for commit d30f20b. Summary will update on new commits.