Skip to content

Fix kafka plugin converting characters to Unicode escape sequences (CMEM-8068) - #197

Merged
msaipraneeth merged 1 commit into
mainfrom
bugfix/UnicodeEscape-CMEM-8068
Sep 8, 2026
Merged

msaipraneeth merged 1 commit into
mainfrom
bugfix/UnicodeEscape-CMEM-8068

Conversation

@msaipraneeth

Copy link
Copy Markdown
Contributor

Summary

  • json.dumps() defaults to ensure_ascii=True, which silently escaped non-ASCII characters (e.g. ö, ü) into \uXXXX sequences in three places:
    • get_message_with_json_wrapper() — consumed Kafka messages aggregated into the output JSON dataset file (Consumer flow, JSON dataset target)
    • KafkaJSONDataHandler._split_data() — JSON dataset content re-serialized into a produced Kafka message value (Producer flow, JSON dataset source)
    • KafkaEntitiesDataHandler._split_data() — entity values serialized into a produced Kafka message value (Producer flow, entities source)
  • Added ensure_ascii=False to all three so the JSON representation of the original content is preserved.

All three move original user content into a new artifact (a dataset file or a Kafka message payload) — the escaping kept the JSON technically valid, but changed the textual representation of that content, which is exactly what CMEM-8068 reports.

Test plan

  • Added 3 regression tests exercising _split_data()/get_message_with_json_wrapper() directly (no Kafka broker or CMEM connection needed, since these are pure in-memory transforms); confirmed each fails against the pre-fix code with the exact ö/ü pattern, then passes with the fix.
  • ruff check and mypy (package + tests) pass.
  • Full pytest run: all tests that don't require a live CMEM connection pass; the CMEM-gated integration tests error with a pre-existing, unrelated 401 from this environment's local DataIntegration instance.

json.dumps() defaults to ensure_ascii=True, which escaped non-ASCII
characters (e.g. ö, ü) into unicode escape sequences in three places:

- get_message_with_json_wrapper(): consumed Kafka messages aggregated
  into the output JSON dataset file
- KafkaJSONDataHandler._split_data(): JSON dataset content
  re-serialized into a produced Kafka message value
- KafkaEntitiesDataHandler._split_data(): entity values serialized
  into a produced Kafka message value

All three move original user content into a new artifact (a dataset
file or a Kafka message payload), so the escaping changed the
representation of that content even though it stayed valid JSON.
@msaipraneeth msaipraneeth changed the title Fix jq plugin converting characters to Unicode escape sequences (CMEM-8068) Fix kafka plugin converting characters to Unicode escape sequences (CMEM-8068) Sep 7, 2026
@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown

Coverage

Coverage Report
File Stmts Miss Cover Missing
init.py 0 0 100%
constants.py 24 0 100%
kafka_handlers.py 196 15 92% 75 104 127 190 196-197 233 246 296-298 310 388-389 394
utils.py 221 53 76% 76 94 108 151-158 160-164 166-169 176 192-193 279 332 365-366 371-375 395-396 401 410-412 414-417 419-423 425 427 434 436-438
workflow/init.py 0 0 100%
workflow/consumer.py 87 8 91% 177 208-210 214-216 235
workflow/producer.py 72 5 93% 184 215-217 233
TOTAL 600 81 87%  

Tests Skipped Failures Errors Time
23 1 💤 0 ❌ 0 🔥 1717.829 ⏱

@msaipraneeth
msaipraneeth merged commit b21032d into main Sep 8, 2026
2 checks passed
@msaipraneeth
msaipraneeth deleted the bugfix/UnicodeEscape-CMEM-8068 branch September 8, 2026 11:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant