Search before asking
Motivation
Currently, Fluss lookup join does not support lookup custom shuffle. When lookup join is executed with normal hash shuffle, records are distributed by Flink's join-key hash, which is not aligned with Fluss bucket assignment.
This makes it difficult to combine lookup join with full-cache lookup efficiently. Each lookup subtask may receive keys from almost all Fluss buckets, so every subtask may need to load a large portion of the dimension table.
This issue aims to support lookup custom shuffle by Fluss bucket in Flink 2.2, so lookup input records can be shuffled according to Fluss bucket rules.
Solution
No response
Anything else?
No response
Willingness to contribute
Search before asking
Motivation
Currently, Fluss lookup join does not support lookup custom shuffle. When lookup join is executed with normal hash shuffle, records are distributed by Flink's join-key hash, which is not aligned with Fluss bucket assignment.
This makes it difficult to combine lookup join with full-cache lookup efficiently. Each lookup subtask may receive keys from almost all Fluss buckets, so every subtask may need to load a large portion of the dimension table.
This issue aims to support lookup custom shuffle by Fluss bucket in Flink 2.2, so lookup input records can be shuffled according to Fluss bucket rules.
Solution
No response
Anything else?
No response
Willingness to contribute