I’m seeing a repeatable speaker-routing issue with MiniMax H3 Ref2VA in multi-character dialogue scenes.
Visual subject binding works extremely well: multiple characters referenced from separate images remain visually distinct and consistent. However, the audio branch can fail to keep different voice identities attached to the correct subjects. One subject’s voice characteristics can bleed into another subject, or one voice can effectively dominate both speakers.
Example setup:
<Subject 1> = female character from <Picture 1>
Voice A = distinct British female voice anchor
<Subject 2> = female character from <Picture 2>
Voice B = distinct American female voice anchor
Dialogue alternates: Subject 2 → Subject 1 → Subject 2.
The visual identities and speaker/lip assignments remain correct, but Voice B can bleed into Subject 1, causing both subjects to share or partially share the same vocal identity, or Subject 1 uses a completely different voice that wasn't prompted.
Update: I originally reported this as a Ref2VA-specific issue, but I have now reproduced the same multi-speaker voice bleed using FL2VA from a single composed first frame with no last frame.
In FL2VA, explicitly conditioning Subject 2 to speak in an American voice causes Subject 1’s voice to become American as well. Conversely, explicitly conditioning Subject 1 to speak in an Elizabeth Swann / Keira Knightley-style British voice restores Subject 1’s voice, but Subject 2 also shifts to a British voice.
The issue also persists with only one speaker transition (A → B), with explicit tags, and with pure <Subject 1> / <Subject 2> speaker labels.
This suggests the issue may not be specific to Ref2VA reference routing. It may instead involve multi-speaker voice conditioning being applied too globally across the audio stream rather than remaining reliably bound to an individual speaker.
For comparison, other two-character combinations can work correctly, so this does not appear to be a universal inability to generate two distinct voices.
I’m seeing a repeatable speaker-routing issue with MiniMax H3 Ref2VA in multi-character dialogue scenes.
Visual subject binding works extremely well: multiple characters referenced from separate images remain visually distinct and consistent. However, the audio branch can fail to keep different voice identities attached to the correct subjects. One subject’s voice characteristics can bleed into another subject, or one voice can effectively dominate both speakers.
Example setup:
<Subject 1> = female character from <Picture 1>
Voice A = distinct British female voice anchor
<Subject 2> = female character from <Picture 2>
Voice B = distinct American female voice anchor
Dialogue alternates: Subject 2 → Subject 1 → Subject 2.
The visual identities and speaker/lip assignments remain correct, but Voice B can bleed into Subject 1, causing both subjects to share or partially share the same vocal identity, or Subject 1 uses a completely different voice that wasn't prompted.
Update: I originally reported this as a Ref2VA-specific issue, but I have now reproduced the same multi-speaker voice bleed using FL2VA from a single composed first frame with no last frame.
In FL2VA, explicitly conditioning Subject 2 to speak in an American voice causes Subject 1’s voice to become American as well. Conversely, explicitly conditioning Subject 1 to speak in an Elizabeth Swann / Keira Knightley-style British voice restores Subject 1’s voice, but Subject 2 also shifts to a British voice.
The issue also persists with only one speaker transition (A → B), with explicit tags, and with pure <Subject 1> / <Subject 2> speaker labels.
This suggests the issue may not be specific to Ref2VA reference routing. It may instead involve multi-speaker voice conditioning being applied too globally across the audio stream rather than remaining reliably bound to an individual speaker.
For comparison, other two-character combinations can work correctly, so this does not appear to be a universal inability to generate two distinct voices.