Repository navigation
Replies: 1 comment 1 reply
|
Hey, just to say that the Multilingual LibriSpeech dataset would do the trick. It's also from LibriVox and has 45k hours of transcribed speech. |
1 reply
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
I have read through the paper and found out they are using 60k unlabeled audio data to train the AI.
Anyone has any direction we can reproduce the method in the paper that using speech recognition to generate the transcription and increase the accuracy of it?
All reactions