Testing whether related, higher-resource languages can make speech translation better for languages with little data.
I’m an undergraduate researcher in WAVLab (Watanabe’s Audio and Voice Lab) at Carnegie Mellon’s Language Technologies Institute, advised by Shinji Watanabe. The lab works on speech recognition, speech translation, spoken language understanding and machine learning for speech.
Most of the world’s languages have very little recorded, transcribed speech, so speech translation for them is poor. My question: can a closely related language with lots of data measurably improve translation quality for a low-resource one?