UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units

Hirofumi Inaguma; Sravya Popuri; Ilia Kulikov; Peng-Jen Chen; Changhan Wang; Yu-An Chung; Yun Tang; Ann Lee; Shinji Watanabe; Juan Pino

doi:10.18653/v1/2023.acl-long.872

UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units

Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, Juan Pino

Abstract

Direct speech-to-speech translation (S2ST), in which all components can be optimized jointly, is advantageous over cascaded approaches to achieve fast inference with a simplified pipeline. We present a novel two-pass direct S2ST architecture, UnitY, which first generates textual representations and predicts discrete acoustic units subsequently. We enhance the model performance by subword prediction in the first-pass decoder, advanced two-pass decoder architecture design and search strategy, and better training regularization. To leverage large amounts of unlabeled text data, we pre-train the first-pass text decoder based on the self-supervised denoising auto-encoding task. Experimental evaluations on benchmark datasets at various data scales demonstrate that UnitY outperforms a single-pass speech-to-unit translation model by 2.5-4.2 ASR-BLEU with 2.83x decoding speed-up. We show that the proposed methods boost the performance even when predicting spectrogram in the second pass. However, predicting discrete units achieves 2.51x decoding speed-up compared to that case.

Anthology ID:: 2023.acl-long.872
Volume:: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2023
Address:: Toronto, Canada
Editors:: Anna Rogers, Jordan Boyd-Graber, Naoaki Okazaki
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 15655–15680
Language:
URL:: https://aclanthology.org/2023.acl-long.872
DOI:: 10.18653/v1/2023.acl-long.872
Bibkey:
Cite (ACL):: Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, and Juan Pino. 2023. UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15655–15680, Toronto, Canada. Association for Computational Linguistics.
Cite (Informal):: UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units (Inaguma et al., ACL 2023)
Copy Citation:
PDF:: https://aclanthology.org/2023.acl-long.872.pdf

PDF Cite Search