TY - GEN
T1 - DiTVC
T2 - 2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2025
AU - Wang, Yunyun
AU - Su, Jiaqi
AU - Finkelstein, Adam
AU - Kumar, Rithesh
AU - Chen, Ke
AU - Jin, Zeyu
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Traditional zero-shot voice conversion methods typically extract a speaker embedding from a reference recording first and then generate the source speech content in the target speaker's voice by conditioning on that embedding. However, this process often overlooks time-dependent speaker characteristics, such as voice dynamics and speaking rates, as well as environmental acoustic properties of the reference recording. To address these limitations, we propose a one-shot voice conversion framework capable of replicating not only voice timbre but also acoustic properties. Our model is built upon Diffusion Transformers (DiT) and conditioned on a designed content representation for acoustic cloning. Besides, we introduce specific augmentations during training to enable accurate speaking rate cloning. Both objective and subjective evaluations demonstrate that our method outperforms existing approaches in terms of audio quality, speaker similarity, and environmental acoustic similarity, while effectively capturing the speaking rate distribution of target speakers. Audio samples are available at: ditvc.github.io.
AB - Traditional zero-shot voice conversion methods typically extract a speaker embedding from a reference recording first and then generate the source speech content in the target speaker's voice by conditioning on that embedding. However, this process often overlooks time-dependent speaker characteristics, such as voice dynamics and speaking rates, as well as environmental acoustic properties of the reference recording. To address these limitations, we propose a one-shot voice conversion framework capable of replicating not only voice timbre but also acoustic properties. Our model is built upon Diffusion Transformers (DiT) and conditioned on a designed content representation for acoustic cloning. Besides, we introduce specific augmentations during training to enable accurate speaking rate cloning. Both objective and subjective evaluations demonstrate that our method outperforms existing approaches in terms of audio quality, speaker similarity, and environmental acoustic similarity, while effectively capturing the speaking rate distribution of target speakers. Audio samples are available at: ditvc.github.io.
UR - https://www.scopus.com/pages/publications/105026955205
UR - https://www.scopus.com/pages/publications/105026955205#tab=citedBy
U2 - 10.1109/WASPAA66052.2025.11230986
DO - 10.1109/WASPAA66052.2025.11230986
M3 - Conference contribution
AN - SCOPUS:105026955205
T3 - IEEE Workshop on Applications of Signal Processing to Audio and Acoustics
BT - Proceedings of the 2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 12 October 2025 through 15 October 2025
ER -