Skip to main navigation Skip to search Skip to main content

DiTVC: One-Shot Voice Conversion via Diffusion Transformer with Environment and Speaking Rate Cloning

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Traditional zero-shot voice conversion methods typically extract a speaker embedding from a reference recording first and then generate the source speech content in the target speaker's voice by conditioning on that embedding. However, this process often overlooks time-dependent speaker characteristics, such as voice dynamics and speaking rates, as well as environmental acoustic properties of the reference recording. To address these limitations, we propose a one-shot voice conversion framework capable of replicating not only voice timbre but also acoustic properties. Our model is built upon Diffusion Transformers (DiT) and conditioned on a designed content representation for acoustic cloning. Besides, we introduce specific augmentations during training to enable accurate speaking rate cloning. Both objective and subjective evaluations demonstrate that our method outperforms existing approaches in terms of audio quality, speaker similarity, and environmental acoustic similarity, while effectively capturing the speaking rate distribution of target speakers. Audio samples are available at: ditvc.github.io.

Original languageEnglish (US)
Title of host publicationProceedings of the 2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331537456
DOIs
StatePublished - 2025
Event2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2025 - Tahoe City, United States
Duration: Oct 12 2025Oct 15 2025

Publication series

NameIEEE Workshop on Applications of Signal Processing to Audio and Acoustics
ISSN (Print)1931-1168
ISSN (Electronic)1947-1629

Conference

Conference2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2025
Country/TerritoryUnited States
CityTahoe City
Period10/12/2510/15/25

All Science Journal Classification (ASJC) codes

  • Computer Science Applications
  • Electrical and Electronic Engineering

Fingerprint

Dive into the research topics of 'DiTVC: One-Shot Voice Conversion via Diffusion Transformer with Environment and Speaking Rate Cloning'. Together they form a unique fingerprint.

Cite this