티스토리 뷰
Paper/TTS
[Paper 리뷰] OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
feVeRin 2026. 10. 8.반응형
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
- Multilingual zero-shot text-to-speech model이 필요함
- OmniVoice
- Complex 2-stage pipeline 대신 text를 multi-codebook acoustic token으로 direct mapping 하여 performance bottleneck을 완화
- Full-codebook random masking strategy와 LLM initialization을 활용해 multilingual coverage를 향상
- 논문 (INTERSPEECH 2026) : Paper Link
1. Introduction
- Zero-shot Text-to-Speech (TTS) model은 few-second reference audio 만으로도 high-quality speech를 생성할 수 있지만, language limitation으로 인해 low-resource language에서는 사용하기 어려움
- 특히 최근의 discrete-token-based non-autoregressive (NAR) TTS model은 기존의 autoregressive (AR) 방식과 달리 text-to-semantic/semantic-to-acoustic의 2-stage cascaded pipeline으로 구성됨
- BUT, 해당 방식은 error propagation과 information bottleneck으로 인해 intelligibility의 한계가 있음
-> 그래서 cascaded pipeline의 complexity를 개선하고 large-scale multilingual TTS를 지원할 수 있는 OmniVoice를 제안
- OmniVoice
- Bidirectional Transformer 기반의 discrete masked diffusion objective를 사용하여 text를 multi-codebook acoustic token으로 directly mapping
- Full-codebook random masking과 Large Language Model (LLM) initialization을 통해 training efficiency와 intelligibility를 추가적으로 향상
< Overall of OmniVoice >
- Bidirectional Transformer backbone과 full-codebook random masking, LLM initialization을 활용한 single-stage multilingual zero-shot NAR TTS model
- 결과적으로 기존보다 우수한 성능을 달성
2. Method
- Architecture
- OmniVoice는 diffusion language model-style architecture를 가지는 single-stage NAR TTS model로써 discrete diffusion objective와 bidirectional Transformer backbone을 통해 training 됨
- 논문은 특히 text를 multi-codebook acoustic token으로 directly mapping 하여 기존 2-stage cascaded pipeline의 error propagation, information bottleneck 문제를 해결함
- OmniVoice의 input은 다음과 같이 구성됨:
- Text token sequence $Y$
- Linguistic, task-oriented guidance를 제공하는 instruct, transcript token으로 구성된 concatenated sequence - Acoustic token matrix $X$
- Time step $T$, codebook 수 $C$에 대해 $X\in \mathbb{R}^{T\times C}$로 정의되는 multi-codebook matrix
- Text token sequence $Y$
- Acoustic matrix $X$는 temporal dimension을 따라 2가지 segment로 partition 됨
- Prompt segment $X_{prompt}$는 prefix acoustic content를 contain 하고 target masked segment $X_{target}$은 token이 special mask token $[M]$으로 randomly replace 됨
- 결과적으로 OmniVoice는 text condition $Y$, prompt $X_{prompt}$, $X_{target}$의 unmasked token을 활용하여 $X_{target}$의 masked position에 대한 original token을 recover 함
- Training loss는 masked acoustic token position에서 compute 됨:
(Eq. 1) $\mathcal{L}=-\sum_{(t,c)\in\mathcal{M}}\log P(x_{t,c}|X,Y;\theta)$
- $\mathcal{M}$ : target segment 내 masked position에 대한 index $(t,c)$의 set
- $t\in\{T_{p}+1,T_{p}+2,...,T\}$, $c\in\{1,2,...,C\}$
- $x_{t,c}$ : time step $t$, codebook index $c$에서의 ground-truth acoustic token, $P(x_{t,c}|...;\theta)$ : $\theta$로 parameterize 된 model-predicted probability distribution

- Full-Codebook Random Masking
- 기존 multi-codebook acoustic token prediction은 per-layer masking schedule을 채택해 각 sample 마다 single codebook layer $c$ 내에서 random masking을 수행하고 해당 layer에 대한 loss를 exclusively compute 함
- BUT, 해당 방식은 각 iteration 마다 token matrix의 sparse subset만 optimize 하므로 suboptimal 함
- 따라서 논문은 모든 codebook layer에 대해 fully stochastic masking strategy를 적용함
- $T\times C$ token matrix의 각 entry 마다 binary mask $m_{i,j}\sim\text{Bernoulli}(p_{t})$를 independently sample 함
- Masking ratio $p_{t}$는 각 training instance 마다 uniform distriubtion $p_{t}\sim \mathcal{U}(0,1)$에서 drawn 됨 - 결과적으로 token의 $50\%$가 loss computation에 사용되고, 이는 per-layer masking strategy에 비해 $C$배 더 많으므로 convergence를 크게 accelerate 할 수 있음
- $T\times C$ token matrix의 각 entry 마다 binary mask $m_{i,j}\sim\text{Bernoulli}(p_{t})$를 independently sample 함
- LLM Initialization
- NAR TTS model의 intelligibility를 향상하기 위해 model backbone을 pre-trained AR LLM으로 initialize 함
- 특히 OmniVoice backbone은 standard AR LLM과 structurally identical 하므로 direct weight transfer가 가능함
- 해당 initialization을 통해 OmniVoice는 LLM의 linguistic capacity를 text-to-speech mapping을 위한 strong prior로 사용할 수 있음
- Multilingual Scaling
- 논문은 low-resource language를 포함한 600개의 language를 지원하는 것을 목표로 함
- 이를 위해 다양한 dataset을 aggregate 한 다음, speech resotration model과 rule-based filtering을 활용해 noise, invalid transcript를 exclude 함
- 결과적으로 581k hours audio로 구성된 dataset을 얻음 - 추가적으로 low-resource language에 대한 data imbalance를 mitigate 하기 위해 language-level data resampling strategy를 적용함
- Multilingual text processing 시에는 pre-trained LLM의 sub-word tokenizer를 사용하여 grapheme-to-phoneme conversion과 language-specific text normalization을 eliminate 함
- 이를 위해 다양한 dataset을 aggregate 한 다음, speech resotration model과 rule-based filtering을 활용해 noise, invalid transcript를 exclude 함

- Multi-Dimensional Controllability
- 논문은 multilingual coverage 외에도 추가적으로 acoustic, identity, linguistic controllability를 개선함
- Acoustic Control: Prompt Denoising
- Audio prompt는 non-ideal condition에서 record 되는 경우가 많으므로, environmental noise나 reverberation에 의해 compromise 될 수 있음
- 따라서 논문은 prompt denoising을 도입함 - Training 시 data subset에 synthetic noise와 reverberation을 inject 하여 augmentation을 수행하고, 해당 sample을 $\texttt{<|denoise|>}$ instruction token과 함께 pair 함
- Audio prompt는 non-ideal condition에서 record 되는 경우가 많으므로, environmental noise나 reverberation에 의해 compromise 될 수 있음
- Identity Control: Attribute-Guided Voice Design
- Audio prompt 없이도 flexible TTS를 지원하기 위해 speaker-attribute-based voice design을 수행함
- Training instruction sequence에 gender, age, pitch 등과 같은 specific speaker attribute를 incorporate 하여 customized voice를 synthesize 함
- Linguistic Control: Paralinguistics and Phonetics
- Intelligibility와 expressiveness 간의 gap을 bridge 하기 위해 labeled dataset을 활용하여 paralinguistic control을 incorporate 함
- 추가적으로 explicit phonetic override capability를 위해 hybrid text input format을 도입함
- Acoustic Control: Prompt Denoising
3. Experiments
- Settings
- Dataset : Multilingual dataset
- Comparisons : IndexTTS2, CosyVoice3, F5-TTS, MaskGCT, ZipVoice, VoxCPM, Qwen-TTS
- Results
- 전체적으로 OmniVoice의 성능이 가장 우수함

- Multilingual benchmark에서도 뛰어난 성능을 보임

- OmniVoice는 near-ground-truth 수준의 CER을 보임

- Ablation Study
- Full-codebook random masking은 성능 향상에 유효함

- LLM initialization을 사용하면 더 나은 성능을 달성할 수 있음

반응형
'Paper > TTS' 카테고리의 다른 글
댓글
