티스토리 뷰

반응형

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models


  • Multilingual zero-shot text-to-speech model이 필요함
  • OmniVoice
    • Complex 2-stage pipeline 대신 text를 multi-codebook acoustic token으로 direct mapping 하여 performance bottleneck을 완화
    • Full-codebook random masking strategy와 LLM initialization을 활용해 multilingual coverage를 향상
  • 논문 (INTERSPEECH 2026) : Paper Link

1. Introduction

  • Zero-shot Text-to-Speech (TTS) model은 few-second reference audio 만으로도 high-quality speech를 생성할 수 있지만, language limitation으로 인해 low-resource language에서는 사용하기 어려움
    • 특히 최근의 discrete-token-based non-autoregressive (NAR) TTS model은 기존의 autoregressive (AR) 방식과 달리 text-to-semantic/semantic-to-acoustic의 2-stage cascaded pipeline으로 구성됨
    • BUT, 해당 방식은 error propagation과 information bottleneck으로 인해 intelligibility의 한계가 있음

-> 그래서 cascaded pipeline의 complexity를 개선하고 large-scale multilingual TTS를 지원할 수 있는 OmniVoice를 제안

 

  • OmniVoice
    • Bidirectional Transformer 기반의 discrete masked diffusion objective를 사용하여 text를 multi-codebook acoustic token으로 directly mapping
    • Full-codebook random masking과 Large Language Model (LLM) initialization을 통해 training efficiency와 intelligibility를 추가적으로 향상

< Overall of OmniVoice >

  • Bidirectional Transformer backbone과 full-codebook random masking, LLM initialization을 활용한 single-stage multilingual zero-shot NAR TTS model
  • 결과적으로 기존보다 우수한 성능을 달성

2. Method

- Architecture

  • OmniVoice는 diffusion language model-style architecture를 가지는 single-stage NAR TTS model로써 discrete diffusion objective와 bidirectional Transformer backbone을 통해 training 됨
    • 논문은 특히 text를 multi-codebook acoustic token으로 directly mapping 하여 기존 2-stage cascaded pipeline의 error propagation, information bottleneck 문제를 해결함
    • OmniVoice의 input은 다음과 같이 구성됨:
      1. Text token sequence $Y$
        - Linguistic, task-oriented guidance를 제공하는 instruct, transcript token으로 구성된 concatenated sequence
      2. Acoustic token matrix $X$
        - Time step $T$, codebook 수 $C$에 대해 $X\in \mathbb{R}^{T\times C}$로 정의되는 multi-codebook matrix
    • Acoustic matrix $X$는 temporal dimension을 따라 2가지 segment로 partition 됨
      1. Prompt segment $X_{prompt}$는 prefix acoustic content를 contain 하고 target masked segment $X_{target}$은 token이 special mask token $[M]$으로 randomly replace 됨
      2. 결과적으로 OmniVoice는 text condition $Y$, prompt $X_{prompt}$, $X_{target}$의 unmasked token을 활용하여 $X_{target}$의 masked position에 대한 original token을 recover 함
    • Training loss는 masked acoustic token position에서 compute 됨:
      (Eq. 1) $\mathcal{L}=-\sum_{(t,c)\in\mathcal{M}}\log P(x_{t,c}|X,Y;\theta)$
      - $\mathcal{M}$ : target segment 내 masked position에 대한 index $(t,c)$의 set
      - $t\in\{T_{p}+1,T_{p}+2,...,T\}$, $c\in\{1,2,...,C\}$
      - $x_{t,c}$ : time step $t$, codebook index $c$에서의 ground-truth acoustic token, $P(x_{t,c}|...;\theta)$ : $\theta$로 parameterize 된 model-predicted probability distribution

Overview

- Full-Codebook Random Masking

  • 기존 multi-codebook acoustic token prediction은 per-layer masking schedule을 채택해 각 sample 마다 single codebook layer $c$ 내에서 random masking을 수행하고 해당 layer에 대한 loss를 exclusively compute 함
    • BUT, 해당 방식은 각 iteration 마다 token matrix의 sparse subset만 optimize 하므로 suboptimal 함
    • 따라서 논문은 모든 codebook layer에 대해 fully stochastic masking strategy를 적용함
      1. $T\times C$ token matrix의 각 entry 마다 binary mask $m_{i,j}\sim\text{Bernoulli}(p_{t})$를 independently sample 함
        - Masking ratio $p_{t}$는 각 training instance 마다 uniform distriubtion $p_{t}\sim \mathcal{U}(0,1)$에서 drawn 됨
      2. 결과적으로 token의 $50\%$가 loss computation에 사용되고, 이는 per-layer masking strategy에 비해 $C$배 더 많으므로 convergence를 크게 accelerate 할 수 있음

- LLM Initialization

  • NAR TTS model의 intelligibility를 향상하기 위해 model backbone을 pre-trained AR LLM으로 initialize 함
    • 특히 OmniVoice backbone은 standard AR LLM과 structurally identical 하므로 direct weight transfer가 가능함
    • 해당 initialization을 통해 OmniVoice는 LLM의 linguistic capacity를 text-to-speech mapping을 위한 strong prior로 사용할 수 있음

- Multilingual Scaling

  • 논문은 low-resource language를 포함한 600개의 language를 지원하는 것을 목표로 함
    • 이를 위해 다양한 dataset을 aggregate 한 다음, speech resotration model과 rule-based filtering을 활용해 noise, invalid transcript를 exclude 함
      - 결과적으로 581k hours audio로 구성된 dataset을 얻음
    • 추가적으로 low-resource language에 대한 data imbalance를 mitigate 하기 위해 language-level data resampling strategy를 적용함
    • Multilingual text processing 시에는 pre-trained LLM의 sub-word tokenizer를 사용하여 grapheme-to-phoneme conversion과 language-specific text normalization을 eliminate 함

Multilingual Training Dataset

- Multi-Dimensional Controllability

  • 논문은 multilingual coverage 외에도 추가적으로 acoustic, identity, linguistic controllability를 개선함
    • Acoustic Control: Prompt Denoising
      1. Audio prompt는 non-ideal condition에서 record 되는 경우가 많으므로, environmental noise나 reverberation에 의해 compromise 될 수 있음
        - 따라서 논문은 prompt denoising을 도입함
      2. Training 시 data subset에 synthetic noise와 reverberation을 inject 하여 augmentation을 수행하고, 해당 sample을 $\texttt{<|denoise|>}$ instruction token과 함께 pair 함
    • Identity Control: Attribute-Guided Voice Design
      1. Audio prompt 없이도 flexible TTS를 지원하기 위해 speaker-attribute-based voice design을 수행함
      2. Training instruction sequence에 gender, age, pitch 등과 같은 specific speaker attribute를 incorporate 하여 customized voice를 synthesize 함
    • Linguistic Control: Paralinguistics and Phonetics
      1. Intelligibility와 expressiveness 간의 gap을 bridge 하기 위해 labeled dataset을 활용하여 paralinguistic control을 incorporate 함
      2. 추가적으로 explicit phonetic override capability를 위해 hybrid text input format을 도입함

3. Experiments

- Settings

- Results

  • 전체적으로 OmniVoice의 성능이 가장 우수함

Model 성능 비교

  • Multilingual benchmark에서도 뛰어난 성능을 보임

Multilingual 성능

  • OmniVoice는 near-ground-truth 수준의 CER을 보임

CER 비교

  • Ablation Study
    • Full-codebook random masking은 성능 향상에 유효함

Masking 효과

  • LLM initialization을 사용하면 더 나은 성능을 달성할 수 있음

Initialization 효과

 

반응형
댓글
최근에 올라온 글
최근에 달린 댓글
«   2026/10   »
일 월 화 수 목 금 토
1 2 3
4 5 6 7 8 9 10
11 12 13 14 15 16 17
18 19 20 21 22 23 24
25 26 27 28 29 30 31
Total
Today
Yesterday