티스토리 뷰

반응형

CycleCodec: Distillation-Free Factorized Neural Speech Codec via Cycle-Consistent Speaker Swapping


  • Factorized neural codec은 pre-trained Self-Supervised Learning encoder나 Automatic Speech Recognition system에 의존하므로 low-resource language에서 활용하기 어려움
  • CycleCodec
    • Cycle-consistent speaker swapping을 활용해 speaker-conditioned generation의 cross-stream leakage를 방지
    • 추가적으로 speaker-content disentanglement를 향상하기 위해 query-based Transformer aggregator와 speaker contrastive loss를 도입
  • 논문 (INTERSPEECH 2026) : Paper Link

1. Introduction

  • Neural speech codec은 general-purpose tokenization 외에도 codec representation을 time-varying/invariant stream으로 factorizing 하는 데 사용할 수 있음
    • BUT, 기존의 factorized neural codec은 pre-trained teacher model에서 distill된 knowledge에 의존함
      - 즉, content space가 Automatic Speech Recognition (ASR) objective나 Wav2Vec2.0과 같은 Self-Supervised Learning (SSL) model의 phone-like representation으로 mapping 됨
    • 결과적으로 unseen language를 robsutly transfer 할 수 없고 external teacher supervision으로 인한 cross-stream leakage 문제가 나타남

-> 그래서 external distillation을 제거한 factorized neural codec인 CycleCodec을 제안

 

  • CycleCodec
    • Cycle-consistent speaker swapping을 통해 codec-internal self-supervision signal을 제공하여 pre-trained teacher model에 대한 의존성을 제거
    • Query-based Transformer aggregator와 cycle-consistency reconstruction constraint를 활용해 speaker-content disentanglement를 향상

< Overall of CycleCodec >

  • Cycle-consistent speaker swapping을 활용한 distillation-free factorized neural speech codec
  • 결과적으로 기존보다 우수한 성능을 달성

2. Method

  • CycleCodec은 distillation-free factorized neural speech codec으로써 각 utterance를 2개의 complementary stream으로 decompose 함
    • 즉, Time-varying discrete sequence는 content, prosody를 capture 하고 Time-invariant continuous global embedding은 speaker trait을 capture 함
    • 구조적으로는 TiCodec을 기반으로 encoder-decoder backbone, time-varying quantization pipeline, global feature injection을 사용함
      1. Utterance $\mathbf{y}$가 주어지면 encoder는 frame-level feature $\mathbf{z}\in\mathbb{R}^{T\times d_{z}}$와 global embedding $\mathbf{g}\in\mathbb{R}^{d_{g}}$를 생성함
      2. 이때 time-varying stream은 discrete feature sequence $\mathbf{q}\in\mathbb{R}^{T\times d_{q}}$로 further quantize 되고, decoder는 $\mathbf{g}$에 condition 된 quantized feature $\mathbf{q}$로부터 speech $\hat{\mathbf{y}}=D(\mathbf{q},\mathbf{g})$를 reconstruct 함

- Architectural Design for Factorization

  • Representation interface에서 cross-stream leakage를 줄이기 위해 2가지 supporting architecture를 도입함
    • 먼저 quantizer의 codebook size를 줄여 time-varying stream의 information capacity를 constrain 함
      - Overly expressive codebook은 residual speaker information까지 encode 할 수 있기 때문
    • 추가로 query-based Transformer aggregator를 사용해 global speaker modeling을 strengthen 함
      1. Speaker cue는 time에 따라 unevenly distribute 되므로 모든 frame에 대해 uniform pooling을 수행하면 weak global speaker representation이 생성될 수 있음
      2. 반면 query-based Transformer aggregator를 사용하면 long context에 대한 speaker cue를 selectively summeraize 할 수 있음
    • $\mathbf{H}\in\mathbb{R}^{T\times d_{h}}$를 global speaker modeling을 위해 encoder에서 추출된 intermediate frame-level feature라고 하자
      1. 여기서 논문은 $N$ learnable query token $\mathbf{Q}\in\mathbb{R}^{N\times d_{h}}$를 temporal dimension을 따라 $\mathbf{H}$와 concatenate 하여 $N+T$ length의 sequence를 구성함
        - 해당 sequence는 $L$-stacked self-attention block을 가지는 Transformer encoder로 process 됨
      2. 이후 query position의 output만 speaker summary token으로 취급하여 flatten 한 다음, lightweight projection head를 적용해 utternace-level embedding $\mathbf{g}$를 얻음
        - Lightweight projection head는 single linear layer와 LeakyReLU로 구성됨
    • 추가적으로 segment-level consistency obejctive 만으로는 speaker-discriminative global embedding을 보장하기 어려우므로 auxiliary regularizer로 standard contrastive loss를 도입함:
      (Eq. 1) $\mathcal{L}_{spk}=-\frac{1}{B}\sum_{i=1}^{B} \frac{1}{|P(i)|} \sum_{p\in P(i)}\log \frac{\exp(\text{sim}(\mathbf{g}_{i},\mathbf{g}_{p})/\tau)}{\sum_{a\neq i}\exp( \text{sim}(\mathbf{g}_{i},\mathbf{g}_{a})/\tau)}$
      - $\text{sim}(\cdot, \cdot)$ : $l_{2}$-normalized embedding 간의 cosine-similarity, $\tau$ : temperature, $P(i)$ : anchor $i$와 동일한 speaker label을 가지는 positives

Overview

- Cycle-Consistent Speaker Swapping

  • Speaker-conditioned generation에서 content preservation을 보장하기 위해 cycle-consistent speaker swapping을 도입함
    • Cycle-consistent speaker swapping은 external supervision 없이 codec 자체로 $\mathbf{q},\mathbf{g}$를 verify 함
    • Cycle procedure는 두 개의 forward pass로 구성됨
      1. 먼저 batch 내에서 각 source utterance $\mathbf{y}_{src}$가 다른 speaker의 target $\mathbf{y}_{tgt}$와 match 되도록 source-target pair를 구성함
      2. Swap forward pass에서는 shared encoder가 $\mathbf{y}_{src}$에서 source time-varying feature $\mathbf{q}_{src}$를, $\mathbf{y}_{tgt}$에서 target speaker embedding $\mathbf{g}_{tgt}$을 추출함
      3. 이후 decoder를 source time-varying feature와 target speaker embedding에 condition 하여 swapped utterance $\mathbf{y}_{swap}=D(\mathbf{q}_{src},\mathbf{g}_{tgt})$을 생성함
        - 이때 $\mathbf{y}_{swap}$은 $\mathbf{q}_{src}$의 content, prosody를 preserve 하면서 $\mathbf{g}_{tgt}$의 speaker identity를 가져야 함
    • Self-supervised consistency signal을 제공하기 위해 $\mathbf{y}_{swap}$을 re-analysis한 다음, swap-back decoding을 통해 source utterance를 reconstruct하는 cycle forward pass를 도입함
      1. 이를 위해 $\mathbf{y}_{swap}$을 shared encoder, quantizer에 다시 pass 하여 $\mathbf{q}_{swap},\mathbf{g}_{swap}$을 얻은 다음, 두 가지 complementary constraint를 impose 함:
        (Eq. 2) $\mathcal{L}_{swap}^{g}=1-\cos(\mathbf{g}_{swap}, \mathbf{g}_{tgt})$
        (Eq. 3) $\mathcal{L}_{swap}^{q}=\left|\left| \mathbf{q}_{swap}-\mathbf{q}_{src}\right|\right|_{2}^{2}$
        - (Eq. 2)는 swapped utterance가 global space에서 target speaker와 match 되도록 함
        - (Eq. 3)는 speaker condition이 swapping 되어도 recovered time-varying feature가 unchange 되도록 함
      2. 해당 loss를 통해 $\mathbf{g}$가 speaker identity를 control 하고 $\mathbf{q}$가 time-varying dynamics를 preserve 하도록 유도하여 cross-stream leakage를 방지할 수 있음
    • 한편 $(\mathbf{q},\mathbf{g})$에 대한 feature-level constraint 만으로는 waveform-level content distortion을 방지할 수 없음
      1. 따라서 논문은 cycle forward pass에서 swap-back reconstruction을 포함하여 original source utterance를 re-synthesize 하고 content drift에 대한 direct signal을 제공함
      2. 즉, recovered time-varying feature를 source speaker embedding과 combine 하여 cycle utterance $\mathbf{y}_{cycle}=D(\mathbf{q}_{swap},\mathbf{g}_{src})$를 decode 하고 mel-spectrogram reconstruction loss를 적용함:
        (Eq. 4) $\mathcal{L}_{cycle}^{mel}=\left|\left| \text{Mel}(\mathbf{y}_{cycle})-\text{Mel}(\mathbf{y}_{src})\right|\right|_{1}$

- Objective and Optimization

  • 논문은 2-stage training strategy를 채택함
    • First stage에서는 original codec objective $\mathcal{L}_{codec}$과 speaker-discriminative contrastive term $\mathcal{L}_{spk}$를 train 하여 time-varying/global stream 간의 factorization을 지원함
    • Second stage에서는 cycle-consistent speaker swapping을 codec-internal self-supervision signal로 사용하여 speaker-conditioned generation 시 content preservation을 보장함
      - Encoder, quantizer는 frozen 됨
    • 결과적으로 fine-tuning objective는:
      (Eq. 5) $\mathcal{L}=\mathcal{L}_{codec}+\lambda_{spk}\mathcal{L}_{spk}+\lambda_{q}\mathcal{L}_{swap}^{q}+\lambda_{g}\mathcal{L}_{swap}^{g}+\lambda_{m}\mathcal{L}_{cycle}^{mel}$
      - $\lambda_{spk}, \lambda_{q},\lambda_{g},\lambda_{m}$ : hyperparameter

3. Experiments

- Settings

  • Dataset : LibriTTS, SeedTTS, VieNeuTTS
  • Comparisons : TiCodec, LSCodec

- Results

  • 전체적으로 CycleCodec의 성능이 가장 우수함

Model 성능 비교

  • Zero-shot VC 측면에서도 CycleCodec이 가장 우수한 성능을 보임

Zero-Shot VC에서의 성능

  • CycleCodec은 strong content preservation이 가능함

Iterative Evaluation

 

반응형
댓글
최근에 올라온 글
최근에 달린 댓글
«   2026/10   »
일 월 화 수 목 금 토
1 2 3
4 5 6 7 8 9 10
11 12 13 14 15 16 17
18 19 20 21 22 23 24
25 26 27 28 29 30 31
Total
Today
Yesterday