티스토리 뷰
Paper/Neural Codec
[Paper 리뷰] CycleCodec: Distillation-Free Factorized Neural Speech Codec via Cycle-Consistent Speaker Swapping
feVeRin 2026. 10. 7.반응형
CycleCodec: Distillation-Free Factorized Neural Speech Codec via Cycle-Consistent Speaker Swapping
- Factorized neural codec은 pre-trained Self-Supervised Learning encoder나 Automatic Speech Recognition system에 의존하므로 low-resource language에서 활용하기 어려움
- CycleCodec
- Cycle-consistent speaker swapping을 활용해 speaker-conditioned generation의 cross-stream leakage를 방지
- 추가적으로 speaker-content disentanglement를 향상하기 위해 query-based Transformer aggregator와 speaker contrastive loss를 도입
- 논문 (INTERSPEECH 2026) : Paper Link
1. Introduction
- Neural speech codec은 general-purpose tokenization 외에도 codec representation을 time-varying/invariant stream으로 factorizing 하는 데 사용할 수 있음
- BUT, 기존의 factorized neural codec은 pre-trained teacher model에서 distill된 knowledge에 의존함
- 즉, content space가 Automatic Speech Recognition (ASR) objective나 Wav2Vec2.0과 같은 Self-Supervised Learning (SSL) model의 phone-like representation으로 mapping 됨 - 결과적으로 unseen language를 robsutly transfer 할 수 없고 external teacher supervision으로 인한 cross-stream leakage 문제가 나타남
- BUT, 기존의 factorized neural codec은 pre-trained teacher model에서 distill된 knowledge에 의존함
-> 그래서 external distillation을 제거한 factorized neural codec인 CycleCodec을 제안
- CycleCodec
- Cycle-consistent speaker swapping을 통해 codec-internal self-supervision signal을 제공하여 pre-trained teacher model에 대한 의존성을 제거
- Query-based Transformer aggregator와 cycle-consistency reconstruction constraint를 활용해 speaker-content disentanglement를 향상
< Overall of CycleCodec >
- Cycle-consistent speaker swapping을 활용한 distillation-free factorized neural speech codec
- 결과적으로 기존보다 우수한 성능을 달성
2. Method
- CycleCodec은 distillation-free factorized neural speech codec으로써 각 utterance를 2개의 complementary stream으로 decompose 함
- 즉, Time-varying discrete sequence는 content, prosody를 capture 하고 Time-invariant continuous global embedding은 speaker trait을 capture 함
- 구조적으로는 TiCodec을 기반으로 encoder-decoder backbone, time-varying quantization pipeline, global feature injection을 사용함
- Utterance $\mathbf{y}$가 주어지면 encoder는 frame-level feature $\mathbf{z}\in\mathbb{R}^{T\times d_{z}}$와 global embedding $\mathbf{g}\in\mathbb{R}^{d_{g}}$를 생성함
- 이때 time-varying stream은 discrete feature sequence $\mathbf{q}\in\mathbb{R}^{T\times d_{q}}$로 further quantize 되고, decoder는 $\mathbf{g}$에 condition 된 quantized feature $\mathbf{q}$로부터 speech $\hat{\mathbf{y}}=D(\mathbf{q},\mathbf{g})$를 reconstruct 함
- Architectural Design for Factorization
- Representation interface에서 cross-stream leakage를 줄이기 위해 2가지 supporting architecture를 도입함
- 먼저 quantizer의 codebook size를 줄여 time-varying stream의 information capacity를 constrain 함
- Overly expressive codebook은 residual speaker information까지 encode 할 수 있기 때문 - 추가로 query-based Transformer aggregator를 사용해 global speaker modeling을 strengthen 함
- Speaker cue는 time에 따라 unevenly distribute 되므로 모든 frame에 대해 uniform pooling을 수행하면 weak global speaker representation이 생성될 수 있음
- 반면 query-based Transformer aggregator를 사용하면 long context에 대한 speaker cue를 selectively summeraize 할 수 있음
- $\mathbf{H}\in\mathbb{R}^{T\times d_{h}}$를 global speaker modeling을 위해 encoder에서 추출된 intermediate frame-level feature라고 하자
- 여기서 논문은 $N$ learnable query token $\mathbf{Q}\in\mathbb{R}^{N\times d_{h}}$를 temporal dimension을 따라 $\mathbf{H}$와 concatenate 하여 $N+T$ length의 sequence를 구성함
- 해당 sequence는 $L$-stacked self-attention block을 가지는 Transformer encoder로 process 됨 - 이후 query position의 output만 speaker summary token으로 취급하여 flatten 한 다음, lightweight projection head를 적용해 utternace-level embedding $\mathbf{g}$를 얻음
- Lightweight projection head는 single linear layer와 LeakyReLU로 구성됨
- 여기서 논문은 $N$ learnable query token $\mathbf{Q}\in\mathbb{R}^{N\times d_{h}}$를 temporal dimension을 따라 $\mathbf{H}$와 concatenate 하여 $N+T$ length의 sequence를 구성함
- 추가적으로 segment-level consistency obejctive 만으로는 speaker-discriminative global embedding을 보장하기 어려우므로 auxiliary regularizer로 standard contrastive loss를 도입함:
(Eq. 1) $\mathcal{L}_{spk}=-\frac{1}{B}\sum_{i=1}^{B} \frac{1}{|P(i)|} \sum_{p\in P(i)}\log \frac{\exp(\text{sim}(\mathbf{g}_{i},\mathbf{g}_{p})/\tau)}{\sum_{a\neq i}\exp( \text{sim}(\mathbf{g}_{i},\mathbf{g}_{a})/\tau)}$
- $\text{sim}(\cdot, \cdot)$ : $l_{2}$-normalized embedding 간의 cosine-similarity, $\tau$ : temperature, $P(i)$ : anchor $i$와 동일한 speaker label을 가지는 positives
- 먼저 quantizer의 codebook size를 줄여 time-varying stream의 information capacity를 constrain 함

- Cycle-Consistent Speaker Swapping
- Speaker-conditioned generation에서 content preservation을 보장하기 위해 cycle-consistent speaker swapping을 도입함
- Cycle-consistent speaker swapping은 external supervision 없이 codec 자체로 $\mathbf{q},\mathbf{g}$를 verify 함
- Cycle procedure는 두 개의 forward pass로 구성됨
- 먼저 batch 내에서 각 source utterance $\mathbf{y}_{src}$가 다른 speaker의 target $\mathbf{y}_{tgt}$와 match 되도록 source-target pair를 구성함
- Swap forward pass에서는 shared encoder가 $\mathbf{y}_{src}$에서 source time-varying feature $\mathbf{q}_{src}$를, $\mathbf{y}_{tgt}$에서 target speaker embedding $\mathbf{g}_{tgt}$을 추출함
- 이후 decoder를 source time-varying feature와 target speaker embedding에 condition 하여 swapped utterance $\mathbf{y}_{swap}=D(\mathbf{q}_{src},\mathbf{g}_{tgt})$을 생성함
- 이때 $\mathbf{y}_{swap}$은 $\mathbf{q}_{src}$의 content, prosody를 preserve 하면서 $\mathbf{g}_{tgt}$의 speaker identity를 가져야 함
- Self-supervised consistency signal을 제공하기 위해 $\mathbf{y}_{swap}$을 re-analysis한 다음, swap-back decoding을 통해 source utterance를 reconstruct하는 cycle forward pass를 도입함
- 이를 위해 $\mathbf{y}_{swap}$을 shared encoder, quantizer에 다시 pass 하여 $\mathbf{q}_{swap},\mathbf{g}_{swap}$을 얻은 다음, 두 가지 complementary constraint를 impose 함:
(Eq. 2) $\mathcal{L}_{swap}^{g}=1-\cos(\mathbf{g}_{swap}, \mathbf{g}_{tgt})$
(Eq. 3) $\mathcal{L}_{swap}^{q}=\left|\left| \mathbf{q}_{swap}-\mathbf{q}_{src}\right|\right|_{2}^{2}$
- (Eq. 2)는 swapped utterance가 global space에서 target speaker와 match 되도록 함
- (Eq. 3)는 speaker condition이 swapping 되어도 recovered time-varying feature가 unchange 되도록 함 - 해당 loss를 통해 $\mathbf{g}$가 speaker identity를 control 하고 $\mathbf{q}$가 time-varying dynamics를 preserve 하도록 유도하여 cross-stream leakage를 방지할 수 있음
- 이를 위해 $\mathbf{y}_{swap}$을 shared encoder, quantizer에 다시 pass 하여 $\mathbf{q}_{swap},\mathbf{g}_{swap}$을 얻은 다음, 두 가지 complementary constraint를 impose 함:
- 한편 $(\mathbf{q},\mathbf{g})$에 대한 feature-level constraint 만으로는 waveform-level content distortion을 방지할 수 없음
- 따라서 논문은 cycle forward pass에서 swap-back reconstruction을 포함하여 original source utterance를 re-synthesize 하고 content drift에 대한 direct signal을 제공함
- 즉, recovered time-varying feature를 source speaker embedding과 combine 하여 cycle utterance $\mathbf{y}_{cycle}=D(\mathbf{q}_{swap},\mathbf{g}_{src})$를 decode 하고 mel-spectrogram reconstruction loss를 적용함:
(Eq. 4) $\mathcal{L}_{cycle}^{mel}=\left|\left| \text{Mel}(\mathbf{y}_{cycle})-\text{Mel}(\mathbf{y}_{src})\right|\right|_{1}$
- Objective and Optimization
- 논문은 2-stage training strategy를 채택함
- First stage에서는 original codec objective $\mathcal{L}_{codec}$과 speaker-discriminative contrastive term $\mathcal{L}_{spk}$를 train 하여 time-varying/global stream 간의 factorization을 지원함
- Second stage에서는 cycle-consistent speaker swapping을 codec-internal self-supervision signal로 사용하여 speaker-conditioned generation 시 content preservation을 보장함
- Encoder, quantizer는 frozen 됨 - 결과적으로 fine-tuning objective는:
(Eq. 5) $\mathcal{L}=\mathcal{L}_{codec}+\lambda_{spk}\mathcal{L}_{spk}+\lambda_{q}\mathcal{L}_{swap}^{q}+\lambda_{g}\mathcal{L}_{swap}^{g}+\lambda_{m}\mathcal{L}_{cycle}^{mel}$
- $\lambda_{spk}, \lambda_{q},\lambda_{g},\lambda_{m}$ : hyperparameter
3. Experiments
- Settings
- Results
- 전체적으로 CycleCodec의 성능이 가장 우수함

- Zero-shot VC 측면에서도 CycleCodec이 가장 우수한 성능을 보임

- CycleCodec은 strong content preservation이 가능함

반응형
'Paper > Neural Codec' 카테고리의 다른 글
댓글
