티스토리 뷰

반응형

HoliTok: A Continuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding


  • Unified speech modeling을 위해서는 holistic tokenization space가 필요함
  • HoliTok
    • Signal-level fidelity를 preserve 하면서 semantic information을 incorporate 하는 progressive training을 도입
    • Unified AR+DiT model을 활용해 latent sequence가 generation-specific, unified generation-understanding task를 모두 지원하도록 구성
  • 논문 (EMNLP 2026) : Paper Link

1. Introduction

  • Multi-modal foundation model은 unified understanding과 generation을 요구함
    • 특히 speech task에서는 tokenizer를 통해 simultaneously decodable, learnable, informative 한 continuous space를 reprsent 할 수 있어야 함
    • EnCodec, WavTokenizer와 같은 discrete codec-based tokenizer를 활용하면 language-model-friendly symbol을 얻을 수 있지만 multi-codebook design으로 인한 information loss와 complexity 문제가 발생함
      - Continuous tokenizer는 quantization loss를 줄일 수 있지만 understanding model에는 적합하지 않음
    • 기존의 unified representation 역시 task-specific system에서만 evaluate되므로 shared modeling의 한계가 있음

-> 그래서 unified speech generation, understanding modeling을 지원하는 tokenizer인 HoliTok을 제안

 

  • HoliTok
    • Progressive training을 활용하여 learnable, semantically informative latent space를 구성
    • High-level feature distillation과 audio-language supervision을 반영해 variational regularization을 strengthen하고 latent space를 refine
    • 추가적으로 autoregressive (AR) + DiT architecture에 기반한 unified generation-understanding model을 구축하여 unified speech language modeling을 지원

< Overall of HoliTok >

  • Progressive training과 audio-language supervision을 활용한 unified speech tokenizer
  • 결과적으로 기존보다 우수한 성능을 달성

2. Method

- Main Architecture

  • Encoder
    • Encoder는 구조적으로 1-dimensional convolutional projection 뒤에 6개의 strided causal convolutional downsampling block이 이어짐
      1. 이때 channel width는 각 block마다 double 되어 $12$에서 $768$까지 증가하고, kernel size는 $[4,4,4,8,12,20]$, downsampling rate는 $[2,2,2,4,6,10]$을 사용함
      2. 결과적으로 total hop size는 $1920$이 되어 48kHz audio에 대해 25Hz latent sequence를 생성함
    • 각 downsampling block 다음에는 dilated causal convolution으로 구성된 residual stack이 추가되고, final encoder projection은 hidden sequence를 128-dimensional acoustic representation으로 mapping 함
  • Temporal Variational Bottleneck
    • Convolutional encoder에 4-layer LSTM block과 projection-in/out layer로 구성된 bottleneck layer를 추가하고, 이후 $1\times 1$ convolution을 통해 diagonal Gaussian posterior의 mean/log-scale을 predict 함
      1. 이때 latent sequence는 reparameterization trick으로 sampling 됨
      2. 추가적으로 standard Normal prior에 대한 KL regularization 계산 시, normalizing flow를 적용하여 latent distribution의 expressiveness를 향상함
    • 결과적으로 sampled latent sequence는 model dimension으로 project back 되고 decoding 전에 encoder-side bottleneck의 mirrored structure로 process 됨
  • Decoder
    • Decoder는 BigVGAN-style generator를 활용하여 25Hz latent sequence로부터 48kHz waveform을 reconstruct 함
      - Upsampling module은 encoder downsampling structure를 mirror 하고 BigVGAN을 따라 SnakeBeta activation을 가진 AMPBlock으로 refine 함
    • Final projection은 hidden feature를 single-channel wavevform으로 mapping 함
  • Supervision Network
    • Supervision network는 encoder-decoder design을 따라 0.6B Transformer encoder와 pre-trained Qwen2.5-0.5B decoder로 구성됨
    • Encoder는 latent sample을 process 하고, 해당 sample은 task-label과 concatenate 되어 language model-decoder로 전달됨
      - 해당 decoder는 Stage 3 supervision에서만 사용되고 이후에는 discard 됨

Overview

- Stage 1 & 2: Progressive Training of High-Fidelity Variational Latent Space

  • VAE training에서 strong KL constraint를 impose 하면 structured latent distribution을 promote 할 수 있지만, decoder가 high-fidelity reconstruction manifold를 학습하기 전에 representation에서 acoustic detail이 discard 됨
    • 따라서 논문은 해당 fidelity loss를 mitigate 하기 위해 HoliTok latent space를 progressively shaping 함
      1. Stage 1에서는 deterministic autoencoder를 training 해 high-fidelity acoustic autoencoding space를 구축함
      2. Stage 2에서는 pre-trained encoder/decoder를 freeze 하고 weak KL regularization을 적용해 temporal variational bottleneck만 training 하여 autoencoding space를 stochastic latent space로 convert 함
    • 해당 progressive training을 통해 latent trajectory를 reliable decoding region과 가깝게 유지할 수 있음
  • Stage 1: Reconstruction-Oriented Autoencoder Pre-training
    • Encoder $E_{\phi}$는 input waveform $\mathbf{x}$를 low-rate acoustic representation $\mathbf{z}_{AE}=E_{\phi}(\mathbf{x})$로 mapping 하고, decoder $G_{\psi}$는 $\hat{\mathbf{x}}_{AE}=G_{\psi}(\mathbf{z}_{AE})$와 같이 waveform을 reconstruct 함
    • 이때 Stage 1은 reconstruction-oriented generator objective로 training 됨:
      (Eq. 1) $ \mathcal{L}_{I}=\mathbb{E}_{\mathbf{x}}\left[\ell_{gen}\left( \mathbf{x},G_{\psi}\left( E_{\phi}(\mathbf{x})\right)\right)\right]$
    • $\ell_{gen}$은 generator-side waveform generation loss로써, multi-scale spectral reconstruction, adversarial supervision, discriminator feature matching loss를 combine 하여 얻어짐:
      (Eq. 2) $\ell_{gen} = \lambda_{spec}\mathcal{L}_{spec}+\lambda_{adv}\mathcal{L}_{adv}^{G}+\lambda_{fm}\mathcal{L}_{fm}$
      - $\mathcal{L}_{spec}$ : multi-scale mel-spectral reconstruction loss, $\mathcal{L}_{adv}^{G}$ : generator-side adversarial loss, $\mathcal{L}_{fm}$ : feature matching loss
    • 결과적으로 Stage 1은 variational regularization 이전에 high-fidelity reconstruction manifold를 구성함
  • Stage 2: Autoencoding-to-Variational Latent Transfer
    • Stage 2는 pre-trained autoencoder에서 encoder $E_{\phi}$와 decoder $G_{\psi}$를 freeze 하고 temporal variational bottleneck만 training 함
      1. Deterministic acoustic representation $\mathbf{z}_{AE}=E_{\phi}(\mathbf{x})$가 주어지면, bottleneck은 stochastic latent에 대한 posterior $q_{\eta}(\mathbf{z}_{VAE}|\mathbf{z}_{AE})$를 define 함
        - 이는 reparameterization trick으로 sampling 되고 frozen decoder로 decode 됨
      2. 이때 논문은 reconstruction-dominated VAE objective를 optimize 함:
        (Eq. 3) $\mathcal{L}_{II}=\mathbb{E}_{\mathbf{x}}\left[\mathbb{E}_{\mathbf{z}_{VAE} \sim q_{\eta}(\cdot |\mathbf{z}_{VAE})}\left[ \ell_{gen}\left(\mathbf{x},G_{\psi}(\mathbf{z}_{VAE})\right)\right]+ \beta_{low}D_{KL}\left( q_{\eta}\left( \mathbf{z}_{VAE}|\mathbf{z}_{AE}\right) ||p(\mathbf{z})\right)\right]$
        - $p(\mathbf{z})=\mathcal{N}(0,I)$
    • Small KL weight는 bottleneck이 reconstruction-critical acoustic detail을 discard 하는 것을 방지하면서 distributional regularity를 encourage 함
    • $E_{\phi}, G_{\psi}$가 fix 되어 있으므로 Stage 2는 deterministic autoencoding space를 variational latent space로 transfer 하면서 sampled latent를 decoder의 high-fidelity reconstruction region과 close 할 수 있음
  • Implicit Fidelity Transfer
    • Progressive Stage 1, 2 design은 implict fidelity-transfer effect를 제공함
    • Frozen pre-trained decoder와 reconstruction-dominated objective는 Stage 2 variational sample이 high-fidelity autoencoding manifold 주변에 stay 하도록 constrain 함
    • Expected waveform distortion은 Stage 1 autoencoder distortion과 AE-to-VAE latent shift에 의해서만 control 되므로, pre-trained decoder를 freeze 하고 small KL weight로만 temporal variational bottleneck을 training 할 수 있음

- Stage 3: Downstream-aware Enrichment of the Tokenization Space

  • Stage 3에서는 VAE latent space를 pre-trained speech representation과 task-conditioned supervision으로 further enrich 하여 tokenization space가 downstream speech language modeling에 informative 하도록 함
    - 이때 full VAE posterior $q_{\theta}(\mathbf{z}|\mathbf{x})=q_{\eta '}(\mathbf{z}|E_{\phi '}(\mathbf{z}))$는 Stage 2 posterior $q_{\eta}(\mathbf{z}_{VAE}| E_{\phi}(\mathbf{x}))$와 동일한 bottleneck architecture를 inherit 하여 initialize 됨
  • Multi-Granularity Representation Distillation
    • 논문은 VAE latent space를 enrich 하기 위해 multi-granularity representation distillation을 도입함
    • $\mathbf{z}\sim q_{\theta}(\mathbf{z}|\mathbf{x})$가 주어지면, latent sequence를 frame, utterance level의 pre-trained speech model에서 얻어진 frozen teacher reprsentation 모두 대해 align 함
      1. Frame-level distillation의 경우, WavLM을 contextual teacher로 사용하고 prediction head를 적용해 latent sequence를 23rd-layer hidden representation으로 mapping 함
        - Frame rate가 다른 경우 temporal interpolation을 사용함
      2. Utternace-level distillation의 경우, latent sequence를 utterance-level representation으로 aggregate 한 다음 X-vector speaker embedding에 대해 align 함
    • 그러면 unified distillation objective는:
      (Eq. 4) $\mathcal{L}_{distill}=\sum_{r\in\mathcal{R}}\lambda_{r}\left[ 1-\cos\left( H_{r}\left( A_{r}(\mathbf{z})\right), \text{sg}\left(F_{r}(\mathbf{z})\right)\right)\right]$
      - $\mathcal{R}$ : teacher representation set, $F_{r}$ : frozen teacher, $A_{r}$ : frame-/utterance-level teacher에 대한 temporal alignment, $H_{r}$ : adapted latent representation에 대한 teacher space mapping, $\text{sg}(\cdot)$ : stop-gradient
      - Frame-level teacher에서 $\cos$ term은 temporal alignment 이후에 compute 되고 time에 따라 average 됨 
  • Multi-task Language-Modeling Supervision
    • Task-conditioned supervision network를 활용해 latent representation을 downstream supervision에 expose 함
    • Task type $\tau \in\mathcal{T}$와 해당 target output $\mathbf{y}^{\tau}$가 주어졌을 때, 다음의 unified language modeling objective를 optimize 함:
      (Eq. 5) $\mathcal{L}_{sup}=-\mathbb{E}_{(\mathbf{x},\tau,\mathbf{y}^{\tau})}\mathbb{E}_{\mathbf{z}\sim q_{\theta}(\cdot |\mathbf{x})}\left[\log p_{\omega}(\mathbf{y}^{\tau}|\mathbf{z},\tau)\right]$
      - 이를 통해 latent space는 waveform reconstruction에는 불필요하지만 speech/audio understanding에는 필요한 information을 retain 할 수 있음
    • Waveform reconstruction, variational regularization, representation distillation, downstream supervision을 combine 하여 얻어지는 Stage 3 objective는:
      (Eq. 6) $\mathcal{L}_{III}=\mathcal{L}_{gen}+\beta_{high}\mathcal{L}_{KL}+\mathcal{L}_{distill}+\lambda_{sup}\mathcal{L}_{sup}$
      - $\mathcal{L}_{gen}$ : generator-side waveform generation loss, $\mathcal{L}_{KL}$ : VAE regularization
  • Variational Interpretation
    • Stage 3는 downstream-aware variational surrogate를 optimizing 하는 것으로 볼 수 있음
    • $\mathbf{u}_{r}=F_{r}(\mathbf{x})$를 frozen teacher representation이라 하면, latent variable $\mathbf{z}$는 waveform, teacher representation, task target을 jointly explaining 함:
      (Eq. 7) $p\left(\mathbf{x},\{\mathbf{u}_{r}\}_{r\in\mathcal{R}},\mathbf{y}^{\tau}\right) =\int p(\mathbf{z})p_{\psi}(\mathbf{x}|\mathbf{z})p_{\omega}(\mathbf{y}^{\tau}|\mathbf{z},\tau) \prod_{r\in\mathcal{R}}p_{r}(\mathbf{u}_{r}|\mathbf{z})d\mathbf{z}$
    • Variational posterior $q_{\theta}(\mathbf{z}|\mathbf{x})$에 대해, weighted ELBO-style objective를 얻을 수 있음:
      (Eq. 8) $\mathcal{I}_{III}=\mathbb{E}_{\mathbf{z}\sim q_{\theta}(\cdot |\mathbf{x})}\left[ \log p_{\psi}(\mathbf{x}|\mathbf{z})+\lambda_{sup}\log p_{\omega}(\mathbf{y}^{\tau}|\mathbf{z},\tau) + \sum_{r\in\mathcal{R}}\lambda_{r}\log p_{r}(\mathbf{u}_{r}|\mathbf{z})\right] - \beta_{high}D_{KL}\left( q_{\theta}(\mathbf{z}|\mathbf{x})||p(\mathbf{z})\right)$
      - 결과적으로 $\mathcal{L}_{III}$를 minimize 하는 것은 practical waveform, distillation, supervision, KL term에 대한 surrogate를 maximizing 하는 것과 같음

- Downstream Unified Spoken Language Modeling

  • 논문은 DiTAR를 따라 AR+DiT design을 활용하여 speech understanding, speech generation을 모두 지원하는 downstream spoken language model을 구축함
    • Autoregressive (AR) language model은 mixed text-audio embedding sequence를 process 하고 DiT-based flow-matching module은 speech generation을 위한 continuous latent patch를 predict 함
    • Speech Understanding Objective
      1. Audio latent patch를 $\mathbf{z}_{audio}$, 해당 language-model embedding을 $\mathbf{e}_{audio}$라고 하자
      2. Textual context $\mathbf{c}$, target text $\mathbf{y}_{text}$가 주어지면, autoregressive Cross-Entropy objective를 optimize 할 수 있음:
        (Eq. 9) $\mathcal{L}_{understand}=-\sum_{j}\log p_{\theta}\left(y_{j}|\mathbf{y}_{<j},\mathbf{e}_{audio}, \mathbf{c}\right)$
    • Speech Generation Objective
      1. Language model은 available text와 audio history를 causal hidden state로 summarize 하고 DiT flow-matching module은 historical latent와 previous hidden을 condition으로 각 future latent를 predict 함
      2. 이때 conditional generation process는:
        (Eq. 10) $p_{\theta}\left(\mathbf{z}_{1:K}|\mathbf{c}\right)=\prod_{k=1}^{K}p_{\theta}\left(\mathbf{z}_{k} |\mathbf{h}_{\leq k},\mathbf{z}_{<k}\right)$
        - $\mathbf{h}_{\leq k}$ : $k$ patch prediction을 위한 causal language-model hidden state sequence, $\mathbf{z}_{<k}$ : previously generated audio latent patch
      3. 각 conditional patch distribution은 flow-matching objective로 학습됨:
        (Eq. 11) $\mathcal{L}_{FM}=\mathbb{E}_{k,t}\left[\left|\left| v_{\theta}\left( \mathbf{z}_{k,t}, t|\mathbf{h}_{\leq k},\mathbf{z}_{<k}\right)-\mathbf{u}_{k,t}\right|\right|_{2}^{2}\right]$
        - $\mathbf{z}_{k,t}$ : timestamp $t$의 $k$-th latent patch에 대한 interpolated noisy state, $\mathbf{u}_{k,t}$ : 해당 target velocity
      4. 추가적으로 binary Cross-Entropy EOS loss를 사용하여 audio termination을 supervise 함:
        (Eq. 12) $\mathcal{L}_{generate}=\mathcal{L}_{FM}+\lambda_{eos}\mathcal{L}_{eos}$
        - Generated latent patch는 latent sequence로 assemble 된 후, frozen decoder를 통해 waveform으로 decode 됨

3. Experiments

- Settings

  • Dataset : AISHELL-3, HiFi-TTS, HiFi-TTS2, VCTK, AudioSet, VGGSound, VocalSound, FSD50k, MusicCaps, WavCaps
  • Comparisons : Semantic-VAE, MingTok

- Results

  • 전체적으로 HoliTok의 성능이 가장 뛰어남

Reconstruction 성능

  • Speech synthesis 측면에서도 우수한 성능을 보임

Speech Synthesis 성능

  • HoliTok은 zero-shot TTS에서도 효과적임

Zero-Shot TTS 성능

  • Unified spoken language modeling에서도 뛰어난 성능을 달성함

Unified Spoken Language Modeling

 

반응형
댓글
최근에 올라온 글
최근에 달린 댓글
«   2026/10   »
일 월 화 수 목 금 토
1 2 3
4 5 6 7 8 9 10
11 12 13 14 15 16 17
18 19 20 21 22 23 24
25 26 27 28 29 30 31
Total
Today
Yesterday