티스토리 뷰

반응형

TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling


  • Speech token을 활용하면 joint text-speech modeling을 향상할 수 있음
  • TASTE
    • Tokenization stage에서 speech token을 text transcription과 align 하여 modality gap을 완화
    • Attention-based aggregation과 speech reconstruction을 training objective로 사용
  • 논문 (ICLR 2026) : Paper Link

1. Introduction

  • Spoken Language Modeling (SLM)을 위해서는 speech tokenization이 필요함
    • 이를 위해 Wav2Vec2와 같은 Self-Supervised Learning (SSL) representation이나 EnCodec과 같은 neural codec model을 활용할 수 있음
      - BUT, 해당 speech token은 semantic fidelity가 부족하다는 한계가 있음
    • 한편으로 TWIST와 같이 speech, text token을 joint modeling 하는 방법을 고려할 수도 있음
      - BUT, 이 경우 두 modality 간의 length mismatch를 해결해야 함

Concept

-> 그래서 기존 text-speech joint modeling의 misalignment 문제를 개선한 Text-Aligned Speech Tokenization and Embedding (TASTE)를 제안

 

  • TASTE
    • Automatic Speech Recognition (ASR) model을 활용해 textual transcription을 얻고 cross-attention mechanism을 통해 transcription 기반의 speech token을 derive
    • 추가적으로 end-to-end manner로 동작하여 explicit speech-text alignment에 대한 의존성을 제거

< Overall of TASTE >

  • ASR model과 cross-attention 기반의 joint modeling을 활용한 speech tokenization mechanism
  • 결과적으로 기존보다 우수한 성능을 달성

2. Method

- TASTE Speech Tokenizer

  • Speech utterance $\mathbf{u}$, 해당 textual transcription $\mathbf{v}$에 대해, TASTE speech tokenizer $\text{Tokenizer}(\cdot)$는 speech-text pair $X=(\mathbf{u},\mathbf{v})$를 input으로 사용하여 text-aligned speech tokenization, embedding을 생성함
    • 구조적으로 TASTE speech tokenizer는 Encoder, Aggregator, Quantizer로 구성됨
    • 먼저 Encoder $\text{Encoder}(\cdot)$은 $L$ layer Transformer encoder block으로 구성되고 high-dimensional speech representation을 추출함
      1. 이를 위해 논문은 pre-trained Whisper encoder를 도입하고 training 시에는 freeze 함
      2. Input speech utterance $\mathbf{u}$에 대해 encoder는 각 layer $[\mathbf{h}^{(1)},\mathbf{h}^{(2)},...,\mathbf{h}^{(L)}]$의 hidden state sequence를 생성함
      3. 이때 last hidden representation $\mathbf{h}^{(L)}$과 encoder hidden representation의 first-half인 shallow representation $\mathbf{h}^{(l)}$은 retain 됨:
        (Eq. 1) $\mathbf{h}^{(L)},\mathbf{h}^{(l)}=\text{Encoder}(\mathbf{u}),\,\,\,\text{where}\,\,1\leq l \leq \left\lfloor \frac{L}{2}\right\rfloor$
        - $\mathbf{h}^{(L)},\mathbf{h}^{(l)}\in \mathbb{R}^{T\times d_{h}}$, $T$ : length, $d_{h}$ : hidden dimension
    • Encoder에서 추출한 hidden representation은 Aggregator로 전달되고, Aggregator는 text transcription $\mathbf{v}$에 대한 length-aligned compressed speech representation $\mathbf{z}$를 생성함
      1. Length $N$의 text token sequence $\mathbf{v}=[v_{1},v_{2},...,v_{N}],\,\,\, v_{i}\in\mathbb{V}$에 대해, Aggregator의 input/output은:
        (Eq. 2) $\mathbf{z}=\text{Aggregator}(\mathbf{v},\mathbf{h}^{(L)},\mathbf{h}^{(l)}),\,\,\, \text{where}\,\, \mathbf{z}\in\mathbb{R}^{N\times d_{z}},\mathbf{v}\in\mathbb{V}^{N},\,\, \text{and}\,\,\mathbf{h}^{(L)},\mathbf{h}^{(l)}\in\mathbb{R}^{T\times d_{h}}$
      2. 논문은 speech representation $\mathbf{z}$를 text-align 하기 위해 attention mechanism을 도입함
      3. Multi-head attention을 $\text{MultiHead}(Q,K,V)$라 할 때 Aggregator의 first layer attention은:
        (Eq. 3) $Q=\text{text transcription}\,\,\mathbf{v},\,\, K=\text{encoder last hidden}\,\,\mathbf{h}^{(L)},\,\, V=\text{encoder shallow hidden}\,\,\mathbf{h}^{(l)}$
        - 그러면 first multi-head attention output length는 text transcription $\mathbf{v}$를 따라야 함
    • 이후 Quantizer $\text{Quantizer}(\cdot)$을 사용해 text-aligned representation을 discretize 함
      1. 논문은 Residual Vector Quantization (RVQ)를 채택하여 coarse-to-fine quantization을 수행함
      2. Text-aligned speech representation $\mathbf{z}$와 $R$ RVQ layer를 가진 Quantizer는:
        (Eq. 4) $\mathbf{q},\hat{\mathbf{z}}=\text{Quantizer}(\mathbf{z}),\,\,\,\mathbf{q}=[\mathbf{q}^{(1)},\mathbf{q}^{(2)},...,\mathbf{q}^{(R)}],\,\,\hat{\mathbf{z}}=\sum_{r=1}^{R}\hat{\mathbf{z}}^{(r)}$
        - $\mathbf{q}^{(r)}\in\mathbb{C}^{N}$ : code set $\mathbb{C}$를 가지는 $r$-th layer code sequence
      3. Quantized embedding $\hat{\mathbf{z}}$는 codebook vector의 각 layer를 summation하여 얻어짐
        - Code sequence와 quantized speech embedding $\hat{\mathbf{z}}$는 text-aligned이고, length $N$을 가짐

Overview

- TASTE Speech Decoder

  • Speech decoder는 text token sequence와 text-aligned speech tokenization을 기반으로 speech reoncstruction을 수행함
    • 이때 text, speech token은 length-align 되고 autoregressive manner로 weight sum 된 후, speech decoder에 전달됨
    • 구조적으로 speech decoder는 Unit Decoder, Unit-to-Speech Vocoder로 구성됨
      1. Unit Decoder $\text{UnitDecoder}(\cdot)$은 Transformer-based decoder로써, text token sequence $\mathbf{v}$, aligned speech embedding $\hat{\mathbf{z}}$를 condition으로 speech unit $\mathbf{y}$를 predict 함:
        (Eq. 5) $\mathbf{y}=\text{UnitDecoder}(\hat{\mathbf{z}},\mathbf{v})$
      2. Speech unit $\mathbf{y}$를 생성한 다음, Unit-to-Speech Vocoder를 사용하여 unit을 speech로 reconstruct 함

- Training Objective

  • 논문은 original speech $\mathbf{u}$에서 length $T'$의 speech unit $\mathbf{y}^{target}$을 추출하여 Speech Tokenizer, Speech Decoder의 target unit으로 사용함
    • Text transcription $\mathbf{v}$, TASTE speech embedding $\hat{\mathbf{z}}$, original speech unit $\mathbf{y}^{target}$에 대해, $\theta$로 parameterize 된 speech reconstruction은 다음 Cross-Entropy loss를 minimize 하는 것으로 볼 수 있음:
      (Eq. 6) $\mathcal{L}_{ce}(\theta)=\frac{1}{|T'|}\sum_{t=1}^{T'}-\log p_{\theta}\left( y_{t}^{target}|\hat{\mathbf{z}},\mathbf{v};\mathbf{y}_{<t}^{target}\right)$
    • 추가적으로 논문은 Encoder-Aggregator에서 추출된 continuous representation $\mathbf{z}$를 tokenize 하기 위해 다음의 commitment loss를 사용함:
      (Eq. 7) $\mathcal{L}_{rvq}(\theta)=\sum_{r=1}^{R}\left|\left| \mathbf{z}^{(r)}-\hat{\mathbf{z}}^{(r)}\right|\right|$
      - $\mathbf{z}^{(r)}$ : $r$-th residual, $\hat{\mathbf{z}}^{(r)}$ : $r$-th quantized residual
    • 결과적으로 TASTE의 overall training loss는:
      (Eq. 8) $\mathcal{L}_{taste}=\mathcal{L}_{ce}+\mathcal{L}_{rvq}$

- Modeling TASTE Token

  • TASTE 기반의 spoken language modeling을 위해 Text-Aligned Spoken Language Model (TASLM)을 고려함
    • 먼저 RVQ quantizer에서 derive 된 speech token은 $R$ layer code를 포함하고 있으므로, $\text{TASTLM}_{token}$은 $R$ linear head를 사용하여 multi-head prediction을 수행함
      - 즉, $\text{TASTLM}_{token}$은 각 step에서 next text token과 해당 $R$ layer의 speech token을 simultaneously predict 함
    • Text transcription $\mathbf{v}$와 $R$ layer의 quantized RVQ code $\mathbf{q}$가 주어졌을 때, multi-head next-token prediction training objective는:
      (Eq. 9) $\mathcal{L}_{token}(\phi)=\frac{1}{|N|}\sum_{i=1}^{N}\left(-\log p_{\phi}^{text}\left(v_{i}|\mathbf{v}_{<i},\mathbf{q}_{<i}\right)+ \sum_{r=1}^{R}-\log p_{\phi}^{(r)}\left( q_{i}^{(r)}|\mathbf{v}_{<i},\mathbf{q}_{<i}\right)\right)$
      - $\phi$ : $\text{TASLM}_{token}$의 parameter, $p^{(r)}$ : $r$-th RVQ code의 $r$-th probability prediction
    • 추론 시에는 code, text를 directly sampling 하고 code를 Speech Decoder의 embedding으로 transform 함

- Modeling TASTE Embedding

  • 추가적으로 latent modeling을 위해 $i$-th sequence frame에 대해, mean vector $\mu_{i}$, log-magnitude variance vector $\log \sigma_{i}^{2}$을 predict 하는 linear layer를 고려할 수 있음
    • 그러면 $i$ frame의 final predicted latent는 $e_{i}=\mu_{i}+\sigma_{i}\odot \epsilon$과 같음
      - $\epsilon\sim\mathcal{N}(0,I)$
    • Training 시에는 backpropagation을 위해 straight-through estimator를 사용하고, latent prediction을 위해 regularization loss와 Kullback-Leibler (KL) divergence loss를 도입함:
      (Eq. 10) $\mathcal{L}_{reg}(\psi)=||\mathbf{e}_{\psi}-\hat{\mathbf{z}}||_{2}^{2},\,\,\,\mathcal{L}_{KL}=\frac{1}{2}\sum_{i=1}^{N}\sum_{j=1}^{d_{z}}\left( \sigma_{i}[j]+\left( \mu_{i}[j]-\hat{z}_{i}[j]\right)^{2}-1-\log \sigma^{2}_{i}[j]\right)$
      - $\psi$ : $\text{TASLM}_{emb}$ parameter, $d_{z}$ : text-aligned embedding $\hat{\mathbf{z}}$의 dimension,
      $\mathcal{L}_{reg}$ : regularization loss, $\mathcal{L}_{KL}$ : KL-divergence loss
    • 이때 target distribution은 MELLE를 따라 $\mathcal{N}(\hat{\mathbf{z}}_{i},I)$로 select 함
      - 그러면 $\mathcal{L}_{KL}$을 simplify 할 수 있고, predicted vector $\mu_{i},\sigma_{i}$, target embedding $\hat{\mathbf{z}}_{i}$로 approximate 할 수 있음
    • 결과적으로 overall loss는:
      (Eq. 11) $\mathcal{L}_{emb}(\psi)=\lambda_{reg}\cdot\mathcal{L}_{reg}+\lambda_{KL}\cdot\mathcal{L}_{KL}+\frac{1}{|N|}\sum_{i=1}^{N}-\log p_{\psi}^{text}\left(v_{i}|\mathbf{v}_{<i},\hat{\mathbf{z}}_{<i}\right)$
      - $\lambda_{reg},\lambda_{KL}$ : coefficient

3. Experiments

- Settings

- Results

  • 전체적으로 TASTE의 성능이 가장 우수함

Model 성능 비교

  • Spoken Language Modeling
    • TASTE를 사용한 SLM이 더 나은 성능을 보임

SLM에서의 성능

  • Text-Aligned Speech Editing
    • Specific word position의 token을 swapping 하면 해당 word에서 duration shift가 나타남

Text-Aligned Speech Editing

  • Spoken Question Answering
    • Spoken Question Answering task에 대해서도 우수한 성능을 보임

Spoken Question Answering

  • Ablation Study
    • 각 component는 성능 향상에 유효함

Ablation Study

 

반응형
댓글
최근에 올라온 글
최근에 달린 댓글
«   2026/10   »
일 월 화 수 목 금 토
1 2 3
4 5 6 7 8 9 10
11 12 13 14 15 16 17
18 19 20 21 22 23 24
25 26 27 28 29 30 31
Total
Today
Yesterday