티스토리 뷰
Paper/Neural Codec
[Paper 리뷰] TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
feVeRin 2026. 9. 21.반응형
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
- Speech token을 활용하면 joint text-speech modeling을 향상할 수 있음
- TASTE
- Tokenization stage에서 speech token을 text transcription과 align 하여 modality gap을 완화
- Attention-based aggregation과 speech reconstruction을 training objective로 사용
- 논문 (ICLR 2026) : Paper Link
1. Introduction
- Spoken Language Modeling (SLM)을 위해서는 speech tokenization이 필요함

-> 그래서 기존 text-speech joint modeling의 misalignment 문제를 개선한 Text-Aligned Speech Tokenization and Embedding (TASTE)를 제안
- TASTE
- Automatic Speech Recognition (ASR) model을 활용해 textual transcription을 얻고 cross-attention mechanism을 통해 transcription 기반의 speech token을 derive
- 추가적으로 end-to-end manner로 동작하여 explicit speech-text alignment에 대한 의존성을 제거
< Overall of TASTE >
- ASR model과 cross-attention 기반의 joint modeling을 활용한 speech tokenization mechanism
- 결과적으로 기존보다 우수한 성능을 달성
2. Method
- TASTE Speech Tokenizer
- Speech utterance $\mathbf{u}$, 해당 textual transcription $\mathbf{v}$에 대해, TASTE speech tokenizer $\text{Tokenizer}(\cdot)$는 speech-text pair $X=(\mathbf{u},\mathbf{v})$를 input으로 사용하여 text-aligned speech tokenization, embedding을 생성함
- 구조적으로 TASTE speech tokenizer는 Encoder, Aggregator, Quantizer로 구성됨
- 먼저 Encoder $\text{Encoder}(\cdot)$은 $L$ layer Transformer encoder block으로 구성되고 high-dimensional speech representation을 추출함
- 이를 위해 논문은 pre-trained Whisper encoder를 도입하고 training 시에는 freeze 함
- Input speech utterance $\mathbf{u}$에 대해 encoder는 각 layer $[\mathbf{h}^{(1)},\mathbf{h}^{(2)},...,\mathbf{h}^{(L)}]$의 hidden state sequence를 생성함
- 이때 last hidden representation $\mathbf{h}^{(L)}$과 encoder hidden representation의 first-half인 shallow representation $\mathbf{h}^{(l)}$은 retain 됨:
(Eq. 1) $\mathbf{h}^{(L)},\mathbf{h}^{(l)}=\text{Encoder}(\mathbf{u}),\,\,\,\text{where}\,\,1\leq l \leq \left\lfloor \frac{L}{2}\right\rfloor$
- $\mathbf{h}^{(L)},\mathbf{h}^{(l)}\in \mathbb{R}^{T\times d_{h}}$, $T$ : length, $d_{h}$ : hidden dimension
- Encoder에서 추출한 hidden representation은 Aggregator로 전달되고, Aggregator는 text transcription $\mathbf{v}$에 대한 length-aligned compressed speech representation $\mathbf{z}$를 생성함
- Length $N$의 text token sequence $\mathbf{v}=[v_{1},v_{2},...,v_{N}],\,\,\, v_{i}\in\mathbb{V}$에 대해, Aggregator의 input/output은:
(Eq. 2) $\mathbf{z}=\text{Aggregator}(\mathbf{v},\mathbf{h}^{(L)},\mathbf{h}^{(l)}),\,\,\, \text{where}\,\, \mathbf{z}\in\mathbb{R}^{N\times d_{z}},\mathbf{v}\in\mathbb{V}^{N},\,\, \text{and}\,\,\mathbf{h}^{(L)},\mathbf{h}^{(l)}\in\mathbb{R}^{T\times d_{h}}$ - 논문은 speech representation $\mathbf{z}$를 text-align 하기 위해 attention mechanism을 도입함
- Multi-head attention을 $\text{MultiHead}(Q,K,V)$라 할 때 Aggregator의 first layer attention은:
(Eq. 3) $Q=\text{text transcription}\,\,\mathbf{v},\,\, K=\text{encoder last hidden}\,\,\mathbf{h}^{(L)},\,\, V=\text{encoder shallow hidden}\,\,\mathbf{h}^{(l)}$
- 그러면 first multi-head attention output length는 text transcription $\mathbf{v}$를 따라야 함
- Length $N$의 text token sequence $\mathbf{v}=[v_{1},v_{2},...,v_{N}],\,\,\, v_{i}\in\mathbb{V}$에 대해, Aggregator의 input/output은:
- 이후 Quantizer $\text{Quantizer}(\cdot)$을 사용해 text-aligned representation을 discretize 함
- 논문은 Residual Vector Quantization (RVQ)를 채택하여 coarse-to-fine quantization을 수행함
- Text-aligned speech representation $\mathbf{z}$와 $R$ RVQ layer를 가진 Quantizer는:
(Eq. 4) $\mathbf{q},\hat{\mathbf{z}}=\text{Quantizer}(\mathbf{z}),\,\,\,\mathbf{q}=[\mathbf{q}^{(1)},\mathbf{q}^{(2)},...,\mathbf{q}^{(R)}],\,\,\hat{\mathbf{z}}=\sum_{r=1}^{R}\hat{\mathbf{z}}^{(r)}$
- $\mathbf{q}^{(r)}\in\mathbb{C}^{N}$ : code set $\mathbb{C}$를 가지는 $r$-th layer code sequence - Quantized embedding $\hat{\mathbf{z}}$는 codebook vector의 각 layer를 summation하여 얻어짐
- Code sequence와 quantized speech embedding $\hat{\mathbf{z}}$는 text-aligned이고, length $N$을 가짐

- TASTE Speech Decoder
- Speech decoder는 text token sequence와 text-aligned speech tokenization을 기반으로 speech reoncstruction을 수행함
- 이때 text, speech token은 length-align 되고 autoregressive manner로 weight sum 된 후, speech decoder에 전달됨
- 구조적으로 speech decoder는 Unit Decoder, Unit-to-Speech Vocoder로 구성됨
- Unit Decoder $\text{UnitDecoder}(\cdot)$은 Transformer-based decoder로써, text token sequence $\mathbf{v}$, aligned speech embedding $\hat{\mathbf{z}}$를 condition으로 speech unit $\mathbf{y}$를 predict 함:
(Eq. 5) $\mathbf{y}=\text{UnitDecoder}(\hat{\mathbf{z}},\mathbf{v})$ - Speech unit $\mathbf{y}$를 생성한 다음, Unit-to-Speech Vocoder를 사용하여 unit을 speech로 reconstruct 함
- Unit Decoder $\text{UnitDecoder}(\cdot)$은 Transformer-based decoder로써, text token sequence $\mathbf{v}$, aligned speech embedding $\hat{\mathbf{z}}$를 condition으로 speech unit $\mathbf{y}$를 predict 함:
- Training Objective
- 논문은 original speech $\mathbf{u}$에서 length $T'$의 speech unit $\mathbf{y}^{target}$을 추출하여 Speech Tokenizer, Speech Decoder의 target unit으로 사용함
- Text transcription $\mathbf{v}$, TASTE speech embedding $\hat{\mathbf{z}}$, original speech unit $\mathbf{y}^{target}$에 대해, $\theta$로 parameterize 된 speech reconstruction은 다음 Cross-Entropy loss를 minimize 하는 것으로 볼 수 있음:
(Eq. 6) $\mathcal{L}_{ce}(\theta)=\frac{1}{|T'|}\sum_{t=1}^{T'}-\log p_{\theta}\left( y_{t}^{target}|\hat{\mathbf{z}},\mathbf{v};\mathbf{y}_{<t}^{target}\right)$ - 추가적으로 논문은 Encoder-Aggregator에서 추출된 continuous representation $\mathbf{z}$를 tokenize 하기 위해 다음의 commitment loss를 사용함:
(Eq. 7) $\mathcal{L}_{rvq}(\theta)=\sum_{r=1}^{R}\left|\left| \mathbf{z}^{(r)}-\hat{\mathbf{z}}^{(r)}\right|\right|$
- $\mathbf{z}^{(r)}$ : $r$-th residual, $\hat{\mathbf{z}}^{(r)}$ : $r$-th quantized residual - 결과적으로 TASTE의 overall training loss는:
(Eq. 8) $\mathcal{L}_{taste}=\mathcal{L}_{ce}+\mathcal{L}_{rvq}$
- Text transcription $\mathbf{v}$, TASTE speech embedding $\hat{\mathbf{z}}$, original speech unit $\mathbf{y}^{target}$에 대해, $\theta$로 parameterize 된 speech reconstruction은 다음 Cross-Entropy loss를 minimize 하는 것으로 볼 수 있음:
- Modeling TASTE Token
- TASTE 기반의 spoken language modeling을 위해 Text-Aligned Spoken Language Model (TASLM)을 고려함
- 먼저 RVQ quantizer에서 derive 된 speech token은 $R$ layer code를 포함하고 있으므로, $\text{TASTLM}_{token}$은 $R$ linear head를 사용하여 multi-head prediction을 수행함
- 즉, $\text{TASTLM}_{token}$은 각 step에서 next text token과 해당 $R$ layer의 speech token을 simultaneously predict 함 - Text transcription $\mathbf{v}$와 $R$ layer의 quantized RVQ code $\mathbf{q}$가 주어졌을 때, multi-head next-token prediction training objective는:
(Eq. 9) $\mathcal{L}_{token}(\phi)=\frac{1}{|N|}\sum_{i=1}^{N}\left(-\log p_{\phi}^{text}\left(v_{i}|\mathbf{v}_{<i},\mathbf{q}_{<i}\right)+ \sum_{r=1}^{R}-\log p_{\phi}^{(r)}\left( q_{i}^{(r)}|\mathbf{v}_{<i},\mathbf{q}_{<i}\right)\right)$
- $\phi$ : $\text{TASLM}_{token}$의 parameter, $p^{(r)}$ : $r$-th RVQ code의 $r$-th probability prediction - 추론 시에는 code, text를 directly sampling 하고 code를 Speech Decoder의 embedding으로 transform 함
- 먼저 RVQ quantizer에서 derive 된 speech token은 $R$ layer code를 포함하고 있으므로, $\text{TASTLM}_{token}$은 $R$ linear head를 사용하여 multi-head prediction을 수행함
- Modeling TASTE Embedding
- 추가적으로 latent modeling을 위해 $i$-th sequence frame에 대해, mean vector $\mu_{i}$, log-magnitude variance vector $\log \sigma_{i}^{2}$을 predict 하는 linear layer를 고려할 수 있음
- 그러면 $i$ frame의 final predicted latent는 $e_{i}=\mu_{i}+\sigma_{i}\odot \epsilon$과 같음
- $\epsilon\sim\mathcal{N}(0,I)$ - Training 시에는 backpropagation을 위해 straight-through estimator를 사용하고, latent prediction을 위해 regularization loss와 Kullback-Leibler (KL) divergence loss를 도입함:
(Eq. 10) $\mathcal{L}_{reg}(\psi)=||\mathbf{e}_{\psi}-\hat{\mathbf{z}}||_{2}^{2},\,\,\,\mathcal{L}_{KL}=\frac{1}{2}\sum_{i=1}^{N}\sum_{j=1}^{d_{z}}\left( \sigma_{i}[j]+\left( \mu_{i}[j]-\hat{z}_{i}[j]\right)^{2}-1-\log \sigma^{2}_{i}[j]\right)$
- $\psi$ : $\text{TASLM}_{emb}$ parameter, $d_{z}$ : text-aligned embedding $\hat{\mathbf{z}}$의 dimension, $\mathcal{L}_{reg}$ : regularization loss, $\mathcal{L}_{KL}$ : KL-divergence loss - 이때 target distribution은 MELLE를 따라 $\mathcal{N}(\hat{\mathbf{z}}_{i},I)$로 select 함
- 그러면 $\mathcal{L}_{KL}$을 simplify 할 수 있고, predicted vector $\mu_{i},\sigma_{i}$, target embedding $\hat{\mathbf{z}}_{i}$로 approximate 할 수 있음 - 결과적으로 overall loss는:
(Eq. 11) $\mathcal{L}_{emb}(\psi)=\lambda_{reg}\cdot\mathcal{L}_{reg}+\lambda_{KL}\cdot\mathcal{L}_{KL}+\frac{1}{|N|}\sum_{i=1}^{N}-\log p_{\psi}^{text}\left(v_{i}|\mathbf{v}_{<i},\hat{\mathbf{z}}_{<i}\right)$
- $\lambda_{reg},\lambda_{KL}$ : coefficient
- 그러면 $i$ frame의 final predicted latent는 $e_{i}=\mu_{i}+\sigma_{i}\odot \epsilon$과 같음
3. Experiments
- Settings
- Dataset : Emilia, LibriTTS
- Comparisons
- Tokenizer : EnCodec, SpeechTokenizer, CosyVoice, Mimi
- SLM : TWIST, SpiritLM
- Results
- 전체적으로 TASTE의 성능이 가장 우수함

- Spoken Language Modeling
- TASTE를 사용한 SLM이 더 나은 성능을 보임

- Text-Aligned Speech Editing
- Specific word position의 token을 swapping 하면 해당 word에서 duration shift가 나타남

- Spoken Question Answering
- Spoken Question Answering task에 대해서도 우수한 성능을 보임

- Ablation Study
- 각 component는 성능 향상에 유효함

반응형
'Paper > Neural Codec' 카테고리의 다른 글
댓글