HoliTok: A Continuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and UnderstandingUnified speech modeling을 위해서는 holistic tokenization space가 필요함HoliTokSignal-level fidelity를 preserve 하면서 semantic information을 incorporate 하는 progressive training을 도입Unified AR+DiT model을 활용해 latent sequence가 generation-specific, unified generation-understanding task를 모두 지원하도록 구성논문 (EM..
이달의 슈게이즈 9회 - 26년 9월* 업로드 당일 기준 작성자 레이더망에 걸린 것들만 올리니 놓치는게 있을 수도 있습니다. 1. 눈부신 청춘을 향해 8월 28일에 떨어진 탓에 지난 이달슈에서 미처 말하지 못한 중국 슈게이즈 밴드 一點alittlebit의 데뷔 EP 로 시작해 봅시다. 전체적으로 앨범은 더 단단한 어른이 되어가는 여정을 담고 있습니다. 특히 사운드적인 측면에서 밴드는 층층이 쌓인 레이어와 감미로운 여성 보컬을 조합하여, 앨범 제목처럼 비 온 뒤의 햇살을 맞이하는 듯한 아름다운 사운드를 들려줍니다.一點alittlebit - '樹'2. 춤추는 소용돌이 4일 뉴욕에서는 Diary가 신보 를 발매했습니다. 쟁글팝, 매드체스터 기반의 댄서블한 슈게이즈 사운드를 지향하는 앨범으로써, 해당 장르의 산..
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language ModelingSpeech token을 활용하면 joint text-speech modeling을 향상할 수 있음TASTETokenization stage에서 speech token을 text transcription과 align 하여 modality gap을 완화Attention-based aggregation과 speech reconstruction을 training objective로 사용논문 (ICLR 2026) : Paper Link1. IntroductionSpoken Language Modeling (SLM)을 위해서는 speech tokenization이 필요..
EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis기존의 emotional Text-to-Speech는 emotion의 temporal nature와 misalign 함EmoTra-TTSMulti-pass flow blending pipeline을 활용해 frame-aligned transition audio를 반영Dual-stage Valence-Arousal-Dominance conditioning을 통해 LLM과 Flow Decoder를 guideDirection-magnitude decoupled injection을 사용하여 content degradation을 방지논문 (EMNLP 2026) : Paper Lin..
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal GroundingMulti-speaker automatic speech recognition과 diarization을 위한 model이 필요함TagSpeechSerialized Output Training을 통해 semantic, speaker stream을 decouple 하고 turn-taking dynamics를 학습Interleaved time anchor mechanism을 활용해 fine-grained timestamp prediction을 지원논문 (ACL 2026) : Paper Link1. Introduction최근에는 Large Langua..
Data-Efficient Targeted Token-Level Preference Optimization for LLM-based Text-to-SpeechPreference optimization을 통해 system output을 human feedback과 align 할 수 있음TKTOPaired data에 의존하지 않고 token-level unit을 target으로 사용해 data-efficient training을 지원Token-level annotation 없이 fine-grained alignment signal을 제공논문 (ACL 2026) : Paper Link1. IntroductionText-to-Speech (TTS) model은 Grapheme-to-Phoneme (G2P)를 사..
POWSM: A Phonetic Open Whisper-Style Speech Foundation ModelAutomatic Speech Recognition, Phoneme Recognition, Grapheme-to-Phoneme, Phoneme-to-Grapheme conversion은 각각의 task-specific architecture에 의존함POWSMMultiple phone-related task를 jointly perform 하는 unified frameworkAudio, text, phone 간의 seamless conversion을 지원논문 (ACL 2026) : Paper Link1. IntroductionPhone은 모든 language에 대해 International Phonet..
Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-SpeechHigh-arousal emotion을 생성하는 것은 여전히 어려움Rectifying the Emotional FlowEmotion Rectified Noise Prior를 활용해 initialization 시 semantic gradient를 injectLikelihood-Inverse Guidance를 통해 guidance를 adaptively schedule논문 (ACL 2026) : Paper Link1. IntroductionText-to-Speech (TTS) model은 stability-expressive bottl..
이달의 슈게이즈 8회 - 26년 8월 * 업로드 당일 기준 작성자 레이더망에 걸린 것들만 올리니 놓치는게 있을 수도 있습니다. 1. 돌고 돌아 잉글랜드 영국 사우샘프턴에서 활동하는 Shadow Flowers의 데뷔앨범으로 시작해 봅시다. 지난 7일에 발매된 은 근본과 전통의 영국씬답게 Slowdive, Mercury Rev 등을 계승한 몽환적인 슈게이즈 사운드를 들려줍니다. 전반적으로 완벽하게 다듬어진 앨범은 아닙니다만, 윙윙거리는 소음과 사이키델릭한 선율이 조화된 3번 트랙 'Touching' 하나만큼은 올해 최고의 트랙이라고 할 수 있습니다.Shadow Flowers - 'Touching'2. 실연에도 국경이 있나요 한편 같은 날 폴란드에서는 Duszno의 새 EP 가 발매되었습니다. 제목을 번역기에..