Data-Efficient Targeted Token-Level Preference Optimization for LLM-based Text-to-SpeechPreference optimization을 통해 system output을 human feedback과 align 할 수 있음TKTOPaired data에 의존하지 않고 token-level unit을 target으로 사용해 data-efficient training을 지원Token-level annotation 없이 fine-grained alignment signal을 제공논문 (ACL 2026) : Paper Link1. IntroductionText-to-Speech (TTS) model은 Grapheme-to-Phoneme (G2P)를 사..
POWSM: A Phonetic Open Whisper-Style Speech Foundation ModelAutomatic Speech Recognition, Phoneme Recognition, Grapheme-to-Phoneme, Phoneme-to-Grapheme conversion은 각각의 task-specific architecture에 의존함POWSMMultiple phone-related task를 jointly perform 하는 unified frameworkAudio, text, phone 간의 seamless conversion을 지원논문 (ACL 2026) : Paper Link1. IntroductionPhone은 모든 language에 대해 International Phonet..
Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-SpeechHigh-arousal emotion을 생성하는 것은 여전히 어려움Rectifying the Emotional FlowEmotion Rectified Noise Prior를 활용해 initialization 시 semantic gradient를 injectLikelihood-Inverse Guidance를 통해 guidance를 adaptively schedule논문 (ACL 2026) : Paper Link1. IntroductionText-to-Speech (TTS) model은 stability-expressive bottl..
이달의 슈게이즈 8회 - 26년 8월 * 업로드 당일 기준 작성자 레이더망에 걸린 것들만 올리니 놓치는게 있을 수도 있습니다. 1. 돌고 돌아 잉글랜드 영국 사우샘프턴에서 활동하는 Shadow Flowers의 데뷔앨범으로 시작해 봅시다. 지난 7일에 발매된 은 근본과 전통의 영국씬답게 Slowdive, Mercury Rev 등을 계승한 몽환적인 슈게이즈 사운드를 들려줍니다. 전반적으로 완벽하게 다듬어진 앨범은 아닙니다만, 윙윙거리는 소음과 사이키델릭한 선율이 조화된 3번 트랙 'Touching' 하나만큼은 올해 최고의 트랙이라고 할 수 있습니다.Shadow Flowers - 'Touching'2. 실연에도 국경이 있나요 한편 같은 날 폴란드에서는 Duszno의 새 EP 가 발매되었습니다. 제목을 번역기에..
The Otals - SuperhyperdestructiveIloveyou!!!- Date: 2026.08.11.- Location: Japan, Tokyo, Shibuya, Club Quattro 이 이야기는 지난 나고야 라이브에서 시작됩니다. FAXxxxxx : 우리 라이브 한 번 더 한다! 마지막 투어 보러 왔더니 추가 공연 소식에 장소는 무려 시부야 콰트로 라는데 이걸 마다할 이유가 있나요. 무조건 콜이죠. 아무튼 그렇게 들뜬 마음을 안고 예매처를 확인했는데... 로치케였습니다. 사실 여기서 티켓을 사려면 일본 전화번호가 필요해서 외국인 입장에서는 방법이 없습니다. 그렇다고 한국 기준으론 아예 불가능한 건 아닌데, 일본 유심을 구해서 날씨 좋은 날에 부산 앞바다로 간 다음, 거기서 대마도 회..
이달의 슈게이즈 7회 - 26년 7월* 업로드 당일 기준 작성자 레이더망에 걸린 것들만 올리니 놓치는게 있을 수도 있습니다. 1. 경이로운 완급조절 호주의 Swapmeet이 발표한 신보 이야기로 시작해 봅시다. 17일에 발매된 는 Sonic Youth, Pavement와 같은 노이즈팝 사운드를 기반으로 청년기의 불안감과 감정적인 인간관계를 다룬 앨범입니다. 무엇보다도 앨범의 사운드 구성을 칭찬하지 않을 수가 없는데, 자칫 지루해지기 쉬운 느릿한 템포를 폭발적인 노이즈 활용과 진한 숙취의 멜로디로 완벽하게 조절하면서 상당한 흡입력을 만들어냈습니다.Swapmeet - 'Seeds'2. 대서양을 건너 3일에는 캐나다의 Sophie Brightman과 스페인의 Javier Manriquedelara가 결성한 듀..
Self-Guidance: Enhancing Neural Codecs via Decoder Manifold AlignmentNeural codec은 quantization error로 인한 reconstruction fidelity의 한계가 존재함Self-GuidanceQuantized token, continuous embedding을 process 할 때 internal decoder feature manifold를 align이를 위해 lightweight feature mapping loss를 도입논문 (ICML 2026) : Paper Link1. IntroductionSoundStream, EnCodec과 같은 Vector-Quantized Variational AutoEncoder (VQ-VA..
AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech기존 Text-to-Speech model은 composite instruction에 대한 fine-grained control이 어려움AgentSteerTTSAdversarial Disetanglement Agent를 활용해 speaker-emotion leakage를 방지Dual-Stream Anchoring Controller를 통해 abstract intent를 반영하고 Fast-Slow Feedback Agent를 통해 semantic-acoustic mismatch를 resolve논문 (ICML 2026) : Paper Link1. ..
SPEAR: A Unified SSL Framework for Learning Speech and Audio RepresentationsSelf-Supervised Learning은 speech, audio event understanding에 대한 gap이 존재함SPEARContinuous teacher representation에 multi-codebook vector quantization을 적용하여 semantic, acoustic information을 capture추가적으로 asymmetric pre-training loss와 token mixing을 활용해 robustness를 향상논문 (ICML 2026) : Paper Link1. IntroductionSpeech processing을 위..
