Self-Guidance: Enhancing Neural Codecs via Decoder Manifold AlignmentNeural codec은 quantization error로 인한 reconstruction fidelity의 한계가 존재함Self-GuidanceQuantized token, continuous embedding을 process 할 때 internal decoder feature manifold를 align이를 위해 lightweight feature mapping loss를 도입논문 (ICML 2026) : Paper Link1. IntroductionSoundStream, EnCodec과 같은 Vector-Quantized Variational AutoEncoder (VQ-VA..
AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech기존 Text-to-Speech model은 composite instruction에 대한 fine-grained control이 어려움AgentSteerTTSAdversarial Disetanglement Agent를 활용해 speaker-emotion leakage를 방지Dual-Stream Anchoring Controller를 통해 abstract intent를 반영하고 Fast-Slow Feedback Agent를 통해 semantic-acoustic mismatch를 resolve논문 (ICML 2026) : Paper Link1. ..
SPEAR: A Unified SSL Framework for Learning Speech and Audio RepresentationsSelf-Supervised Learning은 speech, audio event understanding에 대한 gap이 존재함SPEARContinuous teacher representation에 multi-codebook vector quantization을 적용하여 semantic, acoustic information을 capture추가적으로 asymmetric pre-training loss와 token mixing을 활용해 robustness를 향상논문 (ICML 2026) : Paper Link1. IntroductionSpeech processing을 위..
DisCo-Speech: Controllable Zero-Shot Speech Generation with a Disentangled Speech Codec기존 codec은 timbre, prosody의 entanglement로 인해 independent control이 어려움DisCo-SpeechParallel encoder와 hybrid loss를 사용하여 speech를 content, prosody, timbre의 tri-factor로 disentangleUnified content-prosody token을 구성해 disentanglement-reconstruction trade-off를 balance논문 (ACL 2026) : Paper Link1. IntroductionCodec-based L..
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-TrainingFine-grained speaking style을 modeling 하는 것은 어려움CLSP47k hours speech, 19M fine-grained caption을 포함한 FCaps dataset을 구축FCaps dataset을 기반으로 global, fine-grained supervision을 integrate 한 contrastive language-speech pre-trained modeling을 수행논문 (ACL 2026) : Paper Link1. IntroductionSpeaking style은 gender, age와 같은 speaker-int..
ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation AlignmentEnvironmental audio와 함께 speech를 jointly generate 하는 것은 어려움ImmersiveTTSMutimodal diffusion Transformer를 기반으로 transcript-aligned speech latent와 text-conditioned environmental context를 joint attention으로 fuseSemantic consistency를 향상하기 위해 domain-specific representation alig..