The Otals - SuperhyperdestructiveIloveyou!!!- Date: 2026.08.11.- Location: Japan, Tokyo, Shibuya, Club Quattro 이 이야기는 지난 나고야 라이브에서 시작됩니다. FAXxxxxx : 우리 라이브 한 번 더 한다! 마지막 투어 보러 왔더니 추가 공연 소식에 장소는 무려 시부야 콰트로 라는데 이걸 마다할 이유가 있나요. 무조건 콜이죠. 아무튼 그렇게 들뜬 마음을 안고 예매처를 확인했는데... 로치케였습니다. 사실 여기서 티켓을 사려면 일본 전화번호가 필요해서 외국인 입장에서는 방법이 없습니다. 그렇다고 한국 기준으론 아예 불가능한 건 아닌데, 일본 유심을 구해서 날씨 좋은 날에 부산 앞바다로 간 다음, 거기서 대마도 회..
이달의 슈게이즈 7회 - 26년 7월* 업로드 당일 기준 작성자 레이더망에 걸린 것들만 올리니 놓치는게 있을 수도 있습니다. 1. 경이로운 완급조절 호주의 Swapmeet이 발표한 신보 이야기로 시작해 봅시다. 17일에 발매된 는 Sonic Youth, Pavement와 같은 노이즈팝 사운드를 기반으로 청년기의 불안감과 감정적인 인간관계를 다룬 앨범입니다. 무엇보다도 앨범의 사운드 구성을 칭찬하지 않을 수가 없는데, 자칫 지루해지기 쉬운 느릿한 템포를 폭발적인 노이즈 활용과 진한 숙취의 멜로디로 완벽하게 조절하면서 상당한 흡입력을 만들어냈습니다.Swapmeet - 'Seeds'2. 대서양을 건너 3일에는 캐나다의 Sophie Brightman과 스페인의 Javier Manriquedelara가 결성한 듀..
Self-Guidance: Enhancing Neural Codecs via Decoder Manifold AlignmentNeural codec은 quantization error로 인한 reconstruction fidelity의 한계가 존재함Self-GuidanceQuantized token, continuous embedding을 process 할 때 internal decoder feature manifold를 align이를 위해 lightweight feature mapping loss를 도입논문 (ICML 2026) : Paper Link1. IntroductionSoundStream, EnCodec과 같은 Vector-Quantized Variational AutoEncoder (VQ-VA..
AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech기존 Text-to-Speech model은 composite instruction에 대한 fine-grained control이 어려움AgentSteerTTSAdversarial Disetanglement Agent를 활용해 speaker-emotion leakage를 방지Dual-Stream Anchoring Controller를 통해 abstract intent를 반영하고 Fast-Slow Feedback Agent를 통해 semantic-acoustic mismatch를 resolve논문 (ICML 2026) : Paper Link1. ..
SPEAR: A Unified SSL Framework for Learning Speech and Audio RepresentationsSelf-Supervised Learning은 speech, audio event understanding에 대한 gap이 존재함SPEARContinuous teacher representation에 multi-codebook vector quantization을 적용하여 semantic, acoustic information을 capture추가적으로 asymmetric pre-training loss와 token mixing을 활용해 robustness를 향상논문 (ICML 2026) : Paper Link1. IntroductionSpeech processing을 위..
DisCo-Speech: Controllable Zero-Shot Speech Generation with a Disentangled Speech Codec기존 codec은 timbre, prosody의 entanglement로 인해 independent control이 어려움DisCo-SpeechParallel encoder와 hybrid loss를 사용하여 speech를 content, prosody, timbre의 tri-factor로 disentangleUnified content-prosody token을 구성해 disentanglement-reconstruction trade-off를 balance논문 (ACL 2026) : Paper Link1. IntroductionCodec-based L..
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-TrainingFine-grained speaking style을 modeling 하는 것은 어려움CLSP47k hours speech, 19M fine-grained caption을 포함한 FCaps dataset을 구축FCaps dataset을 기반으로 global, fine-grained supervision을 integrate 한 contrastive language-speech pre-trained modeling을 수행논문 (ACL 2026) : Paper Link1. IntroductionSpeaking style은 gender, age와 같은 speaker-int..
ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation AlignmentEnvironmental audio와 함께 speech를 jointly generate 하는 것은 어려움ImmersiveTTSMutimodal diffusion Transformer를 기반으로 transcript-aligned speech latent와 text-conditioned environmental context를 joint attention으로 fuseSemantic consistency를 향상하기 위해 domain-specific representation alig..
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream QuantizationSpeech codec은 semantically-rich representation과 high-quality reconstruction에 대한 trade-off가 존재함SACSemantic-acoustic dual-stream quantization을 활용해 semantic, acoustic modeling을 disentangling두 개의 dedicated stream을 각각의 respective role에 맞게 optimize논문 (ACL 2026) : Paper Link1. IntroductionSpeech tokenizer는 continuous speech wavefor..