论文解读

LMPAN: A Lightweight Multi-Path Alignment Network for Joint Full-Duplex Acoustic Echo Cancellation and Noise Suppression

语音增强 | 6.2/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8743 字 阅读 →
论文解读

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track

语音翻译 | 6.2/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7251 字 阅读 →
论文解读

Neural Audio Codec with Adjustable Token Temporal Resolution Using Sampling-Frequency-Independent Convolutional Layers

Neural Audio Codec with Adjustable Token Temporal Resolution Using Sampling-Frequency-Independent Convolutional Layers

 · 更新于 2026-09-06 · 约 11 分钟 · 5152 字 阅读 →
论文解读

Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack

语音唤醒 | 6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6409 字 阅读 →
论文解读

Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score

音频检索 | 5.2/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6672 字 阅读 →
论文解读

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

音视频理解 | 7.2/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7157 字 阅读 →
论文解读

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

语音识别 | 5.6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6166 字 阅读 →
论文解读

RT-Tango: Real-Time Distributed Binaural Speech Enhancement for Low-Power Hearing Aid Devices

语音增强 | 5.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6468 字 阅读 →
论文解读

SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios

声源定位 | 7.1/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7824 字 阅读 →
论文解读

Self-Supervised Test-Time Tuning for Packet Loss Concealment

音频修复 | 7.4/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7565 字 阅读 →
论文解读

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

语音合成 | 5.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5960 字 阅读 →
论文解读

Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition

声源定位 | 4.1/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5837 字 阅读 →
论文解读

Speaker head orientation estimation with a single microphone array using phase spectrogram features

声源定位 | 5.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5887 字 阅读 →
论文解读

Towards a Phonology-Informed Evaluation of Multilingual TTS

语音质量评估 | 5.7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6691 字 阅读 →
论文解读

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

语音交互 | 7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5205 字 阅读 →
论文解读

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

语音交互 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5306 字 阅读 →
论文解读

Using embeddings to predict spoken word duration and pitch in Mandarin monosyllabic words

语音合成 | 4/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7644 字 阅读 →
论文解读

UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation

音乐生成 | 4.1/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6889 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-03

共分析 31 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 101 分钟 · 50504 字 阅读 →
论文解读

A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

语音合成 | 6.6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6328 字 阅读 →