论文解读

Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization

语音识别 说话人日志 | 5.5/10

 · 更新于 2026-10-01 · 约 21 分钟 · 10162 字 阅读 →
论文解读

ChladniSonify: A Visual-Acoustic Mapping Method for Chladni Patterns in New Media Art Creation

音频生成 | 6.0/10

 · 更新于 2026-10-01 · 约 18 分钟 · 8585 字 阅读 →
论文解读

CORTEG: Foundation Models Enable Cross-Modality Representation Transfer from Scalp to Intracranial Brain Recordings

脑机接口 | 6.5/10

 · 更新于 2026-10-01 · 约 19 分钟 · 9166 字 阅读 →
论文解读

DiffVQE: Hybrid Diffusion Voice Quality Enhancement Under Acoustic Echo and Noise

语音增强 | 6.2/10

 · 更新于 2026-10-01 · 约 17 分钟 · 8037 字 阅读 →
论文解读

Dolphin-CN-Dialect: Where Chinese Dialects Matter

语音识别 | 5.5/10

 · 更新于 2026-10-01 · 约 18 分钟 · 8629 字 阅读 →
论文解读

Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs

音乐生成 | 4.0/10

 · 更新于 2026-10-01 · 约 19 分钟 · 9279 字 阅读 →
论文解读

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

音频事件检测 | 5.8/10

 · 更新于 2026-10-01 · 约 19 分钟 · 9380 字 阅读 →
论文解读

Encoding and Decoding Temporal Signals with Spiking Bandpass Wavelets

音频编码 | 7.0/10

 · 更新于 2026-10-01 · 约 17 分钟 · 8486 字 阅读 →
论文解读

Evaluating the Expressive Appropriateness of Speech in Rich Contexts

语音质量评估 | 7.2/10

 · 更新于 2026-10-01 · 约 20 分钟 · 9625 字 阅读 →
论文解读

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries

音频检索 | 6.0/10

 · 更新于 2026-10-01 · 约 19 分钟 · 9517 字 阅读 →
论文解读

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue

语音对话系统 | 6.0/10

 · 更新于 2026-10-01 · 约 18 分钟 · 8653 字 阅读 →
论文解读

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

语音合成 | 5.5/10

 · 更新于 2026-10-01 · 约 20 分钟 · 9695 字 阅读 →
论文解读

Latent Secret Spin: Keyed Orthogonal Rotations for Blind Speech Watermarking in Anisotropic Latent Spaces

音频水印 | 5.5/10

 · 更新于 2026-10-01 · 约 16 分钟 · 7659 字 阅读 →
论文解读

Low-Cost Detection of Degraded Voice Clones via Source-Output Acoustic Consistency

语音伪造检测 | 5.3/10

 · 更新于 2026-10-01 · 约 16 分钟 · 7610 字 阅读 →
论文解读

Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent Recognition

意图识别 | 7.0/10

 · 更新于 2026-10-01 · 约 19 分钟 · 9463 字 阅读 →
论文解读

Multi-layer attentive probing improves transfer of audio representations for bioacoustics

生物声学 音频分类 | 4.0/10

 · 更新于 2026-10-01 · 约 16 分钟 · 7879 字 阅读 →
论文解读

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

基准测试 | 6.0/10

 · 更新于 2026-10-01 · 约 16 分钟 · 7890 字 阅读 →
论文解读

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

基准测试 | 6.5/10

 · 更新于 2026-10-01 · 约 18 分钟 · 8876 字 阅读 →
论文解读

Online Segmented Beamforming via Dynamic Programming

声源定位 | 6.0/10

 · 更新于 2026-10-01 · 约 14 分钟 · 6970 字 阅读 →
论文解读

PoDAR: Power-Disentangled Audio Representation for Generative Modeling

语音合成 | 7.3/10

 · 更新于 2026-10-01 · 约 18 分钟 · 8991 字 阅读 →