论文解读

Dolphin-CN-Dialect: Where Chinese Dialects Matter

语音识别 | 5.5/10

 · 更新于 2026-10-02 · 约 18 分钟 · 8629 字 阅读 →
论文解读

Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs

音乐生成 | 4.0/10

 · 更新于 2026-10-02 · 约 19 分钟 · 9279 字 阅读 →
论文解读

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

音频事件检测 | 5.8/10

 · 更新于 2026-10-02 · 约 19 分钟 · 9380 字 阅读 →
论文解读

Encoding and Decoding Temporal Signals with Spiking Bandpass Wavelets

音频编码 | 7.0/10

 · 更新于 2026-10-02 · 约 17 分钟 · 8486 字 阅读 →
论文解读

Evaluating the Expressive Appropriateness of Speech in Rich Contexts

语音质量评估 | 7.2/10

 · 更新于 2026-10-02 · 约 20 分钟 · 9625 字 阅读 →
论文解读

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries

音频检索 | 6.0/10

 · 更新于 2026-10-02 · 约 19 分钟 · 9517 字 阅读 →
论文解读

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue

语音对话系统 | 6.0/10

 · 更新于 2026-10-02 · 约 18 分钟 · 8653 字 阅读 →
论文解读

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

语音合成 | 5.5/10

 · 更新于 2026-10-02 · 约 20 分钟 · 9695 字 阅读 →
论文解读

Latent Secret Spin: Keyed Orthogonal Rotations for Blind Speech Watermarking in Anisotropic Latent Spaces

音频水印 | 5.5/10

 · 更新于 2026-10-02 · 约 16 分钟 · 7659 字 阅读 →
论文解读

Low-Cost Detection of Degraded Voice Clones via Source-Output Acoustic Consistency

语音伪造检测 | 5.3/10

 · 更新于 2026-10-02 · 约 16 分钟 · 7610 字 阅读 →
论文解读

Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent Recognition

意图识别 | 7.0/10

 · 更新于 2026-10-02 · 约 19 分钟 · 9463 字 阅读 →
论文解读

Multi-layer attentive probing improves transfer of audio representations for bioacoustics

生物声学 音频分类 | 4.0/10

 · 更新于 2026-10-02 · 约 16 分钟 · 7879 字 阅读 →
论文解读

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

基准测试 | 6.0/10

 · 更新于 2026-10-02 · 约 16 分钟 · 7890 字 阅读 →
论文解读

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

基准测试 | 6.5/10

 · 更新于 2026-10-02 · 约 18 分钟 · 8876 字 阅读 →
论文解读

Online Segmented Beamforming via Dynamic Programming

声源定位 | 6.0/10

 · 更新于 2026-10-02 · 约 14 分钟 · 6970 字 阅读 →
论文解读

PoDAR: Power-Disentangled Audio Representation for Generative Modeling

语音合成 | 7.3/10

 · 更新于 2026-10-02 · 约 18 分钟 · 8991 字 阅读 →
论文解读

Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration

音乐生成 | 7.5/10

 · 更新于 2026-10-02 · 约 18 分钟 · 8830 字 阅读 →
论文解读

Probing Cross-modal Information Hubs in Audio-Visual LLMs

模型分析 | 6.5/10

 · 更新于 2026-10-02 · 约 19 分钟 · 9363 字 阅读 →
论文解读

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations

音频深度伪造检测 | 6.0/10

 · 更新于 2026-10-02 · 约 16 分钟 · 7522 字 阅读 →
论文解读

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation

语音增强 | 7.2/10

 · 更新于 2026-10-02 · 约 21 分钟 · 10260 字 阅读 →