论文解读

Sandboxed Coding Agents are Competitive Omni-modal Task Solvers

强化学习 | 7.9/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5961 字 阅读 →
论文解读

SegTune: Structured and Fine-Grained Control for Song Generation

音乐生成 | 8.5/10

 · 更新于 2026-09-30 · 约 14 分钟 · 6788 字 阅读 →
论文解读

SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment

语音识别 | 7/10

 · 更新于 2026-09-30 · 约 14 分钟 · 6619 字 阅读 →
论文解读

SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling

音乐生成 | 8.6/10

 · 更新于 2026-09-30 · 约 26 分钟 · 12623 字 阅读 →
论文解读

SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription

语音识别 | 8.8/10

 · 更新于 2026-09-30 · 约 14 分钟 · 6767 字 阅读 →
论文解读

SpeakerCard-1M: An Evidence-Grounded Speaker Card Corpus for In-the-Wild Speaker Verification

说话人验证 | 7.4/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6193 字 阅读 →
论文解读

Speech Emotion Recognition using Attention-based LSTM-Network with Residual Connection

语音情感识别 | 7.5/10

 · 更新于 2026-09-30 · 约 10 分钟 · 4523 字 阅读 →
论文解读

Stable Hybrid Cross-Attention Fusion for Audio-Visual Event Recognition

自监督学习 | 6.7/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5944 字 阅读 →
论文解读

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

语音识别 | 8.7/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5528 字 阅读 →
论文解读

The DeepSpeak-Agentic Dataset

语音合成 | 8.7/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6420 字 阅读 →
论文解读

Tonal parsimony in chord-sequence analysis: combining modulation cost and tonal vocabulary

音乐信息检索 | 8.1/10

 · 更新于 2026-09-30 · 约 9 分钟 · 4455 字 阅读 →
论文解读

Wavelet as Tokenizer: Preliminary Results on a Shared Wavelet Token Schema for Natural Signals

多模态模型 | 5.4/10

 · 更新于 2026-09-30 · 约 11 分钟 · 5061 字 阅读 →
论文解读

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

语音合成 | 9.2/10

 · 更新于 2026-09-30 · 约 12 分钟 · 5546 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-03

共分析 40 篇语音/AI 论文

 · 更新于 2026-09-30 · 约 123 分钟 · 61249 字 阅读 →
论文解读

A 1000-hour EEG-EMG-audio dataset of Japanese speech production

A 1000-hour EEG-EMG-audio dataset of Japanese speech production

 · 更新于 2026-09-30 · 约 18 分钟 · 8937 字 阅读 →
论文解读

A Lightweight Slot-Attention Framework for Multi-Instrument Multi-Pitch Estimation

音乐信息检索 | 6.7/10

 · 更新于 2026-09-30 · 约 13 分钟 · 6340 字 阅读 →
论文解读

Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning

语音增强 | 7.1/10

 · 更新于 2026-09-30 · 约 17 分钟 · 8092 字 阅读 →
论文解读

AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course

AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course

 · 更新于 2026-09-30 · 约 13 分钟 · 6238 字 阅读 →
论文解读

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

多模态模型 | 7/10

 · 更新于 2026-09-30 · 约 15 分钟 · 7267 字 阅读 →
论文解读

Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty

语音识别 | 5.5/10

 · 更新于 2026-09-30 · 约 10 分钟 · 4920 字 阅读 →