论文解读

Ecologically-Constrained Task Arithmetic for Multi-Taxa Bioacoustic Classifiers Without Shared Data

生物声学 | 8.0/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6468 字 阅读 →
论文解读

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework

说话头伪造检测 | 7.5/10

 · 更新于 2026-09-10 · 约 14 分钟 · 6781 字 阅读 →
论文解读

Learning Generalizable Action Representations via Pre-training AEMG

生物声学 | 7.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5844 字 阅读 →
论文解读

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model

语音对话系统 | 7.5/10

 · 更新于 2026-09-10 · 约 26 分钟 · 12846 字 阅读 →
论文解读

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection

语音生物标志物 | 8.0/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5797 字 阅读 →
论文解读

PHALAR: Phasors for Learned Musical Audio Representations

音乐信息检索 | 8.0/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5654 字 阅读 →
论文解读

Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4325 字 阅读 →
论文解读

ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval

音频检索 | 7.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5435 字 阅读 →
论文解读

Smart Passive Acoustic Monitoring: Embedding a Classifier on AudioMoth Microcontroller

生物声学 | 7.5/10

 · 更新于 2026-09-10 · 约 6 分钟 · 2580 字 阅读 →
论文解读

Stage Light is Sequence$^2$: Multi-Light Control via Imitation Learning

音乐信息检索 | 7.5/10

 · 更新于 2026-09-10 · 约 16 分钟 · 7610 字 阅读 →
论文解读

The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail

语音识别 | 8.5/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6496 字 阅读 →
论文解读

Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts

多模态模型 | 7.0/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6070 字 阅读 →
论文解读

Towards Open World Sound Event Detection

音频事件检测 | 8.5/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6369 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-06

共分析 23 篇语音/AI 论文

 · 更新于 2026-09-10 · 约 81 分钟 · 40463 字 阅读 →
论文解读

Artificial intelligence language technologies in multilingual healthcare: Grand challenges ahead

多语言健康沟通 | 6.5/10

 · 更新于 2026-09-10 · 约 5 分钟 · 2457 字 阅读 →
论文解读

BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios

基准测试 | 7.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3895 字 阅读 →
论文解读

Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning

音视频 | 7.5/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6444 字 阅读 →
论文解读

Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models

语音识别 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4665 字 阅读 →
论文解读

Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks

大语言模型 | 8.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5782 字 阅读 →
论文解读

HARMES: A Multi-Modal Dataset for Wearable Human Activity Recognition with Motion, Environmental Sensing and Sound

音频分类 | 8.0/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6251 字 阅读 →