论文解读

A 399uW 114.3 dB DR Companding Readout ASIC for MEMS Microphones Employing a Multirate Time-Domain ADC

A 399uW 114.3 dB DR Companding Readout ASIC for MEMS Microphones Employing a Multirate Time-Domain ADC

 · 更新于 2026-09-07 · 约 10 分钟 · 5000 字 阅读 →
论文解读

A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models

多模态模型 | 6.6/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5184 字 阅读 →
论文解读

A Neuromorphic Trigger for Efficient Audio Event Detection

音频事件检测 | 6.2/10

 · 更新于 2026-09-07 · 约 26 分钟 · 12543 字 阅读 →
论文解读

AI-based Cognitive-linguistic Features for Dementia Assessment in Picture Description

语音识别 | 5.8/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5664 字 阅读 →
论文解读

An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages

语音识别 | 5.7/10

 · 更新于 2026-09-07 · 约 25 分钟 · 12144 字 阅读 →
论文解读

Are you speaking my languages? On spoken language adherence in multimodal LLMs

语音识别 | 8/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4530 字 阅读 →
论文解读

Decision-Driven Geosteering Under Uncertainty: A Unified Framework for Sequential Decision Optimization

强化学习 | 7.8/10

 · 更新于 2026-09-07 · 约 40 分钟 · 19751 字 阅读 →
论文解读

Descriptor: Certus Caliber Classification Gunshot Dataset (C3GD)

音频分类 | 5.9/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4488 字 阅读 →
论文解读

DeSRPA: Decoupled Speech Role-Playing Agent via Inference-Time Intervention

语音合成 | 7.3/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5315 字 阅读 →
论文解读

Direction of arrival estimation from distant microphone data using single frequency filtering

语音活动检测 | 7.3/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5615 字 阅读 →
论文解读

ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation

ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation

 · 更新于 2026-09-07 · 约 22 分钟 · 10954 字 阅读 →
论文解读

Embedded Machine Learning for Microcontroller-Class Edge Devices: Data, Feature, Evaluation, and Deployment Pipelines

Embedded Machine Learning for Microcontroller-Class Edge Devices: Data, Feature, Evaluation, and Deployment Pipelines

 · 更新于 2026-09-07 · 约 20 分钟 · 9943 字 阅读 →
论文解读

From Signals to Patterns: Non-Invasive Tuberculosis Detection from Cough Audio using Bandit Weighted Hyperbolic Prototypes

From Signals to Patterns: Non-Invasive Tuberculosis Detection from Cough Audio using Bandit Weighted Hyperbolic Prototypes

 · 更新于 2026-09-07 · 约 12 分钟 · 5966 字 阅读 →
论文解读

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning

语音识别 | 8.5/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5868 字 阅读 →
论文解读

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

语音识别 | 7.5/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4484 字 阅读 →
论文解读

Intelligibility of Speech in Noise: Investigating Contribution of Magnitude and Phase Spectra

Intelligibility of Speech in Noise: Investigating Contribution of Magnitude and Phase Spectra

 · 更新于 2026-09-07 · 约 11 分钟 · 5195 字 阅读 →
论文解读

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

语音合成 | 7.7/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4464 字 阅读 →
论文解读

L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification

说话人验证 | 7.1/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5871 字 阅读 →
论文解读

Learning task-specific subspaces via interventional post-training of speech foundation models

自监督学习 | 6.2/10

 · 更新于 2026-09-07 · 约 21 分钟 · 10293 字 阅读 →
论文解读

MLLP-VRAIN UPV system for the IWSLT 2026 Simultaneous Speech Translation task

语音识别 | 6.9/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7669 字 阅读 →