论文解读

Text-Dependent Speaker Verification (TdSV) Challenge 2024: Team Naive System Report

说话人验证 | 5.5/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9548 字 阅读 →
论文解读

Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8776 字 阅读 →
论文解读

WARDEN: Endangered Indigenous Language Transcription and Translation with 6 Hours of Training Data

语音识别 | 7.0/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7464 字 阅读 →
论文解读

A Semi-Supervised Framework for Speech Confidence Detection using Whisper

语音自信度检测 | 6.5/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9646 字 阅读 →
论文解读

jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition

多模态检索 | 7.5/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8909 字 阅读 →
论文解读

Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization

语音识别 说话人日志 | 5.5/10

 · 更新于 2026-09-24 · 约 21 分钟 · 10162 字 阅读 →
论文解读

CORTEG: Foundation Models Enable Cross-Modality Representation Transfer from Scalp to Intracranial Brain Recordings

脑机接口 | 6.5/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9166 字 阅读 →
论文解读

Empirical Study of Pop and Jazz Mix Ratios for Genre-Adaptive Chord Generation

音乐生成 | 7.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5401 字 阅读 →
论文解读

The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail

语音识别 | 8.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6296 字 阅读 →
论文解读

Spoken Language Identification with Pre-trained Models and Margin Loss

说话人识别 | 7.5/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3725 字 阅读 →
论文解读

Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?

音乐生成 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4931 字 阅读 →
论文解读

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

音频分类 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4990 字 阅读 →
论文解读

From Birdsong to Rumbles: Classifying Elephant Calls with Out-of-Species Embeddings

音频分类 | 6.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6394 字 阅读 →
论文解读

MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation

语音增强 | 6.0/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5552 字 阅读 →
论文解读

OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging

模型比较 | 7.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4240 字 阅读 →
论文解读

SumRA: Parameter Efficient Fine-tuning with Singular Value Decomposition and Summed Orthogonal Basis

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5875 字 阅读 →
论文解读

Beyond the Baseband: Adaptive Multi-Band Encoding for Full-Spectrum Bioacoustics Classification

生物声学 | 7.0/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6321 字 阅读 →
论文解读

Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device

语音生物标志物 | 7.0/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6353 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5558 字 阅读 →
论文解读

A Parameter-Efficient Multi-Scale Convolutional Adapter for Synthetic Speech Detection

A Parameter-Efficient Multi-Scale Convolutional Adapter for Synthetic Speech Detection

 · 更新于 2026-09-24 · 约 9 分钟 · 4431 字 阅读 →