Moving Horizon Estimation for Underwater Target Tracking Based on Time-Difference-of-Arrival Measurements

📄 Moving Horizon Estimation for Underwater Target Tracking Based on Time-Difference-of-Arrival Measurements 标签:#声源定位 #模型比较 #鲁棒性 #实时处理 6.5/10 | 创新 0.9/2 | 严谨 1.1/1.5 | 实验 1/1.5 | 清晰 0.9/1 | 影响 0.2/1.5 | 开源 1.2/1.5 | 复现 0.3/0.5 | 工程 0.9/1.5 ✅ 6.5/10 | 前50% | 文档类型:方法研究 | 评分置信度:中 | #声源定位 | #模型比较 | #鲁棒性 #实时处理 | arxiv 👥 作者与机构 第一作者:Anton Tolstonogov(Institute for Systems and Robotics (ISR), LARSyS, Instituto Superior Técnico (IST), University of Lisbon, Portugal) 通讯作者:未显式标注;论文给出作者邮箱 作者列表:Anton Tolstonogov、David Cabecinhas、Pedro Batista、Antonio Pascoal(均为 Institute for Systems and Robotics (ISR), LARSyS, Instituto Superior Técnico (IST), University of Lisbon, Portugal) 💡 毒舌点评 论文在稀疏 TDoA 测量、离群值和较远初始化条件下,确实展示了 MHE 相比朴素 EKF 的显著鲁棒性优势;运行时间测试也表明该方法在 1 Hz 声学更新周期内具备实时可行性。但整体证据仍停留在 2D 合成仿真,基线只有 EKF,未与 UKF、粒子滤波、鲁棒 EKF 或带约束的滤波方案比较,也没有真实水下声学数据验证。精度 RMSE 的统计波动和离群值生成机制都未交代,落地工程证据偏弱。作为面向语音/音乐/音频领域读者的分析,其直接影响力更低。 ...

2026-08-18 · 更新于 2026-09-04 · 4 min · 719 words

Numerical and perceptual validity of synthetic Head-Related Transfer Functions at scale

📄 Numerical and perceptual validity of synthetic Head-Related Transfer Functions at scale 标签:#声源定位 #生成模型 #空间音频 #模型评估 #数据集 7.6/10 | 创新 1.2/2 | 严谨 1.2/1.5 | 实验 1/1.5 | 清晰 0.8/1 | 影响 1/1.5 | 开源 1.2/1.5 | 复现 0.3/0.5 | 工程 0.9/1.5 ✅ 7.6/10 | 前25% | 文档类型:应用研究 | 评分置信度:高 | #声源定位 | #生成模型 | #空间音频 #模型评估 | arxiv 👥 作者与机构 第一作者:Katarina C. Poole(Imperial College London, Dyson School of Design Engineering) 通讯作者:Katarina C. Poole(Imperial College London, Dyson School of Design Engineering) 作者列表:Katarina C. Poole(Imperial College London, Dyson School of Design Engineering)、Lorenzo Picinali(Imperial College London, Dyson School of Design Engineering) 💡 毒舌点评 这是一篇在空间音频领域做得相当扎实的大规模评估工作,用 200 名受试者的 HRTF 数据把“合成 HRTF 是否可替代实测 HRTF”这个问题问到了数值、模型和行为三个层面,结论(合成 HRTF 在行为上追平实测、KEMAR 显著更差)对个性化空间音频的规模化落地有直接价值。但论文的硬伤也很明显:数值发现的偏差集中在低仰角后方,行为错误却全挤在前后中线,这种错位暴露了现有数值/模型度量与真实感知之间的断层——作者诚实承认了这一点,并在讨论中列举了几条可能的改进方向(如基于临界带的感知加权、整合中耳/内耳听觉模型、更大规模行为数据集),但没有任何一条被实际验证或给出预测指标,使得结论止步于“哪里对不上”而非“该怎么解决”。此外,随机 HRTF 意外低 LSD 的反直觉结果被一带而过,Barumerli 模型在空间相关性上的系统性失败(对 KEMAR 的极角精度甚至显著负相关)也没有被深究结构原因,这些未消化的事实削弱了论文在感知有效性评估层面的方法论贡献。 ...

2026-08-18 · 更新于 2026-09-04 · 4 min · 846 words

Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions

📄 Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions 标签:#声源定位 #CNN #模型评估 6.4/10 | 创新 1.3/2 | 严谨 1.2/1.5 | 实验 1/1.5 | 清晰 0.8/1 | 影响 1/1.5 | 开源 0/1.5 | 复现 0.3/0.5 | 工程 0.8/1.5 ✅ 6.4/10 | 前50% | 文档类型:方法研究 | 评分置信度:中 | #声源定位 | #CNN | #模型评估 | arxiv 👥 作者与机构 第一作者:Mikko Heikkinen(未说明) 通讯作者:未说明 作者列表:Mikko Heikkinen(未说明)、Archontis Politis(未说明)、Konstantinos Drossos(未说明)、Tuomas Virtanen(未说明) 💡 毒舌点评 用复值ATF替代几何坐标作为阵列条件,是解决非自由场、带壳体散射设备泛化的正确方向,手机型阵列上的泛化结果也确有价值。但全篇只有仿真实验,缺少与Neural-SRP、GI-DOAEnet、PhaseCoder等同赛道方法的对比,也没有证明"ATF优于坐标"的消融;CoordConv、跨通道共享注意力等关键设计同样没有组件级证据;不公开代码和数据让"array-generic"更像一个精致演示,而不是可验证的结论。 📌 核心摘要 论文提出利用复值方向性阵列传递函数(ATF)作为阵列元数据,配合多通道信号谱图进行神经网络声源到达方向估计,以解决大多数深度学习DoA方法只能用于训练阵列、难以泛化到未见设备的问题。核心网络由信号卷积编码器、ATF直接性卷积编码器、跨注意力融合和DoA解码器组成,输出多源笛卡尔向量;所有二维卷积均替换为CoordConv。区别于用麦克风坐标作为元数据的方法,ATF能描述非自由场、非全向、含壳体散射的实际设备。2D单源任务中,方法F1@5°=0.75、LE=4.0°,低于MUSIC(0.97/1.5°),但LE明显优于FCGA(9.8°);3D多源任务中随着源数增加性能衰减比MUSIC更平缓,且mobile-like阵列泛化与单一阵列训练差距不大。意义在于提供了一条将物理阵列模型嵌入神经DoA估计、增强设备泛化性的路径。局限主要包括仅仿真验证、未与近年前沿阵列泛化DNN对比,且未提供代码或数据。 下文表格只列论文正文明确给出的代表性设置与结果;未报告的基线数字不作推断。 🔗 开源详情 代码:论文中未提及代码链接。论文仅给出 arXiv ID 2608.09425v1,未提供 GitHub、HuggingFace、ModelScope、项目主页或 Demo 地址。 ...

2026-08-11 · 更新于 2026-09-04 · 2 min · 376 words

Physics-Informed Learning for Robust Acoustic Localization with Calibrated Uncertainty

📄 Physics-Informed Learning for Robust Acoustic Localization with Calibrated Uncertainty 标签:#声源定位 #Transformer #鲁棒性 #多通道 6.2/10 | 创新 1.3/2 | 严谨 1.1/1.5 | 实验 0.9/1.5 | 清晰 0.8/1 | 影响 1/1.5 | 开源 0/1.5 | 复现 0.1/0.5 | 工程 1/1.5 ✅ 6.2/10 | 前50% | 文档类型:方法研究 | 评分置信度:中 | #声源定位 | #Transformer | #鲁棒性 #多通道 | arxiv 👥 作者与机构 第一作者:Jennifer N. Kampe(University of Jyväskylä, Department of Biological and Environmental Science;Duke University, Department of Statistical Science) 通讯作者:未说明 作者列表: Jennifer N. Kampe(University of Jyväskylä, Department of Biological and Environmental Science;Duke University, Department of Statistical Science) Changwoo J. Lee(Duke University, Department of Statistical Science) Xin Shen(Duke University, Department of Statistical Science) Ari Lehtiö(University of Jyväskylä, Digital Services) Sandro von Brandenburg(University of Jyväskylä, Digital Services) Ossi Nokelainen(University of Jyväskylä, Department of Biological and Environmental Science;University of Jyväskylä, Open Science Centre) David B. Dunson(Duke University, Department of Statistical Science) Otso Ovaskainen(University of Jyväskylä, Department of Biological and Environmental Science) 💡 毒舌点评 用“保留物理求解器 + 学习校正 + 两层物理门控 + GDOP 缩放不确定性”的思路来抑制双曲定位的灾难性长尾,诊断清楚、设计务实,这是论文最大的亮点。但真实实验只有一个冰冻湖站点、六个已知扬声器位置,森林场景全部是仿真,且没有给出完整混合门控在森林仿真中的直接结果;基线只有经典双曲求解器,未与任何现代深度学习、score-based 或 conformal 声源定位方法对比。整体说服力仍未达到顶会主接收准。 ...

2026-08-11 · 更新于 2026-09-04 · 3 min · 619 words

语音/音乐/音频论文速递 2026-08-11

语音/音乐/音频论文速递 2026-08-11 共分析 40 篇论文 ⚡ 今日概览 📥 抓取 40 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #音视频理解 4篇 ████ #音频生成 4篇 ████ #音乐生成 3篇 ███ #音频字幕生成 3篇 ███ #声源定位 2篇 ██ #语音合成 2篇 ██ #语音情感识别 2篇 ██ #语音编码 2篇 ██ 📊 论文评分排行榜(40 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 AVCap: Reinforcing Audio-Video Joint Caption with Detai 8.4分 前25% 方法研究 #音频字幕生成 🥈 Deferred Audio Pruning with Local Audio-Visual Dynamics 8.4分 前25% 方法研究 #模型剪枝 🥉 PACE: A Playback-Aligned Context Engine for LLM-Based F 8.0分 前25% 系统技术报告 #语音交互 4. REFRAMED: Towards Realistic Audio Description Generatio 8.0分 前25% 数据集与基准 #音频字幕生成 5. SraVaani 1.0: Scaling Inclusive Speech Recognition for 7.9分 前25% 模型报告 #语音识别 6. AudioMap: Cloze-and-Choice Reinforcement Learning for T 7.9分 前25% 方法研究 #音频字幕生成 7. SAMOT: State-Aware Step Modulation and Optimal Transpor 7.7分 前25% 方法研究 #音视频理解 8. ReLMCodec: Designing Predictable Speech Tokens from Pre 7.7分 前25% 方法研究 #语音编码 9. VoxZip: Semantic-Anchored Temporal KV Cache Compression 7.7分 前25% 方法研究 #音频理解 10. Steering dense music retrieval with open-vocabulary con 7.7分 前25% 方法研究 #音乐检索 11. VIOLET: High-Fidelity Violin Synthesis with Techniques 7.6分 前25% 方法研究 #音乐生成 12. SonicWeave: Chunk-Routed Mixture-of-Experts for Unified 7.6分 前25% 模型报告 #音频生成 13. Beyond Reconstruction: Full-Context Generative DiT for 7.5分 前25% 系统技术报告 #音乐生成 14. From Inaudible Inputs to Model Failures: Low-Frequency 7.4分 前50% 方法研究 #音频理解 15. Beyond Piano: Cross-Instrument MIDI Velocity Estimation 7.3分 前50% 方法研究 #音乐理解 16. omni-macos: On-Device Omni-Modal Search on Apple Silico 7.2分 前50% 系统技术报告 #音频检索 17. SCoPE: Training-Free Audio-Visual Event Perception via 7.1分 前50% 方法研究 #音视频理解 18. BAMU: Bitstream-Aware Marginal-Utility Allocation for F 7.1分 前50% 方法研究 #语音编码 19. Dramarrator: Object-Based Audio Editing for Audio Drama 7.0分 前50% 系统技术报告 #音频生成 20. CuteTTS: Efficient and High-Quality Speech Synthesis vi 7.0分 前50% 系统技术报告 #语音合成 21. RAG-Audio: Retrieval-Augmented Generation for Faithful 6.9分 前50% 方法研究 #音频生成 22. Listen, See and Track: Spatio-Temporal Audio-Visual Sou 6.7分 前50% 数据集与基准 #音视频问答 23. From Speech to Interaction: Analyzing Multimodal System 6.6分 前50% 系统技术报告 #音视频语音识别 24. MADBench: A Benchmark for Modality-Aware Audio Deepfake 6.6分 前50% 数据集与基准 #音视频理解 25. A Unifying Perspective on Audio Generative Modeling: La 6.5分 前50% 理论研究 #音频生成 26. Multilingual Emotion Neurons in Large Audio-Language Mo 6.5分 前50% 方法研究 #语音情感识别 27. DAVE: A Decoupled Audio-Visual Enhancement Framework fo 6.5分 前50% 系统技术报告 #音视频语音分离 28. Structured Phonological Representations for Audio-Artic 6.5分 前50% 方法研究 #语音属性识别 29. CtrlSpeech: Coarse-to-Fine Control for Expressive Speec 6.4分 前50% 方法研究 #语音合成 30. Neural Array-Generic Direction-of-Arrival Estimation Ex 6.4分 前50% 方法研究 #声源定位 31. AI-Guided Learning: Research on Knowledge and Skill Acq 6.3分 前50% 系统技术报告 #音视频理解 32. Speaker Role and Language Diarization for Analyzing Mul 6.3分 前50% 应用研究 #说话人日志 33. MusicLayout: Explicit Structural Planning for Controlla 6.3分 前50% 系统技术报告 #音乐生成 34. Mitigating Over-Suppression in Speech Enhancement via I 6.2分 前50% 方法研究 #语音增强 35. Physics-Informed Learning for Robust Acoustic Localizat 6.2分 前50% 方法研究 #声源定位 36. EmoS: A Theory-Grounded Framework for Evaluating and Al 6.2分 前50% 数据集与基准 #语音情感识别 37. Dynamic Clustering for Cross-Segment Permutation Alignm 6.2分 前50% 方法研究 #语音分离 38. The Voiceprint Fallacy: Why Voices Are Not Unique Biome 6.0分 前50% 综述 #说话人验证 39. Investigating Multimodal Informativity under Different 5.9分 前50% 方法研究 #多模态模型 40. Machine-Learning-Based Diagnostic Framework for Passive 5.6分 前50% 应用研究 #音频分类 📋 论文列表 🥇 AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward 8.4/10 | 创新 1.5/2 | 严谨 1/1.5 | 实验 1.2/1.5 | 清晰 0.8/1 | 影响 1/1.5 | 开源 1.2/1.5 | 复现 0.5/0.5 | 工程 1.2/1.5 ...

2026-08-11 · 更新于 2026-09-04 · 33 min · 6905 words

Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence

📄 Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence 标签:#声源定位 #对比学习 #音频理解 #Transformer #模型评估 6.8/10 | 创新 1.4/2 | 严谨 1/1.5 | 实验 0.9/1.5 | 清晰 0.8/1 | 影响 1/1.5 | 开源 0.5/1.5 | 复现 0.3/0.5 | 工程 0.9/1.5 ✅ 6.8/10 | 前50% | 文档类型:方法研究 | 评分置信度:中 | #声源定位 | #对比学习 | #音频理解 #Transformer | arxiv 👥 作者与机构 共同第一作者:Han Hu、Dongheng Lin(论文中以脚注标记Equal contribution) 其他作者:Yuqi Hou、Haotian Li、Hyung Jin Chang、Jianbo Jiao 机构:The MIx Group, School of Computer Science, University of Birmingham, UK(所有作者均属该机构) 通讯作者:未披露 ...

2026-08-07 · 更新于 2026-09-04 · 3 min · 473 words

语音/音乐/音频论文速递 2026-08-07

语音/音乐/音频论文速递 2026-08-07 共分析 21 篇论文 ⚡ 今日概览 📥 抓取 21 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #语音属性识别 2篇 ██ #语音识别 2篇 ██ #音乐生成 2篇 ██ #音视频生成 2篇 ██ #音频生成 2篇 ██ #声源定位 1篇 █ #语音交互 1篇 █ #语音伪造检测 1篇 █ 📊 论文评分排行榜(21 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 AffectDF: The Most Comprehensive Benchmark for Speech D 8.4分 前25% 数据集与基准 #语音伪造检测 🥈 LILAC: An Idempotent Neural Speech Codec 8.1分 前25% 方法研究 #语音编码 🥉 KVAE: Family of Tokenizers for Multimodal Generative Mo 8.1分 前25% 模型报告 #音频编码 4. Decolonizing Linguistic Policies in Automated Speech Re 7.8分 前25% 理论研究 #语音识别 5. Audio-to-Score Transcription using Pre-trained Features 7.7分 前25% 数据集与基准 #音乐转录 6. Diff2Mix: Controllable Music Mixing via Diffusion Model 7.5分 前25% 方法研究 #音频生成 7. Vorch-Streamer: Extending Human Audio-Visual Generation 7.4分 前50% 方法研究 #音视频生成 8. EG-VAE: A Unified Framework for Electric Guitar Tone Tr 7.3分 前50% 方法研究 #音频生成 9. Explicit and Stable Pseudospectral Time-Domain Method f 7.2分 前50% 方法研究 #音频理解 10. FormBharo: Designing and Evaluating a Voice Agent for C 7.0分 前50% 系统技术报告 #语音交互 11. Whence the Voice? Self-supervised Dual-source Audio-Vis 6.8分 前50% 方法研究 #声源定位 12. Bias Analysis of L2 Speaking Assessment Systems Using C 6.7分 前50% 方法研究 #语音质量评估 13. Diff-Symbo: Text-Controlled Long-Duration Symbolic Musi 6.5分 前50% 方法研究 #音乐生成 14. Rethinking Automatic Music Mixing as Sequential Stem Bl 6.5分 前50% 方法研究 #音乐生成 15. How to Recognize New Words: A Comparison Between Contex 6.4分 前50% 方法研究 #语音识别 16. Beyond Residual Connections: Manifold-Constrained Hyper 6.0分 前50% 方法研究 #说话人验证 17. A Study of ASR Adaptation and Representation Dimensiona 5.7分 前50% 方法研究 #语音情感识别 18. PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Hea 5.6分 前50% 方法研究 #音视频生成 19. The interface of intonation and lexical tone: Boundary 5.6分 前50% 综述 #语音属性识别 20. C\(^3\)PO: Evaluating Cross-Modal Composition and Counter 5.5分 前50% 数据集与基准 #音视频问答 21. ECHO: A Locally-Deployable Agentic Health Assistant wit 5.5分 前50% 系统技术报告 #语音属性识别 📋 论文列表 🥇 AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks 8.4/10 | 创新 1.4/2 | 严谨 1.2/1.5 | 实验 1.2/1.5 | 清晰 0.8/1 | 影响 1.3/1.5 | 开源 1.2/1.5 | 复现 0.3/0.5 | 工程 1/1.5 ...

2026-08-07 · 更新于 2026-09-04 · 17 min · 3544 words

An End-to-End Workflow for Fin Whale Song Detection, Note Characterization, and Localization with Distributed Acoustic Sensing

📄 An End-to-End Workflow for Fin Whale Song Detection, Note Characterization, and Localization with Distributed Acoustic Sensing 标签:#音频事件检测 #声源定位 #多通道 #低资源 #音频理解 7.5/10 | 创新 1.2/2 | 严谨 1/1.5 | 实验 0.6/1.5 | 清晰 0.9/1 | 影响 0.8/1.5 | 开源 1.5/1.5 | 复现 0.3/0.5 | 工程 1.2/1.5 ✅ 7.5/10 | 前25% | 文档类型:系统技术报告 | 评分置信度:中 | #音频事件检测 | #声源定位 | #多通道 #低资源 | arxiv 👥 作者与机构 第一作者:Dídac Diego-Tortosa(Institut de Ciències del Mar (ICM-CSIC), Barcelona) 通讯作者:未明确标注 作者列表:Dídac Diego-Tortosa(ICM-CSIC)、Miriam Romagosa(ICM-CSIC)、Arantza Ugalde(ICM-CSIC)、Hugo Latorre(ICM-CSIC)、Sergi Ventosa(ICM-CSIC)、Jose Enrique García(IFIC, UV-CSIC)、Antonio Villaseñor(ICM-CSIC) 💡 毒舌点评 亮点是将经典信号处理方法(KVP小波检测、DBSCAN时空聚类、层次合并、双曲线拟合)巧妙组合成一套无需标注训练的生物声学监测管道,在极度稀疏、高噪声的DAS环境下实现了可用的检测精度,工程落地价值明显。短板是完全没有与任何基线方法(哪怕是最简单的能量检测器或频谱互相关)的对比,0.744的召回率缺乏参照系;定位方案承认了双曲线的左右模糊性但未提出缓解策略,整体定位精度缺乏量化指标,仅靠两个轨迹示例支撑"可推断移动方向"的结论,说服力较弱。 ...

2026-08-04 · 更新于 2026-09-04 · 2 min · 323 words

Embodied Passive Aeroacoustic Perception Enables Relative Sensing and Pursuit Between Aerial Robots

📄 Embodied Passive Aeroacoustic Perception Enables Relative Sensing and Pursuit Between Aerial Robots 标签:#声源定位 #CNN #音频分类 #多通道 #实时处理 7.9/10 | 创新 1.3/2 | 严谨 1/1.5 | 实验 0.8/1.5 | 清晰 0.8/1 | 影响 0.8/1.5 | 开源 1.5/1.5 | 复现 0.5/0.5 | 工程 1.2/1.5 ✅ 7.9/10 | 前25% | 文档类型:系统技术报告 | 评分置信度:高 | #声源定位 | #CNN | #音频分类 #多通道 | arxiv 👥 作者与机构 第一作者:Yanbaihui Liu(杜克大学机械工程与材料科学系) 通讯作者:Boyuan Chen(杜克大学机械工程与材料科学系、电气与计算机工程系、计算机科学系) 作者列表:Yanbaihui Liu(杜克大学机械工程与材料科学系)、Ravi Prakash(杜克大学机械工程与材料科学系)、Li-Yu Lo(杜克大学机械工程与材料科学系)、Nils Roede(杜克大学机械工程与材料科学系)、Boyuan Chen(杜克大学机械工程与材料科学系、电气与计算机工程系、计算机科学系) 💡 毒舌点评 这项工作是机器人声学感知交叉领域一次充满想象力的探索,将多旋翼恼人的自噪声点石成金般地转化为无需协作即可感知的相对信号源,户外双机追踪的完整系统演示令人眼前一亮。然而,诗意归诗意,当方位角误差高达31.2°时,一个所谓的实时追踪系统,与主流视觉/激光雷达方案在更公平条件下的对标却被规避了。作者将方案精确框定在“互补性场景”,但这种精明的定位也掩盖了其核心性能粗糙的现实。论文展示了一个很有潜力的方向,但目前看来,它离成为可部署的实用系统仍然相去甚远。 ...

2026-08-04 · 更新于 2026-09-04 · 2 min · 401 words

语音/音乐/音频论文速递 2026-08-04

语音/音乐/音频论文速递 2026-08-04 共分析 39 篇论文。 ⚡ 今日概览 📥 抓取 39 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #语音合成 4篇 ████ #语音识别 3篇 ███ #音视频生成 3篇 ███ #语音属性识别 2篇 ██ #说话人验证 2篇 ██ #音视频理解 2篇 ██ #音频分类 2篇 ██ #音频理解 2篇 ██ 📊 论文评分排行榜(39 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dat 8.3分 前25% 数据集与基准 #语音识别 🥈 FATE: Frame-Level Audio-Visual Temporal Embedding 8.3分 前25% 方法研究 #音视频理解 🥉 Allocation Before Ranking: Decoupled Token Compression 8.1分 前25% 方法研究 #音视频理解 4. Embodied Passive Aeroacoustic Perception Enables Relati 7.9分 前25% 系统技术报告 #声源定位 5. Interpretable MEG Decoding of Perceived Speech: Cortica 7.9分 前25% 方法研究 #音频检索 6. Beyond Prompt Adherence: Auditing Attribute-Level Voice 7.8分 前25% 方法研究 #语音合成 7. Hidden-Domain Routing for All-Type Audio Deepfake Detec 7.6分 前25% 系统技术报告 #领域适应 8. An End-to-End Workflow for Fin Whale Song Detection, No 7.5分 前25% 系统技术报告 #音频事件检测 9. SAGE: Switch-Aware EEG-Guided Soft Gating for Target Sp 7.3分 前50% 方法研究 #语音分离 10. Experience-Calibrated Contrastive Decoding for Mitigati 7.2分 前50% 方法研究 #语音合成 11. DRONEAUDIONET: Noise Suppression for Drone Audition-bas 7.2分 前50% 方法研究 #音频分离 12. Multi-Backbone Self-Supervised Ensembles for Audio Deep 7.2分 前50% 系统技术报告 #音频伪造检测 13. Hear, Invoke, and Understand: A Skill-Calling Multimoda 7.2分 前50% 方法研究 #音频理解 14. CultureVidBench: Benchmarking Cultural Understanding in 7.2分 前50% 数据集与基准 #音视频生成 15. SwanTale: Unified Multi-Speaker Speech and Audio Genera 7.1分 前50% 系统技术报告 #音频生成 16. Scene2Sound: Auditory-Grounded Soundscape Generation fo 7.0分 前50% 方法研究 #音频生成 17. The Learning Objective Governs Perceptual Narrowing: A 7.0分 前50% 方法研究 #语音属性识别 18. Discriminative Axis, Not Data Volume: What a Contrastiv 6.9分 前50% 方法研究 #语音属性识别 19. P-MUSE: Prompt-MIDI-Optional Model for Unified Instrume 6.9分 前50% 系统技术报告 #音频理解 20. AcoustiTrace: When Plausible Sound Violates Physics 6.9分 前50% 数据集与基准 #音视频生成 21. Gecko: Fast Private Inference via Secure Public Encoder 6.9分 前50% 系统技术报告 #迁移学习 22. LeapTalk: Breaking the Latency-Quality Trade-off in Tal 6.8分 前50% 方法研究 #音视频生成 23. Separate-and-Detect: Unified Drum Transcription and Ste 6.8分 前50% 方法研究 #音乐转录 24. Domain-Specific Evaluation of Text-to-Speech Systems: A 6.8分 前50% 数据集与基准 #语音合成 25. JoyAI-Talker: Full-Duplex Speech Interactive Large Mode 6.5分 前50% 系统技术报告 #语音合成 26. REIMU: Efficient Heterogeneous Hierarchical Reasoning f 6.4分 前50% 方法研究 #语音伪造检测 27. The Role of Disfluencies in Speech Translation 6.4分 前50% 数据集与基准 #语音翻译 28. SCOPE: Entanglement Frontier Escape for Source-Free Cla 6.3分 前50% 方法研究 #说话人验证 29. Uncertainty-Aware Crossmodal Fusion for Classification 6.3分 前50% 方法研究 #音频分类 30. Can Foundation Models Hear What Made That Sound? A Tier 6.3分 前50% 数据集与基准 #音频分类 31. QR-Erase: Efficient Subspace-Based Machine Unlearning w 6.0分 前50% 方法研究 #说话人验证 32. SGAD: A State-Guided Adaptive Decision Framework for Ro 6.0分 前50% 方法研究 #Transformer 33. Normal-Anchored First-Order Model-Agnostic Meta-Learnin 5.9分 前50% 方法研究 #语音识别 34. AnyBand: Unified Multi-Bandwidth Speech Extension via F 5.9分 前50% 方法研究 #语音超分 35. Sounding Canvas: Embedding Algorithms in Networked, Sen 5.7分 前50% 系统技术报告 #RNN 36. UOT-IR: Structured Routing of High-Polyphony Symbolic M 5.5分 前50% 方法研究 #音乐理解 37. Embodied Empathy: A Multimodal AR and LLM-Powered Syste 5.4分 后50% 应用研究 #音频交互 38. Analyzing Speech Condition Effects in Dysarthric ASR: A 5.0分 后50% 方法研究 #语音识别 39. Homebot: A Personal AI Agent for Conversational Home As 3.8分 后50% 系统技术报告 #语音交互 📋 论文列表 🥇 SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces 8.3/10 | 创新 1.5/2 | 严谨 1/1.5 | 实验 1/1.5 | 清晰 0.8/1 | 影响 1/1.5 | 开源 1.5/1.5 | 复现 0.3/0.5 | 工程 1.2/1.5 ...

2026-08-04 · 更新于 2026-09-04 · 28 min · 5820 words