Sensing Bone-Conducted Speech with Earbuds

📄 耳塞在震动什么:骨导语音的低频直线与单轴传感器的容差 英文题目:Sensing Bone-Conducted Speech with Earbuds 一句话:这项工作用两款耳塞的同步测量标定佩戴者语音诱发的机壳振动,证明高功率集中在 400 Hz 以下且沿耳道口内外直线振动,使有利安装的单轴传感器仅损失 0.7 至 1.5 dB,而高频与机型差异仍需多轴与更低本底来覆盖。 标签:#语音增强 | #端到端 | #鲁棒性 | #工业应用 评分:6.3/10 | 创新 1.2/2 | 技术严谨 1.2/1.5 | 实验充分 1/1.5 | 清晰度 0.8/1 | 影响力 0.8/1.5 | 开源 0/1.5 | 可复现 0.3/0.5 | 工程/实践 1/1.5 👥 作者与机构 Christoph Weyer:Institute of Communication Systems (IKS), RWTH Aachen University, Aachen, Germany Peter Jax:Institute of Communication Systems (IKS), RWTH Aachen University, Aachen, Germany 💬 毒舌点评 把耳塞壳体振动这个长期凭经验用的模态做了频域加空间的联合标定,并给出单轴替代三轴的定量衰减依据,工程指向清晰。短板是17人中14男3女且多为欧洲背景、仅两款耳塞、无下游增强或检测验证,高频结论受三轴器件本底限制,向通用形态与真实噪声外推仍单薄。 ...

2026-09-03 · 更新于 2026-09-04 · 5 min · 853 words

Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations

📄 在茶树里听见白蚁:当 81.5% 的准确率比 98% 更诚实 英文题目:Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations 一句话:面对斯里兰卡高海拔茶园中肉眼难辨的活木白蚁侵染,该研究用贴合树干的声学物联网采集与轻量卷积神经网络做田间筛查,在真实噪声下取得 81.5% 准确率与 0.819 的 ROC-AUC,却以 17% 的漏检率和启发式严重度公式暴露了小样本田间研究的边界。 标签:#音频分类 #CNN #音频事件检测 #工业应用 评分:5.1/10 | 创新 1/2 | 技术严谨 1/1.5 | 实验充分 0.6/1.5 | 清晰度 0.7/1 | 影响力 0.6/1.5 | 开源 0/1.5 | 可复现 0.3/0.5 | 工程/实践 0.9/1.5 👥 作者与机构 D.K.C. Senevirathna:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia A.A.E. Nanayakkara:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia H.M.C.K. Kulathunga:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia J.K.D.P. Nadula:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia R.M. Mapatuna:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia Malithi Nawarathne:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia Jaliya L. Wijayaraja:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia P.D. Senanayake:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia Samitha Vidhanaarachchi:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia Kalpani Manathunga:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia 💬 毒舌点评 亮点在于把树干贴合式声学采集、云端传输与地理信息系统(Geographic Information System, GIS)可视化串成可用现场流程,直面隐蔽性活木白蚁难以目视发现的痛点。短板是方法学仍停留在 2 层卷积的基线卷积神经网络(Convolutional Neural Network, CNN)加启发式加权严重度公式,实验仅靠 40 株茶树的单次划分支撑结论,难以让人相信其在真实种植园的泛化能力。 ...

2026-08-31 · 更新于 2026-09-04 · 8 min · 1581 words

Acoustic Echo Control Based on Sound Object Identification for Suppressing Howling Caused by Complicated Acoustic Paths

📄 不估路径,先认出声音:跨终端啸叫为何要改用对象门控 英文题目:Acoustic Echo Control Based on Sound Object Identification for Suppressing Howling Caused by Complicated Acoustic Paths 一句话:这篇论文用“默认静音、确认非同一才放行”的声音对象门控,把难以建模的跨终端回授路径改写为重复对象的本地阻断;它能在特定仿真中压住持续啸叫,却把身份误判与语音可懂度损失暴露为必须正面优化的代价。 标签:#回声消除 #音频事件检测 #实时处理 #工业应用 评分:4.9/10 | 创新 1.1/2 | 严谨 0.8/1.5 | 实验 0.6/1.5 | 清晰 0.9/1 | 影响 0.8/1.5 | 开源 0/1.5 | 复现 0.2/0.5 | 工程 0.5/1.5 💬 毒舌点评 这篇稿最可取的贡献,是把跨终端啸叫从“估计固定路径再相减”改成“默认静音、确认非同一才放行”的对象门控;发送与播放各守相应链段,图 2 也把它嵌入既有通话链路的位置画清楚。它因此给出了值得继续检验的因果单位,而非只把传统 AEC 换名。 但它仍是特定仿真中的可行性演示:没有公开基准、强基线、客观质量指标或部署测量,图 3 的 K–M 还显示误放行与过度静音。持续啸叫受抑不能推出真实会议中的可懂度、鲁棒性或整体体验已经合格;硬静音会切碎目标语音的代价仍需用直接比较来回答。 📌 核心摘要 这篇论文用“默认静音、确认非同一才放行”的声音对象门控,把难以建模的跨终端回授路径改写为重复对象的本地阻断;它能在特定仿真中压住持续啸叫,却把身份误判与语音可懂度损失暴露为必须正面优化的代价。不估计难以建模的端到端路径,而是默认静音,只有当前对象被判为不同于近期缓存对象时才发送或播放。它在发端和收端各设一道门,因此理论上可切断经服务器、编解码和非线性处理绕回来的重复对象。手动静音无法可靠协调多个终端,所以系统需要在信号层面而不是靠用户纪律断开回授。 验证实现只是幅度谱余弦相似度门控,并非训练好的识别模型。在 2 间房、3 个终端仿真中,AEC 约 13 s 收敛后,单讲识别错误相对较少,持续啸叫受抑;未控制条件的啸叫约在 2–3 s 后出现。双讲和 3 人重叠时仍可压住持续啸叫,但误放行会带来短暂回声,保守静音则会破碎目标语音并降低可懂度。 ...

2026-08-27 · 更新于 2026-09-04 · 2 min · 366 words

Domain-Adaptive ASR for Telephony AI Agents: Fine-tuning Canary Flash Models for Enterprise Contact Center Applications

📄 把电话声道和业务词汇分开付账:Canary 的两次适配实验 英文题目:Domain-Adaptive ASR for Telephony AI Agents: Fine-tuning Canary Flash Models for Enterprise Contact Center Applications 一句话:这份报告的可取之处不在于把 Canary 包装成万能电话 ASR,而在于用真实通话与可控增强补电话声道、再以姓名地址阶段换取业务词汇准确率,并把这笔专化回退明确暴露出来。 标签:#语音识别 #领域适应 #实时处理 #工业应用 评分:7.9/10 | 创新 1.3/2 | 严谨 1.1/1.5 | 实验 1.2/1.5 | 清晰 0.9/1 | 影响 1.1/1.5 | 开源 1/1.5 | 复现 0.2/0.5 | 工程 1.1/1.5 💬 毒舌点评 这篇报告最扎实的优点,是把真实通话、提示录音和电话化增强拆成可理解的采集分工,并把姓名地址的 3.78% 与一般电话的 10.89% 一起呈现。姓名地址实验最重要的不是 3.78% 本身,而是它与 10.89% 一起呈现专化代价;加上 180M 的 RTFx 601.8,读者能看到面向生产的真实取舍。 但证据仍像内部上线复盘:内部参与者、内部泰语电话集和单张 A100 不能支撑普遍部署结论。延迟只在单一硬件和 batch size 下测量,生产并发、排队和成本并没有被这张 RTFx 表覆盖;又缺少 real-call、mu-law 与噪声各自贡献的消融,生产上仍可能需要模型路由或混合训练。 ...

2026-08-27 · 更新于 2026-09-04 · 3 min · 428 words

Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate

📄 Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate 标签:#语音交互 #数据集 #模型评估 #工业应用 7.3/10 | 创新 1.4/2 | 严谨 1.2/1.5 | 实验 1.2/1.5 | 清晰 0.9/1 | 影响 1.3/1.5 | 开源 0/1.5 | 复现 0.3/0.5 | 工程 1/1.5 ✅ 7.3/10 | 前50% | 文档类型:应用研究 | 评分置信度:中 | #语音交互 | #数据集 | #模型评估 #工业应用 | arxiv 👥 作者与机构 第一作者:Ethan Traister(scam.ai) 通讯作者:Simiao Ren(论文以 benren@scam.ai 标注通讯邮箱) 作者列表:Ethan Traister、Ankit Raj、Jiaqi Gan、Xingyu Shen、Tyler Wu、Yuchen Zhou、Tommy Duong、Kidus Zewde、Siying Chen、Simiao Ren(机构:scam.ai) ...

2026-08-26 · 更新于 2026-09-04 · 3 min · 496 words

Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study

📄 Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study 标签:#音频事件检测 #预训练 #模型比较 #工业应用 8.8/10 | 创新 1.5/2 | 严谨 1.5/1.5 | 实验 1.5/1.5 | 清晰 0.9/1 | 影响 1.4/1.5 | 开源 0.2/1.5 | 复现 0.5/0.5 | 工程 1.3/1.5 🔥 8.8/10 | 前25% | 文档类型:应用研究 | 评分置信度:中 | #音频事件检测 | #预训练 | #模型比较 #工业应用 | arxiv 👥 作者与机构 第一作者:Guan-Hua Wen(Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei, Taiwan) 通讯作者:正文未明确标注 作者列表:Guan-Hua Wen、Kuan-Yu Chen(机构:Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei, Taiwan) ...

2026-08-25 · 更新于 2026-09-04 · 3 min · 559 words

Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring

📄 Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring 标签:#音频分类 #端到端 #工业应用 #高效推理 7.3/10 | 创新 1.4/2 | 严谨 1.2/1.5 | 实验 1.1/1.5 | 清晰 0.8/1 | 影响 0.8/1.5 | 开源 0.5/1.5 | 复现 0.3/0.5 | 工程 1.2/1.5 ✅ 7.3/10 | 前50% | 文档类型:应用研究 | 评分置信度:中 | #音频分类 | #端到端 | #工业应用 #高效推理 | arxiv 👥 作者与机构 第一作者:Steven C. Nesbit(机构未说明) 通讯作者:未说明 作者列表:Steven C. Nesbit、Victor M. Vergara、Michael A. Felix、Evan T. Kain、Luis R. García Carrillo、Gerd J. Kunde、Andrew T. Sornborger(机构信息未在当前正文中完整说明) ...

2026-08-20 · 更新于 2026-09-04 · 3 min · 637 words

Mitigating Spectral Bias in Neural Operators for Underwater Transmission Loss Prediction

📄 Mitigating Spectral Bias in Neural Operators for Underwater Transmission Loss Prediction 标签:#音频理解 #端到端 #高效推理 #工业应用 7.3/10 | 创新 1.4/2 | 严谨 1.2/1.5 | 实验 1.1/1.5 | 清晰 0.8/1 | 影响 0.8/1.5 | 开源 0.5/1.5 | 复现 0.3/0.5 | 工程 1.2/1.5 ✅ 7.3/10 | 前50% | 文档类型:方法研究 | 评分置信度:中 | #音频理解 | #端到端 | #高效推理 #工业应用 | arxiv 👥 作者与机构 第一作者:Yifan Sun(机构未说明) 通讯作者:未说明 作者列表:Yifan Sun、Shikai Fang、Chao Zhang、Lei Cheng、Jianlong Li、Peter Gerstoft(机构信息未在当前正文中完整说明) 💡 毒舌点评 两阶段神经运营商的水声传输损失恢复在固定设定内做得干净,但泛化声明需要打折:真值全部来自同一南海数据集、固定 200Hz 频率与固定网格区域,跨频率、跨季节、跨海底参数的表现一概未知。对照实验也只有 FNO 与 Hankel-FNO 加不加精修两种组合,缺少 U-Net 等直接竞争者和端到端联合训练的参照,两阶段设计的各部分贡献因此难以量化。第二阶段精修器不再接收原始声速与地形输入的设计选择也未做消融。 ...

2026-08-20 · 更新于 2026-09-04 · 3 min · 506 words

语音/音乐/音频论文速递 2026-08-20

语音/音乐/音频论文速递 2026-08-20 共分析 20 篇论文 ⚡ 今日概览 📥 抓取 20 篇 → 🔬 深度分析完成 🏷️ 热门方向 方向 数量 分布 #多模态模型 3篇 ███ #音频理解 3篇 ███ #语音交互 2篇 ██ #语音识别 2篇 ██ #音乐理解 2篇 ██ #音频分类 2篇 ██ #音频生成 2篇 ██ #语音合成 1篇 █ 📊 论文评分排行榜(20 篇,按分数降序) 排名 论文 总分 分档 文档类型 主任务 🥇 Alignment Is All You Need: Instruction-Free Training fo 8.0分 前25% 模型报告 #音频理解 🥈 Aslema at NADI 2026: Augmentation through Fewshot for S 8.0分 前25% 系统技术报告 #语音交互 🥉 X2Streaming-TTS: Causal Token-Level Text-to-Speech from 7.9分 前25% 模型报告 #语音合成 4. Generalized Audio-Driven Synthesis of Precise Drummer M 7.8分 前25% 模型报告 #音视频生成 5. Computational Features for Symbolic Melody Analysis 7.6分 前25% 系统技术报告 #音乐理解 6. Geometric Iterative Retrieval for Neural Audio Codec Re 7.6分 前25% 方法研究 #音频编码 7. Evaluating Music Context Preservation: A Multi-facet Fr 7.5分 前25% 数据集与基准 #音乐理解 8. Nine Emotion Centroids: A Label-Free Valence Axis That 7.5分 前25% 方法研究 #音频理解 9. Accurate Decoding of Natural Sentences from Non-Invasiv 7.5分 前25% 应用研究 #语音识别 10. Mitigating Spectral Bias in Neural Operators for Underw 7.3分 前50% 方法研究 #音频理解 11. FM Synthesizer Audio-Parameter Shared Embeddings 7.3分 前50% 方法研究 #音频生成 12. Low-Power, Neuromorphic, Acoustic Anomaly Detection for 7.3分 前50% 应用研究 #音频分类 13. Finetuning Strategies for Querying Sounds by Vocal Imit 7.3分 前50% 系统技术报告 #音频检索 14. Understanding Multilingual Medical ASR Adaptation Throu 7.2分 前50% 应用研究 #语音识别 15. Multimodal Rapport Estimation in Real-World HRI 7.0分 前50% 应用研究 #多模态模型 16. StocksTalk: A Voice-Enabled Conversational Agent for St 6.7分 前50% 系统技术报告 #语音交互 17. ChiroEcho: extending automated bat vocalisation classif 6.7分 前50% 方法研究 #音频分类 18. Pedagogical AI in Mental Health: A Tri-Stream Fine-Tune 6.7分 前50% 应用研究 #多模态模型 19. Sounds Uncertain: Exploring the Affective Aspects of So 6.7分 前50% 应用研究 #音频生成 20. Large Language Models in Mental Health: A Systematic Re 6.0分 前50% 综述 #多模态模型 📋 论文列表 🥇 Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models 8.0/10 | 创新 1.7/2 | 严谨 1.2/1.5 | 实验 1.2/1.5 | 清晰 0.8/1 | 影响 1.2/1.5 | 开源 0.5/1.5 | 复现 0.3/0.5 | 工程 1.1/1.5 ...

2026-08-20 · 更新于 2026-09-04 · 17 min · 3435 words

The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report

📄 The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report 标签:#语音伪造检测 #自监督学习 #工业应用 #模型评估 6.6/10 | 创新 1.4/2 | 严谨 1/1.5 | 实验 0.6/1.5 | 清晰 0.9/1 | 影响 1.3/1.5 | 开源 0/1.5 | 复现 0.1/0.5 | 工程 1.3/1.5 ✅ 6.6/10 | 前50% | 文档类型:系统技术报告 | 评分置信度:高 | #语音伪造检测 | #自监督学习 | #工业应用 #模型评估 | arxiv 👥 作者与机构 第一作者:Anton Firc(按作者顺序首位) 通讯作者:未说明 作者列表:Anton Firc、Kamil Malinka、Vojtěch Staněk、Miroslav Hlaváček、Marek Bartoň 机构:Security@FIT, Brno University of Technology, Czech Republic;Phonexia, Brno, Czech Republic 项目资助:Czech Ministry of the Interior,SECTECH 安全研究项目 “Tools to Combat Voice DeepFakes”(项目号 VB02000060);Brno University of Technology 内部项目 FIT-S-26-9011 用户组织:Czech Police 利益冲突:作者包含 Phonexia 员工,Phonexia 是该技术的商业厂商 生成式 AI 使用声明:作者在语言润色和文本修改中使用了 Google Gemini、ChatGPT 和 Grammarly 注:论文未将每位作者与机构逐一对应;正文中以 [University] 和 [Company] 指代合作双方,地址信息显示机构为 Brno University of Technology 与 Phonexia。 ...

2026-08-19 · 更新于 2026-09-04 · 5 min · 1017 words