Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade

📄 先听见,再分清:两级级联如何把稀有虎鲸从海噪里捞出来 英文题目:Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade 一句话:针对连续水听器背景压倒信号与部署域偏移,论文用两级 ResNet-18 级联分离检测与生态型判别,在 DCLDE 上将七分类宏平均 F1 做到 0.933 并把稀有 OKW 从 0.842 拉到 0.930,再经主动学习把 Puget Sound 检测 F1 从 0.405 恢复到 0.755,代价是约 5% 漏检无法挽回且连续流召回仍未验证。 标签:#音频分类 | #CNN | #音频事件检测 | #领域适应 评分:6.7/10 | 创新 1.2/2 | 技术严谨 1.2/1.5 | 实验充分 1.2/1.5 | 清晰度 0.8/1 | 影响力 0.8/1.5 | 开源 0/1.5 | 可复现 0.3/0.5 | 工程/实践 1.2/1.5 ...

2026-09-03 · 更新于 2026-09-04 · 6 min · 1120 words

MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries

📄 别只看频谱长什么样:用 19 维力学视角重写环境声音的描述方式 英文题目:MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries 一句话:MADS 把波形当作阻尼振荡与能量传递过程,用 19 维多视角描述子在经典分类器上以一半维度超越 38 维谱摘要基线,并在 ESC 与 MSoS 上取得 81.00%、52.78% 与 67.48% 的峰值准确率,代价是物理代理仍为启发式且未与深度前端对比。 标签:#音频分类 | #集成学习 | #音频事件检测 | #可解释性 评分:5.7/10 | 创新 1.3/2 | 技术严谨 0.9/1.5 | 实验充分 0.9/1.5 | 清晰度 0.8/1 | 影响力 0.8/1.5 | 开源 0/1.5 | 可复现 0.3/0.5 | 工程/实践 0.7/1.5 👥 作者与机构 Utsab Ghosh:ABV-Indian Institute of Information Technology and Management, Gwalior, India Roshni Chakraborty:ABV-Indian Institute of Information Technology and Management, Gwalior, India 💬 毒舌点评 亮点在于把19维物理启发的多视角声学描述子集(Multi-view Acoustic Descriptor Set, MADS)做得紧凑可解释,在经典机器学习流水线上以一半维度压过38维传统基线;短板是所有物理量(有效质量、动能、PDE残差)均为启发式代理且缺乏消融与深度前端对比,所谓物理性更多是命名包装而非可验证的动力学建模。 ...

2026-09-02 · 更新于 2026-09-04 · 5 min · 989 words

TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models

📄 会听不会定时:当大音频模型在 20 分钟里找不到那 2.7 秒 英文题目:TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models 一句话:TAG-Bench 把时序音频定位从 30 秒单区间拉到 20 分钟一对多枚举,用 1750 条人工校验问答与并集 IoU/计数/格式三层度量揭示即使最强的 FireRedAudio 也仅 31.2 mIoU 且所有模型计数准确率不过 13.2% 的系统性短板。 标签:#音频理解 | #音频大模型 | #音频事件检测 | #基准测试 | #长音频处理 评分:6.9/10 | 创新 1.5/2 | 技术严谨 1.2/1.5 | 实验充分 1.1/1.5 | 清晰度 0.8/1 | 影响力 1/1.5 | 开源 0/1.5 | 可复现 0.3/0.5 | 工程/实践 1/1.5 👥 作者与机构 Yuhang Dai:机构信息未在 arXiv HTML 中可靠披露 Xin Shu:机构信息未在 arXiv HTML 中可靠披露 Zengxi Li:机构信息未在 arXiv HTML 中可靠披露 Lei Xie:机构信息未在 arXiv HTML 中可靠披露 Xiangang Li:机构信息未在 arXiv HTML 中可靠披露 Jianwei Yu:机构信息未在 arXiv HTML 中可靠披露 💬 毒舌点评 亮点在于把音频时序定位从 30 秒短片段与单区间检测拉到 20 分钟长音频与一对多枚举的真实痛点,并用 21 个系统的统一自由文本评测撕开了当前大音频语言模型连时间都说不清的遮羞布;短板是所有结论都建立在规则解析的混合度量上,数据与评测代码承诺尚未兑现,且 1750 条查询在 8 个子集上的诊断粒度与统计效力仍显单薄。 ...

2026-09-02 · 更新于 2026-09-04 · 8 min · 1668 words

Self-Supervised Pretext Tasks for Infant Cry Analysis: A Controlled Comparison and a Cautionary Result on Donateacry

📄 在只有 204 个婴儿的世界里,为什么 97.9% 的准确率不值得庆祝 英文题目:Self-Supervised Pretext Tasks for Infant Cry Analysis: A Controlled Comparison and a Cautionary Result on Donateacry 一句话:论文用同一 1.17M 编码器和 115 小时预算横向对比 6 种自监督预训练任务,发现重建类任务以 0.988 AUC 拿下啼哭检测,却在严格按被试划分下让所有原因分类跌至随机附近,并用同一模型复现出文献 97.9% 准确率来证明其来自划分与增强顺序的泄漏。 标签:#音频分类 #自监督学习 #音频事件检测 #对比学习 评分:8.1/10 | 创新 1.4/2 | 技术严谨 1.2/1.5 | 实验充分 1.2/1.5 | 清晰度 0.8/1 | 影响力 1/1.5 | 开源 1/1.5 | 可复现 0.5/0.5 | 工程/实践 1/1.5 👥 作者与机构 Luigi Simeone:Independent researcher 💬 毒舌点评 亮点在于用同一编码器和同一预算把 6 种自监督预训练任务拉到同一起跑线,并用泄漏阶梯把 donateacry 上 90%+ 准确率的泡沫一针戳破,证据链干净到可以当审稿人教材。短板是所有结论都绑在 1.17M 参数小卷积网络和单一小型志愿者数据集上,对大规模模型和临床标签场景的外推只能靠引用他人结果背书。 ...

2026-09-01 · 更新于 2026-09-04 · 7 min · 1412 words

TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

📄 给声音装上刻度:TEMPO 如何让大音频语言模型学会报时 英文题目:TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models 一句话:针对大音频语言模型只会整段描述、不会精确报时的问题,TEMPO 用原子时间戳、墙钟投影与高斯软标签的三件套监督微调加可验证奖励的 GRPO 精炼,在五项时序任务上把多说话人 ASR 的 WER 从 69.7% 压到 43.5%,代价是 0.1 秒固定分辨率与合成音乐主导的训练分布。 标签:#音频事件检测 #后训练 #说话人日志 #强化学习 评分:7.1/10 | 创新 1.5/2 | 技术严谨 1.2/1.5 | 实验充分 1.1/1.5 | 清晰度 0.8/1 | 影响力 1/1.5 | 开源 0/1.5 | 可复现 0.5/0.5 | 工程/实践 1/1.5 👥 作者与机构 Apoorva Kulkarni:University of Maryland, College Park, USA Kaousheik Jayakumar:University of Maryland, College Park, USA Sreyan Ghosh:University of Maryland, College Park, USA Utathya Aich:University of Maryland, College Park, USA Ramani Duraiswami:University of Maryland, College Park, USA Dinesh Manocha:University of Maryland, College Park, USA 💬 毒舌点评 把原子时间戳、分时墙钟投影和高斯软标签打包成可复用的后训练配方,并在5个跨域任务上用单一解码器打通,SFT阶段25个点的WER下降是硬核说服力;短板是GRPO在4个任务上仅0.8至1.4个点的边际精炼却包装成首个统一强化学习,音乐分支困在合成Slakh2100且和弦F1仅6.2%,0.1秒固定分辨率与冻结Whisper-large前端的根本瓶颈并未触及,离专用系统31.4% WER与19.6% DER仍有明显差距。 ...

2026-09-01 · 更新于 2026-09-04 · 9 min · 1816 words

Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations

📄 在茶树里听见白蚁:当 81.5% 的准确率比 98% 更诚实 英文题目:Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations 一句话:面对斯里兰卡高海拔茶园中肉眼难辨的活木白蚁侵染,该研究用贴合树干的声学物联网采集与轻量卷积神经网络做田间筛查,在真实噪声下取得 81.5% 准确率与 0.819 的 ROC-AUC,却以 17% 的漏检率和启发式严重度公式暴露了小样本田间研究的边界。 标签:#音频分类 #CNN #音频事件检测 #工业应用 评分:5.1/10 | 创新 1/2 | 技术严谨 1/1.5 | 实验充分 0.6/1.5 | 清晰度 0.7/1 | 影响力 0.6/1.5 | 开源 0/1.5 | 可复现 0.3/0.5 | 工程/实践 0.9/1.5 👥 作者与机构 D.K.C. Senevirathna:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia A.A.E. Nanayakkara:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia H.M.C.K. Kulathunga:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia J.K.D.P. Nadula:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia R.M. Mapatuna:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia Malithi Nawarathne:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia Jaliya L. Wijayaraja:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia P.D. Senanayake:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia Samitha Vidhanaarachchi:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia Kalpani Manathunga:organization=Sri Lanka Institute of Information Technology, addressline=New Kandy Road, city=Malabe, country=Sri Lanka;organization=School of Molecular and Life Sciences, Curtin University, addressline=Kent Street, city=Bentley, state=Western Australia, country=Australia;organization=University of Kelaniya, addressline=Kandy Road, Dalugama, city=Kelaniya, country=Sri Lanka;organization=Tea Research Institute of Sri Lanka, city=Talawakelle, country=Sri Lanka;organization=Murdoch University, addressline=90 South St, Murdoch, city=Perth, state=Western Australia, country=Australia 💬 毒舌点评 亮点在于把树干贴合式声学采集、云端传输与地理信息系统(Geographic Information System, GIS)可视化串成可用现场流程,直面隐蔽性活木白蚁难以目视发现的痛点。短板是方法学仍停留在 2 层卷积的基线卷积神经网络(Convolutional Neural Network, CNN)加启发式加权严重度公式,实验仅靠 40 株茶树的单次划分支撑结论,难以让人相信其在真实种植园的泛化能力。 ...

2026-08-31 · 更新于 2026-09-04 · 8 min · 1581 words

From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

📄 咳嗽不是噪声:让对话智能体在轮次间听见呼吸 英文题目:From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents 一句话:为解决远程对话中咳嗽被当作噪声丢弃的问题,HealthCUES 用 Qwen3-Omni 把轮次对齐的滚动缓冲与三路并行结构化预测叠加对话感知门控,在 847 段私有对话上做到 93.0% 咳嗽检测 F1 与 340 ms 平均延迟,代价是依赖大模型推理且罕见亚型仍不可靠。 标签:#音频事件检测 #多模态模型 #流式处理 #医疗音频 #音频理解 评分:6.4/10 | 创新 1.4/2 | 技术严谨 1.1/1.5 | 实验充分 0.9/1.5 | 清晰度 0.8/1 | 影响力 0.9/1.5 | 开源 0/1.5 | 可复现 0.1/0.5 | 工程/实践 1.2/1.5 👥 作者与机构 Tanmay Laud:机构信息未在 arXiv HTML 中可靠披露 Herprit Mahal:机构信息未在 arXiv HTML 中可靠披露 Subhabrata Mukherjee:机构信息未在 arXiv HTML 中可靠披露 💬 毒舌点评 亮点在于把被对话系统当噪声丢弃的咳嗽信号做成带亚型与时长的实时管线,并用对话感知门控解决高频误触发这一真实部署痛点,工程闭环完整。短板是核心证据建立在 847 段私有数据与 3 位护士的定性访谈上,对比的 BEATs / PANNs / CoughVID 均非算力与骨干对齐的受控基线,Qwen3-Omni 的提示与约束细节未公开,复现与外推说服力不足。 ...

2026-08-29 · 更新于 2026-09-04 · 3 min · 633 words

Acoustic Echo Control Based on Sound Object Identification for Suppressing Howling Caused by Complicated Acoustic Paths

📄 不估路径,先认出声音:跨终端啸叫为何要改用对象门控 英文题目:Acoustic Echo Control Based on Sound Object Identification for Suppressing Howling Caused by Complicated Acoustic Paths 一句话:这篇论文用“默认静音、确认非同一才放行”的声音对象门控,把难以建模的跨终端回授路径改写为重复对象的本地阻断;它能在特定仿真中压住持续啸叫,却把身份误判与语音可懂度损失暴露为必须正面优化的代价。 标签:#回声消除 #音频事件检测 #实时处理 #工业应用 评分:4.9/10 | 创新 1.1/2 | 严谨 0.8/1.5 | 实验 0.6/1.5 | 清晰 0.9/1 | 影响 0.8/1.5 | 开源 0/1.5 | 复现 0.2/0.5 | 工程 0.5/1.5 💬 毒舌点评 这篇稿最可取的贡献,是把跨终端啸叫从“估计固定路径再相减”改成“默认静音、确认非同一才放行”的对象门控;发送与播放各守相应链段,图 2 也把它嵌入既有通话链路的位置画清楚。它因此给出了值得继续检验的因果单位,而非只把传统 AEC 换名。 但它仍是特定仿真中的可行性演示:没有公开基准、强基线、客观质量指标或部署测量,图 3 的 K–M 还显示误放行与过度静音。持续啸叫受抑不能推出真实会议中的可懂度、鲁棒性或整体体验已经合格;硬静音会切碎目标语音的代价仍需用直接比较来回答。 📌 核心摘要 这篇论文用“默认静音、确认非同一才放行”的声音对象门控,把难以建模的跨终端回授路径改写为重复对象的本地阻断;它能在特定仿真中压住持续啸叫,却把身份误判与语音可懂度损失暴露为必须正面优化的代价。不估计难以建模的端到端路径,而是默认静音,只有当前对象被判为不同于近期缓存对象时才发送或播放。它在发端和收端各设一道门,因此理论上可切断经服务器、编解码和非线性处理绕回来的重复对象。手动静音无法可靠协调多个终端,所以系统需要在信号层面而不是靠用户纪律断开回授。 验证实现只是幅度谱余弦相似度门控,并非训练好的识别模型。在 2 间房、3 个终端仿真中,AEC 约 13 s 收敛后,单讲识别错误相对较少,持续啸叫受抑;未控制条件的啸叫约在 2–3 s 后出现。双讲和 3 人重叠时仍可压住持续啸叫,但误放行会带来短暂回声,保守静音则会破碎目标语音并降低可懂度。 ...

2026-08-27 · 更新于 2026-09-04 · 2 min · 366 words

Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study

📄 Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study 标签:#音频事件检测 #预训练 #模型比较 #工业应用 8.8/10 | 创新 1.5/2 | 严谨 1.5/1.5 | 实验 1.5/1.5 | 清晰 0.9/1 | 影响 1.4/1.5 | 开源 0.2/1.5 | 复现 0.5/0.5 | 工程 1.3/1.5 🔥 8.8/10 | 前25% | 文档类型:应用研究 | 评分置信度:中 | #音频事件检测 | #预训练 | #模型比较 #工业应用 | arxiv 👥 作者与机构 第一作者:Guan-Hua Wen(Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei, Taiwan) 通讯作者:正文未明确标注 作者列表:Guan-Hua Wen、Kuan-Yu Chen(机构:Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei, Taiwan) ...

2026-08-25 · 更新于 2026-09-04 · 3 min · 559 words

Simulation-to-Real First-Break Segmentation for Efficient Inversion in Musculoskeletal Ultrasound Tomography

📄 Simulation-to-Real First-Break Segmentation for Efficient Inversion in Musculoskeletal Ultrasound Tomography 标签:#音频事件检测 #CNN #迁移学习 #医疗音频 9.8/10 | 创新 1.6/2 | 严谨 1.4/1.5 | 实验 1.4/1.5 | 清晰 0.9/1 | 影响 1.3/1.5 | 开源 1.5/1.5 | 复现 0.4/0.5 | 工程 1.3/1.5 🔥 9.8/10 | 前10% | 文档类型:应用研究 | 评分置信度:中 | #音频事件检测 | #CNN | #迁移学习 #医疗音频 | arxiv 👥 作者与机构 第一作者:Yifei Sun(State Key Laboratory of Acoustics and Marine Information、Laboratory of Ultrasonics, Institute of Acoustics, Chinese Academy of Sciences;University of Chinese Academy of Sciences;Université Bourgogne Europe, IMVIA UR 7535) 通讯作者:Yubing Li、Weijun Lin 作者列表:Yifei Sun、Yubing Li、Yannick Benezeth、Stéphanie Bricq、Yunrong Zhang、Lekang Jiang、Chang Su、Ligang Cui、Weijun Lin(机构:Institute of Acoustics, Chinese Academy of Sciences;University of Chinese Academy of Sciences;Université Bourgogne Europe, IMVIA UR 7535;Department of Ultrasound, Peking University Third Hospital) ...

2026-08-25 · 更新于 2026-09-04 · 3 min · 466 words