No Free Lunch from Audio Pretraining in Bioacoustics: A Benchmark Study of Embeddings

cs.AI updates on arXiv.org 08月15日

No Free Lunch from Audio Pretraining in Bioacoustics: A Benchmark Study of Embeddings

文章探讨了生物声学特征提取中，深度学习模型调优的重要性。研究发现，未经微调的音频预训练模型在任务表现上不如经过微调的模型，且调优后的模型在处理背景声音时表现更佳。

arXiv:2508.10230v1 Announce Type: cross Abstract: Bioacoustics, the study of animal sounds, offers a non-invasive method to monitor ecosystems. Extracting embeddings from audio-pretrained deep learning (DL) models without fine-tuning has become popular for obtaining bioacoustic features for tasks. However, a recent benchmark study reveals that while fine-tuned audio-pretrained VGG and transformer models achieve state-of-the-art performance in some tasks, they fail in others. This study benchmarks 11 DL models on the same tasks by reducing their learned embeddings' dimensionality and evaluating them through clustering. We found that audio-pretrained DL models 1) without fine-tuning even underperform fine-tuned AlexNet, 2) both with and without fine-tuning fail to separate the background from labeled sounds, but ResNet does, and 3) outperform other models when fewer background sounds are included during fine-tuning. This study underscores the necessity of fine-tuning audio-pretrained models and checking the embeddings after fine-tuning. Our codes are available: https://github.com/NeuroscienceAI/Audio\_Embeddings

Fish AI Reader

AI辅助创作，多种专业模板，深度分析，高质量内容生成。从观点提取到深度思考，FishAI为您提供全方位的创作支持。新版本引入自定义参数，让您的创作更加个性化和精准。

FishAI

鱼阅，AI 时代的下一个智能信息助手，助你摆脱信息焦虑

联系邮箱 441953276@qq.com

相关标签

生物声学深度学习模型调优音频预训练特征提取

相关文章

Import AI 363: ByteDance’s 10k GPU training run; PPO vs REINFORCE; and generative everything

xLSTM: Enhancing Long Short-Term Memory LSTM Capabilities for Advanced Language Modeling and Beyond

Optimizing Graph Neural Network Training with DiskGNN: A Leap Toward Efficient Large-Scale Learning

V-JEPA, AI Reasoning from a Non-Generative Architecture with Mido Assran - #677

Transformers On Large-Scale Graphs with Bayan Bruss - #641

Towards Improved Transfer Learning with Hugo Larochelle - #631

Stable Diffusion & Generative AI with Emad Mostaque - #604

Engineering Production NLP Systems at T-Mobile with Heather Nolis - #600

Transformers for Tabular Data at Capital One with Bayan Bruss - #591

100x Improvements in Deep Learning Performance with Sparsity, w/ Subutai Ahmad - #562