Publication
Multi-species Mixing for Weakly Supervised SED Under Domain Shift
Novruz Mammadli; Thiago Gouvea; Daniel Sonntag
In: Diedrich Wolter; Gesina Schwalbe (Hrsg.). KI 2026: Advances in Artificial Intelligence. German Conference on Artificial Intelligence (KI-2026), 49th German Conference on Artificial Intelligence, August 11-14, Bremen, Germany, Pages 269-276, Lecture Notes in Computer Science (LNCS), Vol. 16830, ISBN 978-3-032-32335-4, Springer Nature Switzerland, Cham, 2027.
Abstract
Passive Acoustic Monitoring (PAM) is increasingly used for biodiversity monitoring, but training sound event detectors for PAM remains difficult because strong temporal annotations are expensive to obtain. Archival sound libraries provide weakly labelled focal recordings at scale, yet these recordings differ substantially from PAM soundscapes. Focal recordings typically contain one dominant species under relatively clean conditions, whereas PAM recordings often contain overlapping vocalisations from multiple species together with environmental noise. This mismatch limits the transferability of weakly supervised models trained on sound libraries.
This paper investigates whether synthetic multi-species training augmentation can reduce this transfer gap. We build on a weakly supervised sound event detection pipeline using BirdNET embeddings and a linear classifier, and augment the training data by synthetically mixing focal recordings into artificial multi-species soundscapes with randomized temporal overlap and mixture weights. The PAM evaluation data remain unchanged. Experiments on five anuran species show improvements in transfer performance. Compared with training on original focal recordings only, synthetic augmentation increases recall and F1-score at both bag and segment level, while causing a moderate reduction in precision. The best bag-level micro-F1 improves from 0.6792 to 0.7687, and the best segment-level micro-F1 improves from 0.5473 to 0.6120. These results indicate that simulating multi-species co-occurrence during training is a practical way to improve weakly supervised transfer from sound libraries to PAM data.
