Speaker: Vicente Jose Escarigo Miranda

Abstract: Anuran vocalizations are valuable for biodiversity monitoring, but their automatic analysis is challenging because multiple species often vocalize simultaneously and class distributions are highly imbalanced. AnuraSet provides expert annotations for 42 Neotropical species together with additional unlabeled recordings from U-AnuraSet. This work extends the original multi-label classification task to sound event detection (SED), identifying active species and locating their vocalizations over time by adapting a semi-supervised Mean Teacher framework. The adapted baseline achieves a Polyphonic Sound Event Detection Score, scenario 1 (PSDS1), of 0.198 and an event-based macro F1 of 0.326. To address the characteristics of AnuraSet, the system incorporates focal loss, weighted oversampling, data augmentation, pretrained BEATs and BirdNET embeddings, Squeeze-and-Excitation (SE) attention, and Sound Event Bounding Boxes (SEBBs) for event decoding. Performance is evaluated mainly with PSDS1, supported by event- and segment-based macro F1. The final system combines these improvements, increasing PSDS1 from 0.198 to 0.386 and event-based macro F1 from 0.326 to 0.470. These results show that the Mean Teacher framework can be successfully transferred to the challenging bioacoustic conditions of AnuraSet, establishing a strong SED baseline for the dataset.