This data article describes PeaDetect, the first curated audio dataset specifically designed for detecting Indian peafowl (Pavo cristatus) vocalizations. The dataset contains 2,950 five-second audio clips (4.1 hours total duration), evenly balanced between peafowl presence (1,475 clips) and absence (1,475 clips) classes. Recordings were sourced from two public repositories: Xeno-Canto (332 original peafowl recordings from India and Sri Lanka, 2003-2024, other bird species vocalizations) and Freesound (environmental sounds and insects from Sri Lankan habitats). All audio was standardized to 44.1 kHz, 16-bit, stereo WAV format and segmented into 5-second clips. Each clip was manually validated by an ornithologist, with inter-annotator agreement assessed on 20% of the dataset (Cohen's kappa = 0.905). To prevent data leakage, source-aware 5-fold cross-validation splits are provided, ensuring all clips from the same original recording remain within the same fold. The dataset includes comprehensive metadata with geographic (87.1% India, 11.9% Sri Lanka, 1.0% other regions), temporal (2003-2024, all seasons), and acoustic specifications. This dataset supports research in bioacoustic monitoring, human-wildlife conflict mitigation, and machine learning for conservation.
No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong
Qilu Normal University · Genelibs Bioinformatics Lab
750 Shunhua Rd, Jinan
2F, Bldg F, University Science Park
Tel: 0531-88819269
Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.
Business Email
E-mail: [email protected]