Chimp-prosody v.1.0.0
v1.0.0 — Initial release First public release of the chimp_prosody pipeline: a tested, modular tool for extracting pitch–intensity coordination metrics and ToBI-style prosodic classifications from chimpanzee field recordings, developed for the Gombe vocal development analysis. What's included core.py — segmentation, ToBI-style pitch-accent/boundary-tone classification, octave-error flagging, and pitch-intensity alignment logic, implemented as pure functions over numpy arrays for easy testing and reuse. io_utils.py — memory-safe chunked audio reading via soundfile, with no intermediate files and no external subprocess calls. Handles long field recordings (30+ minutes) without loading the full file into memory at once. pipeline.py — a single manifest-driven entry point (python -m chimp_prosody.pipeline manifest.csv output.csv) that processes any number of subjects from a CSV manifest and produces one combined output table. Subjects with multiple archived tapes (e.g. Flint) are stitched onto a single continuous timeline. manifest_full.csv — the complete 15-subject manifest from the reported analysis, with correct age classes and age ranges. 22 unit tests (tests/test_core.py) against synthetic signals with known correct answers, covering segmentation, contour classification, octave-error flagging, and pitch-intensity alignment. Example figures (three_calls_example.png, six_bouts_dense_example.png, pitch_tracking_validation_grid.png) showing the pipeline's output on real recordings. LICENSE (MIT, code only — see license text for the note on the underlying archival audio), requirements.txt, and .gitignore. Validation Pitch tracking was checked against a stratified random sample of real segments, visually verified against their spectrograms: 80% accurate, 15% ambiguous (low signal-to-noise), 5% problematic (broadband/ percussive events where pitch tracking is not well defined). Re-running this pipeline on a full subject recording (Gilka) reproduces the originally reported results exactly: 148 segments, 65.2% mid-to-final covariation among non-flagged segments. Known limitations Bitonal pitch-accent categories (H+!H*, H*!H*) are harder to detect reliably near a segment's boundary; MIN_PEAK_PROMINENCE_HZ and MIN_FRAMES_FOR_BITONAL_DETECTION set conservative thresholds to avoid false positives, at some cost to sensitivity. This does not affect the paper's primary result (pitch-intensity covariation), only the secondary ToBI-category distribution analysis. The source audio is not included in this repository (see Usage in the README for how to obtain it from the Gombe chimpanzee archive on Dryad, DOI 10.5061/dryad.5tq80). Citation If you use this code, please cite the associated paper and software release (see README for full reference; DOI to be added once available) along with Praat (Boersma & Weenink 2023), Parselmouth (Jadoul, Thompson & de Boer 2018), ToBI (Beckman & Hirschberg 1994), and the Gombe recording archive (Plooij et al. 2015).