Datasets
Parallel Acoustic Confound Corpus
The Parallel Acoustic Confound Corpus (PACC) is one fixed set of 5,992 bona fide speech clips, delivered under many controlled conditions. Every condition contains the same clips, so any clip can be compared with itself before and after the change. What differs between two versions of a clip is the condition, not the talker.
PACC contains no synthetic or spoofed speech. It is a reference for what real-world channels do to real speech.
The two corpora
PACC-T: Telecoms
The same 5,992 bona fide speech clips passed through 34 codecs, 12 tandem chains and 5 resample controls.
51 conditions · CC BY 4.0 · v1.0
PACC-P: Presentation
The same 5,992 bona fide speech clips under 50 presentation conditions: noise, babble, music, rooms, echo, band limiting, pitch and tempo changes.
50 conditions · CC BY 4.0 · v1.0
One set of sources, two families of change
Both corpora ship the same 5,992 source clips, drawn from the AMI Meeting Corpus and the CSTR VCTK Corpus. PACC-T sends them through telecom channels: codecs, tandem chains and resample-only controls. PACC-P renders them under presentation conditions: noise, babble, music, rooms, echo, band limiting, and pitch and tempo changes. Because the sources are identical, effects from the two families can be compared clip for clip. MARE, Moonscape's speech measurement engine, has been run on every condition; see the MARE page.
Full source attribution and licences are on the Methods & Provenance page.