---
license: cc-by-4.0
pretty_name: "PACC-P: Parallel Acoustic Confound Corpus, Presentation"
language:
  - en
size_categories:
  - 100K<n<1M
task_categories:
  - audio-classification
tags:
  - speech
  - noise
  - reverberation
  - bona-fide
  - robustness
  - anti-spoofing
  - parallel-corpus
---

# PACC-P: Parallel Acoustic Confound Corpus, Presentation

PACC-P is 5,992 bona fide speech clips under 50 presentation conditions: additive noise,
babble and music at five SNRs, simulated rooms and echo, band limiting, pitch and tempo changes,
pitch correction, and two compound cafe scenes passed through a codec. Every condition contains the
same 5,992 clips as PACC-T, and every clip has a per-clip record logged at generation time.

It contains no synthetic or spoofed speech. It is a reference for what the acoustic path between
talker and microphone does to real speech: a bona fide baseline, a robustness benchmark and a
condition-shift test set.

| | |
|---|---|
| Conditions | 50 |
| Clips per condition | 5,992 (same clips in every condition) |
| Files | 299,600 FLAC, plus 5,992 sources in `pacc_base.tar` |
| Duration | 356 h in total |
| Size | 34.3 GB of audio; one condition 0.28-1.15 GB; sources 0.89 GB |
| Download | per condition, so any subset can be fetched on its own |

Companion dataset: **PACC-T** (Telecoms), the same 5,992 sources through 34 codecs, 12 tandem
chains and 5 resample controls. doi:10.5281/zenodo.23026393

## Before you measure

### SNR is set against the clean source's speech level
Every SNR is relative to the ITU-T P.56 active speech level of the clean source clip. Noise level
is the RMS over 20 ms frames within 40 dB of the loudest frame. Each mixed-in layer is scaled
independently against the speech level. Speech is never rescaled on its own: if a finished mix
would clip, the whole mix is scaled so its peak is 0.999 of full scale, and that gain is logged per
clip (`safety_gain_db`). Achieved SNR is measured after mixing and logged.

### Duration changes break sample alignment
`tempo_*` and `pitch_resample_*` change clip length. These conditions are parallel by clip, not by
sample: align in time before any frame-by-frame comparison. All other conditions keep the source
length.

### Band limiting and formants
`bandpass_300_3400` and `lowpass_2400` remove content above 3.4 and 2.4 kHz. LPC formant trackers
configured for wideband speech (Praat's default ceiling is 5500 Hz) fit spurious poles in the empty
band, pulling F2 and F3 low with no error flag; on `lowpass_2400`, F3 is physically absent. Set the
analysis ceiling from content bandwidth, or treat F2 and above as unreliable in these conditions.

### Cafe conditions are also codec conditions
`cafe_voip` (Opus 16 kbps, 48 kHz) and `cafe_cellular` (AMR-WB 23.85 kbps, 16 kHz) mix three layers
and then encode. To separate the scene from the codec, compare with `cafe_noise_snr15` and
`music_instrumental_snr20` here, and with the matching codec and resample control in PACC-T.

### Native bandwidth differs by pool
AMI sources are native 16 kHz (content to 8 kHz); VCTK sources are 48 kHz. Outputs keep the source
rate except the cafe conditions. Compare within a pool, or account for this when pooling.

## Sources

The same 5,992 source clips as PACC-T, shipped once in `pacc_base.tar` with per-source metadata.

| Pool | Clips | Native format | Source corpus |
|---|---|---|---|
| AMI | 2,992 (4.45 h) | 16 kHz WAV, headset microphone | AMI Meeting Corpus: 139 meetings, 157 participants |
| VCTK_mic1 | 1,500 (1.36 h) | 48 kHz FLAC | CSTR VCTK Corpus 0.92, mic1: 109 speakers |
| VCTK_mic2 | 1,500 (1.36 h) | 48 kHz FLAC | CSTR VCTK Corpus 0.92, mic2: 108 speakers |

VCTK mic pairing: 1,485 utterances appear in both mic pools; 15 in each pool are unpaired. Joining
mic1 to mic2 on utterance ID gives 1,485 pairs, not 1,500. All sources are mono 16-bit PCM.
AMI file IDs encode meeting, channel, segment index and start/end time in seconds.

Per-source metadata ships as `sources/metadata.csv` in `pacc_base.tar`: speaker, gender and where
the label came from, VCTK age and utterance, AMI meeting, channel, participant and segment times.

### Gender balance

| Pool | Female | Male | Label source |
|---|---|---|---|
| VCTK_mic1 | 750 | 750 | VCTK speaker metadata |
| VCTK_mic2 | 750 | 750 | VCTK speaker metadata |
| AMI | 1,126 (38%) | 1,866 (62%) | AMI participant metadata |

Sources were selected for gender balance. VCTK is balanced exactly. AMI is not: its clips were
selected using gender inferred from the audio (pitch and apparent vocal-tract length), and those
labels proved wrong for 17% of clips, mostly men labelled as women (441 clips, against 68 the other
way). Checked against AMI's own participant metadata, the AMI pool is 62% male. The `gender`
column gives the metadata label; the inferred label is kept, for transparency only, as
`gender_inferred_not_recommended`. For TS meetings, which have no entry in AMI's
`participants.xml`, sex is taken from the M/F prefix of the participant ID, which agrees with
`participants.xml` for all 189 participants listed there.

8 AMI clips with no usable speech are excluded from all of PACC (`excluded_sources.csv`). Earlier
internal builds also held RAVDESS and CREMA-D clips; they were removed so that PACC is CC BY only.

## Conditions

| Condition | Family | Description | Output rate | Rate control |
|---|---|---|---|---|
| `babble_snr00` | babble | babble noise (6 talkers), 0 dB SNR | as source | - |
| `babble_snr05` | babble | babble noise (6 talkers), 5 dB SNR | as source | - |
| `babble_snr10` | babble | babble noise (6 talkers), 10 dB SNR | as source | - |
| `babble_snr15` | babble | babble noise (6 talkers), 15 dB SNR | as source | - |
| `babble_snr20` | babble | babble noise (6 talkers), 20 dB SNR | as source | - |
| `env_noise_snr00` | scene | environment noise scene, 0 dB SNR | as source | - |
| `env_noise_snr05` | scene | environment noise scene, 5 dB SNR | as source | - |
| `env_noise_snr10` | scene | environment noise scene, 10 dB SNR | as source | - |
| `env_noise_snr15` | scene | environment noise scene, 15 dB SNR | as source | - |
| `env_noise_snr20` | scene | environment noise scene, 20 dB SNR | as source | - |
| `music_instrumental_snr00` | music | instrumental music, 0 dB SNR | as source | - |
| `music_instrumental_snr05` | music | instrumental music, 5 dB SNR | as source | - |
| `music_instrumental_snr10` | music | instrumental music, 10 dB SNR | as source | - |
| `music_instrumental_snr15` | music | instrumental music, 15 dB SNR | as source | - |
| `music_instrumental_snr20` | music | instrumental music, 20 dB SNR | as source | - |
| `music_vocal_snr00` | music | vocal music, 0 dB SNR | as source | - |
| `music_vocal_snr05` | music | vocal music, 5 dB SNR | as source | - |
| `music_vocal_snr10` | music | vocal music, 10 dB SNR | as source | - |
| `music_vocal_snr15` | music | vocal music, 15 dB SNR | as source | - |
| `music_vocal_snr20` | music | vocal music, 20 dB SNR | as source | - |
| `white_noise_snr00` | white noise | white noise, 0 dB SNR | as source | - |
| `white_noise_snr05` | white noise | white noise, 5 dB SNR | as source | - |
| `white_noise_snr10` | white noise | white noise, 10 dB SNR | as source | - |
| `white_noise_snr15` | white noise | white noise, 15 dB SNR | as source | - |
| `white_noise_snr20` | white noise | white noise, 20 dB SNR | as source | - |
| `reverb_small_rt030` | reverb | simulated room reverb (small), RT60 0.30 s | as source | - |
| `reverb_medium_rt060` | reverb | simulated room reverb (medium), RT60 0.60 s | as source | - |
| `reverb_large_rt100` | reverb | simulated room reverb (large), RT60 1.00 s | as source | - |
| `echo_d150ms_m10db` | echo | single echo, 150 ms delay, -10 dB | as source | - |
| `echo_d150ms_m20db` | echo | single echo, 150 ms delay, -20 dB | as source | - |
| `echo_d300ms_m10db` | echo | single echo, 300 ms delay, -10 dB | as source | - |
| `echo_d300ms_m20db` | echo | single echo, 300 ms delay, -20 dB | as source | - |
| `bandpass_300_3400` | filter | bandpass filter, order 4, 300-3400 Hz | as source | - |
| `lowpass_2400` | filter | lowpass filter, order 8, 2400 Hz | as source | - |
| `pitch_psola_dn4` | pitch psola | pitch shift, PSOLA (duration kept), -4 semitones | as source | - |
| `pitch_psola_dn2` | pitch psola | pitch shift, PSOLA (duration kept), -2 semitones | as source | - |
| `pitch_psola_up2` | pitch psola | pitch shift, PSOLA (duration kept), +2 semitones | as source | - |
| `pitch_psola_up4` | pitch psola | pitch shift, PSOLA (duration kept), +4 semitones | as source | - |
| `pitch_resample_dn4` | pitch resample | pitch shift by resampling (duration changes), -4 semitones | as source | - |
| `pitch_resample_dn2` | pitch resample | pitch shift by resampling (duration changes), -2 semitones | as source | - |
| `pitch_resample_up2` | pitch resample | pitch shift by resampling (duration changes), +2 semitones | as source | - |
| `pitch_resample_up4` | pitch resample | pitch shift by resampling (duration changes), +4 semitones | as source | - |
| `autotune_hard` | autotune | pitch correction (autotune), retune 0 ms | as source | - |
| `autotune_natural` | autotune | pitch correction (autotune), retune 50 ms | as source | - |
| `tempo_080` | tempo | tempo change, pitch kept, x0.8 | as source | - |
| `tempo_125` | tempo | tempo change, pitch kept, x1.25 | as source | - |
| `tempo_150` | tempo | tempo change, pitch kept, x1.5 | as source | - |
| `cafe_voip` | compound | cafe scene (babble 15 dB, cafe noise 15 dB, music 20 dB) then Opus 16 kbps | 48 kHz | PACC-T resample_ctrl_48k |
| `cafe_cellular` | compound | cafe scene (babble 15 dB, cafe noise 15 dB, music 20 dB) then AMR-WB 23.85 kbps | 16 kHz | PACC-T resample_ctrl_16k |
| `cafe_noise_snr15` | scene | cafe noise scene, 15 dB SNR | as source | - |

## How the material was made

- **Rooms:** pyroomacoustics 0.10.1, image source plus ray tracing, air absorption, speaker to
  microphone 1.0 m at 1.5 m height, 4 layouts per room (one per clip, chosen by seed). Small
  4.0 x 3.5 x 2.7 m (RT60 0.3 s), medium 7.0 x 5.0 x 3.0 m (0.6 s), large 12.0 x 9.0 x 4.0 m (1.0 s).
  The room impulse responses ship in `pacc-p_room_cache.tar`. Each reverb clip's params row names
  its room variant and the variant's `rir_sha256`; convolving the source with that response (recipe
  in the tarball's README) recreates the stored clip bit for bit.
- **Babble:** 6 English talkers (3 female, 3 male) from MUSAN's LibriVox speech, each at its own
  position in the medium room.
- **Scenes:** FSD50K clips, a continuous bed plus timed events at +6 dB relative to the bed.
  Cafe: tableware, pouring, footsteps, doors, light kitchen sounds, distant traffic, no voices,
  0.4 events/s, in the medium room. Environment: home, office, street and in-car background, no
  voices, 0.3 events/s, no room.
- **Music:** MUSAN music (CC BY and public-domain tracks only), played through a loudspeaker in the
  medium room. The cafe's music layer is drawn exactly like `music_instrumental_snr20`.
- **Echo:** one copy delayed 150 or 300 ms at -10 or -20 dB.
- **Filters:** Butterworth bandpass 300-3400 Hz (design order 4) and lowpass 2400 Hz (order 8).
- **Pitch:** PSOLA keeps formants (Praat 6.1.38 via parselmouth 0.4.7, pitch 75-600 Hz);
  resampling moves formants with F0 (librosa 0.11.0, soxr_hq). Autotune snaps F0 to the nearest
  equal-tempered semitone (A4 = 440 Hz), retune 0 ms (hard) or 50 ms (natural).
- **Tempo:** ffmpeg 8.1 `atempo`, pitch unchanged.
- **Randomness:** every clip has its own seed, sha256 of `20260916|<condition>|<file_id>`, so every
  draw (noise clip, room layout, event timing) is reproducible.
- **Software:** Python 3.11.9, numpy 1.24.4, scipy 1.15.3, soundfile 0.13.1, librosa 0.11.0,
  soxr 1.0.0, praat-parselmouth 0.4.7, pyroomacoustics 0.10.1, ffmpeg 8.1.

## Files

| Tarball | Contents |
|---|---|
| `pacc_base.tar` | the 5,992 source clips (`sources/<pool>/`) and `sources/metadata.csv`, identical in PACC-T and PACC-P |
| `pacc-p_<condition>.tar` | one condition: `pacc-p/<condition>/<pool>/<file_id>.flac`, `params.csv`, `README.txt`, and `attribution.csv` where third-party audio is mixed in |
| `pacc-p_sample.tar` | 14 fixed sources (5 VCTK utterances on both mics, 4 AMI clips) through every condition |
| `pacc-p_room_cache.tar` | the 24 simulated room variants (impulse responses, geometry, measured RT60/DRR) with recreation and hash recipes |

All audio is 16-bit mono FLAC. `SHA256SUMS` and `TARBALL_INDEX.csv` list every tarball;
`pacc-p_manifest.csv` lists every file with its sha256. Tarballs are deterministic.

## Per-clip records

Each condition's `params.csv` has one row per clip, logged when the clip was generated: input and
output sample rate and length, source speech level and peak, every target and achieved SNR, layer
gains, safety gain, the near-floor flag, codec and codec length error for the cafe conditions.

## Known edge cases

See `EDGE_CASES.txt`: the cafe conditions' combined-SNR tolerance (three clips deviate by more than
0.5 dB), rows near the 16-bit quantisation floor, inherited AMI clipping, and the excluded sources.

## Motivation

The design of this corpus was motivated by Delgado et al. (ICASSP 2026), who argue that deepfake
detection must account for how audio is presented through real communication channels, and by Lee
et al. (ICASSP 2026), whose noise-aware multi-LoRA framework highlights real-world noise
conditions. Neither group was involved in producing this dataset.

- H. Delgado, G. Ramondetti, E. Dalmasso, G. Karvitsky, D. Colibro, H. Talib. "On Deepfake Voice
  Detection - It's All in the Presentation." ICASSP 2026. arXiv:2509.26471.
- W. Lee, H. Dinh-Xuan, T.-P. Doan, S. Jung. "Dynamic Noise-Aware Multi LoRA Framework Towards
  Real-World Audio Deepfake Detection." ICASSP 2026.

## Licence and credits

PACC-P is released under **CC BY 4.0** by Christopher Kleingertner (Moonscape Software). It is derived from:

- **CSTR VCTK Corpus 0.92.** J. Yamagishi, C. Veaux, K. MacDonald. University of Edinburgh, CSTR,
  2019. doi:10.7488/ds/2645. CC BY 4.0.
- **AMI Meeting Corpus.** University of Edinburgh et al. https://groups.inf.ed.ac.uk/ami/corpus/.
  CC BY 4.0. J. Carletta (2006), "Announcing the AMI Meeting Corpus", ELRA Newsletter 11(1).
- **MUSAN.** D. Snyder, G. Chen, D. Povey (2015), arXiv:1510.08484. https://www.openslr.org/17/.
  CC BY 4.0; only CC BY and public-domain files used.
- **FSD50K.** E. Fonseca et al. (2022), IEEE/ACM TASLP 30. doi:10.5281/zenodo.4060432. CC BY; only
  CC0 and CC BY clips used.

Third-party clips keep their own licences; per-clip credits are in `attribution.csv` (4,696 clips:
CC0, CC BY 3.0/4.0 and public domain). Please credit these sources alongside PACC-P.

## Citation

```bibtex
@dataset{moonscape_pacc_p_2026,
  author    = {Kleingertner, Christopher and {Moonscape Software}},
  title     = {{PACC-P: Parallel Acoustic Confound Corpus, Presentation}},
  year      = {2026},
  version   = {1.0},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.23026395}
}
```
