The same 5,992 bona fide speech clips under 50 presentation conditions: noise, babble, music, rooms, echo, band limiting, pitch and tempo changes.

Version
1.0
Conditions
50
Clips per condition
5,992
Licence
CC BY 4.0
DOI
10.5281/zenodo.23026395
Hugging Face: pending Zenodo: pending Companion: PACC-T

What this is

Between a talker and a microphone, the acoustic path changes speech: background noise, other voices, music, room reverberation, echo and band limiting. Pitch and tempo processing changes it further. PACC-P isolates those effects. It takes one fixed set of 5,992 bona fide speech clips, the same clips as PACC-T, and renders every clip under 50 presentation conditions.

Because every condition holds the same clips, any clip can be compared with itself across all 50. Whatever changes between two versions of the same clip is the presentation, not the talker.

Units you can check

Conditions are named for their physical parameters: SNR in dB, room RT60 in seconds, echo delay in milliseconds and level in dB, pitch in semitones. Every SNR is defined against the ITU-T P.56 active speech level of the clean source, and the achieved SNR after mixing is logged for every clip.

Two things to know before comparing

Tempo and resample-based pitch conditions change clip length, so they are parallel by clip rather than by sample and need aligning in time first. Band-limited conditions can mislead formant trackers that assume wideband speech. The datacard below explains both and how to handle them.

Reproducibility

Every clip has its own seed derived from the condition and file ID, so every random draw (noise clip, room layout, event timing) can be regenerated. The simulated room impulse responses are published with their hashes, and convolving a source with the named response recreates the stored reverb clip bit for bit. Every file is listed with its SHA-256 in the manifest.

What it does not contain

No synthetic or spoofed speech. PACC-P is a bona fide reference: a baseline, a robustness benchmark and a condition-shift test set.

Datacard

Snapshot: PACC-P v1.0 datacard · SHA-256: d63308c7384fbd08789117010c5d7e11ce25c42c44492afd540c9f02ef98c81e · Download the exact file

PACC-P is 5,992 bona fide speech clips under 50 presentation conditions: additive noise, babble and music at five SNRs, simulated rooms and echo, band limiting, pitch and tempo changes, pitch correction, and two compound cafe scenes passed through a codec. Every condition contains the same 5,992 clips as PACC-T, and every clip has a per-clip record logged at generation time.

It contains no synthetic or spoofed speech. It is a reference for what the acoustic path between talker and microphone does to real speech: a bona fide baseline, a robustness benchmark and a condition-shift test set.

Conditions 50
Clips per condition 5,992 (same clips in every condition)
Files 299,600 FLAC, plus 5,992 sources in pacc_base.tar
Duration 356 h in total
Size 34.3 GB of audio; one condition 0.28-1.15 GB; sources 0.89 GB
Download per condition, so any subset can be fetched on its own

Companion dataset: PACC-T (Telecoms), the same 5,992 sources through 34 codecs, 12 tandem chains and 5 resample controls. doi:10.5281/zenodo.23026393

Before you measure

SNR is set against the clean source's speech level

Every SNR is relative to the ITU-T P.56 active speech level of the clean source clip. Noise level is the RMS over 20 ms frames within 40 dB of the loudest frame. Each mixed-in layer is scaled independently against the speech level. Speech is never rescaled on its own: if a finished mix would clip, the whole mix is scaled so its peak is 0.999 of full scale, and that gain is logged per clip (safety_gain_db). Achieved SNR is measured after mixing and logged.

Duration changes break sample alignment

tempo_* and pitch_resample_* change clip length. These conditions are parallel by clip, not by sample: align in time before any frame-by-frame comparison. All other conditions keep the source length.

Band limiting and formants

bandpass_300_3400 and lowpass_2400 remove content above 3.4 and 2.4 kHz. LPC formant trackers configured for wideband speech (Praat's default ceiling is 5500 Hz) fit spurious poles in the empty band, pulling F2 and F3 low with no error flag; on lowpass_2400, F3 is physically absent. Set the analysis ceiling from content bandwidth, or treat F2 and above as unreliable in these conditions.

Cafe conditions are also codec conditions

cafe_voip (Opus 16 kbps, 48 kHz) and cafe_cellular (AMR-WB 23.85 kbps, 16 kHz) mix three layers and then encode. To separate the scene from the codec, compare with cafe_noise_snr15 and music_instrumental_snr20 here, and with the matching codec and resample control in PACC-T.

Native bandwidth differs by pool

AMI sources are native 16 kHz (content to 8 kHz); VCTK sources are 48 kHz. Outputs keep the source rate except the cafe conditions. Compare within a pool, or account for this when pooling.

Sources

The same 5,992 source clips as PACC-T, shipped once in pacc_base.tar with per-source metadata.

Pool Clips Native format Source corpus
AMI 2,992 (4.45 h) 16 kHz WAV, headset microphone AMI Meeting Corpus: 139 meetings, 157 participants
VCTK_mic1 1,500 (1.36 h) 48 kHz FLAC CSTR VCTK Corpus 0.92, mic1: 109 speakers
VCTK_mic2 1,500 (1.36 h) 48 kHz FLAC CSTR VCTK Corpus 0.92, mic2: 108 speakers

VCTK mic pairing: 1,485 utterances appear in both mic pools; 15 in each pool are unpaired. Joining mic1 to mic2 on utterance ID gives 1,485 pairs, not 1,500. All sources are mono 16-bit PCM. AMI file IDs encode meeting, channel, segment index and start/end time in seconds.

Per-source metadata ships as sources/metadata.csv in pacc_base.tar: speaker, gender and where the label came from, VCTK age and utterance, AMI meeting, channel, participant and segment times.

Gender balance

Pool Female Male Label source
VCTK_mic1 750 750 VCTK speaker metadata
VCTK_mic2 750 750 VCTK speaker metadata
AMI 1,126 (38%) 1,866 (62%) AMI participant metadata

Sources were selected for gender balance. VCTK is balanced exactly. AMI is not: its clips were selected using gender inferred from the audio (pitch and apparent vocal-tract length), and those labels proved wrong for 17% of clips, mostly men labelled as women (441 clips, against 68 the other way). Checked against AMI's own participant metadata, the AMI pool is 62% male. The gender column gives the metadata label; the inferred label is kept, for transparency only, as gender_inferred_not_recommended. For TS meetings, which have no entry in AMI's participants.xml, sex is taken from the M/F prefix of the participant ID, which agrees with participants.xml for all 189 participants listed there.

8 AMI clips with no usable speech are excluded from all of PACC (excluded_sources.csv). Earlier internal builds also held RAVDESS and CREMA-D clips; they were removed so that PACC is CC BY only.

Conditions

Condition Family Description Output rate Rate control
babble_snr00 babble babble noise (6 talkers), 0 dB SNR as source -
babble_snr05 babble babble noise (6 talkers), 5 dB SNR as source -
babble_snr10 babble babble noise (6 talkers), 10 dB SNR as source -
babble_snr15 babble babble noise (6 talkers), 15 dB SNR as source -
babble_snr20 babble babble noise (6 talkers), 20 dB SNR as source -
env_noise_snr00 scene environment noise scene, 0 dB SNR as source -
env_noise_snr05 scene environment noise scene, 5 dB SNR as source -
env_noise_snr10 scene environment noise scene, 10 dB SNR as source -
env_noise_snr15 scene environment noise scene, 15 dB SNR as source -
env_noise_snr20 scene environment noise scene, 20 dB SNR as source -
music_instrumental_snr00 music instrumental music, 0 dB SNR as source -
music_instrumental_snr05 music instrumental music, 5 dB SNR as source -
music_instrumental_snr10 music instrumental music, 10 dB SNR as source -
music_instrumental_snr15 music instrumental music, 15 dB SNR as source -
music_instrumental_snr20 music instrumental music, 20 dB SNR as source -
music_vocal_snr00 music vocal music, 0 dB SNR as source -
music_vocal_snr05 music vocal music, 5 dB SNR as source -
music_vocal_snr10 music vocal music, 10 dB SNR as source -
music_vocal_snr15 music vocal music, 15 dB SNR as source -
music_vocal_snr20 music vocal music, 20 dB SNR as source -
white_noise_snr00 white noise white noise, 0 dB SNR as source -
white_noise_snr05 white noise white noise, 5 dB SNR as source -
white_noise_snr10 white noise white noise, 10 dB SNR as source -
white_noise_snr15 white noise white noise, 15 dB SNR as source -
white_noise_snr20 white noise white noise, 20 dB SNR as source -
reverb_small_rt030 reverb simulated room reverb (small), RT60 0.30 s as source -
reverb_medium_rt060 reverb simulated room reverb (medium), RT60 0.60 s as source -
reverb_large_rt100 reverb simulated room reverb (large), RT60 1.00 s as source -
echo_d150ms_m10db echo single echo, 150 ms delay, -10 dB as source -
echo_d150ms_m20db echo single echo, 150 ms delay, -20 dB as source -
echo_d300ms_m10db echo single echo, 300 ms delay, -10 dB as source -
echo_d300ms_m20db echo single echo, 300 ms delay, -20 dB as source -
bandpass_300_3400 filter bandpass filter, order 4, 300-3400 Hz as source -
lowpass_2400 filter lowpass filter, order 8, 2400 Hz as source -
pitch_psola_dn4 pitch psola pitch shift, PSOLA (duration kept), -4 semitones as source -
pitch_psola_dn2 pitch psola pitch shift, PSOLA (duration kept), -2 semitones as source -
pitch_psola_up2 pitch psola pitch shift, PSOLA (duration kept), +2 semitones as source -
pitch_psola_up4 pitch psola pitch shift, PSOLA (duration kept), +4 semitones as source -
pitch_resample_dn4 pitch resample pitch shift by resampling (duration changes), -4 semitones as source -
pitch_resample_dn2 pitch resample pitch shift by resampling (duration changes), -2 semitones as source -
pitch_resample_up2 pitch resample pitch shift by resampling (duration changes), +2 semitones as source -
pitch_resample_up4 pitch resample pitch shift by resampling (duration changes), +4 semitones as source -
autotune_hard autotune pitch correction (autotune), retune 0 ms as source -
autotune_natural autotune pitch correction (autotune), retune 50 ms as source -
tempo_080 tempo tempo change, pitch kept, x0.8 as source -
tempo_125 tempo tempo change, pitch kept, x1.25 as source -
tempo_150 tempo tempo change, pitch kept, x1.5 as source -
cafe_voip compound cafe scene (babble 15 dB, cafe noise 15 dB, music 20 dB) then Opus 16 kbps 48 kHz PACC-T resample_ctrl_48k
cafe_cellular compound cafe scene (babble 15 dB, cafe noise 15 dB, music 20 dB) then AMR-WB 23.85 kbps 16 kHz PACC-T resample_ctrl_16k
cafe_noise_snr15 scene cafe noise scene, 15 dB SNR as source -

How the material was made

  • Rooms: pyroomacoustics 0.10.1, image source plus ray tracing, air absorption, speaker to microphone 1.0 m at 1.5 m height, 4 layouts per room (one per clip, chosen by seed). Small 4.0 x 3.5 x 2.7 m (RT60 0.3 s), medium 7.0 x 5.0 x 3.0 m (0.6 s), large 12.0 x 9.0 x 4.0 m (1.0 s). The room impulse responses ship in pacc-p_room_cache.tar. Each reverb clip's params row names its room variant and the variant's rir_sha256; convolving the source with that response (recipe in the tarball's README) recreates the stored clip bit for bit.
  • Babble: 6 English talkers (3 female, 3 male) from MUSAN's LibriVox speech, each at its own position in the medium room.
  • Scenes: FSD50K clips, a continuous bed plus timed events at +6 dB relative to the bed. Cafe: tableware, pouring, footsteps, doors, light kitchen sounds, distant traffic, no voices, 0.4 events/s, in the medium room. Environment: home, office, street and in-car background, no voices, 0.3 events/s, no room.
  • Music: MUSAN music (CC BY and public-domain tracks only), played through a loudspeaker in the medium room. The cafe's music layer is drawn exactly like music_instrumental_snr20.
  • Echo: one copy delayed 150 or 300 ms at -10 or -20 dB.
  • Filters: Butterworth bandpass 300-3400 Hz (design order 4) and lowpass 2400 Hz (order 8).
  • Pitch: PSOLA keeps formants (Praat 6.1.38 via parselmouth 0.4.7, pitch 75-600 Hz); resampling moves formants with F0 (librosa 0.11.0, soxr_hq). Autotune snaps F0 to the nearest equal-tempered semitone (A4 = 440 Hz), retune 0 ms (hard) or 50 ms (natural).
  • Tempo: ffmpeg 8.1 atempo, pitch unchanged.
  • Randomness: every clip has its own seed, sha256 of 20260916|<condition>|<file_id>, so every draw (noise clip, room layout, event timing) is reproducible.
  • Software: Python 3.11.9, numpy 1.24.4, scipy 1.15.3, soundfile 0.13.1, librosa 0.11.0, soxr 1.0.0, praat-parselmouth 0.4.7, pyroomacoustics 0.10.1, ffmpeg 8.1.

Files

Tarball Contents
pacc_base.tar the 5,992 source clips (sources/<pool>/) and sources/metadata.csv, identical in PACC-T and PACC-P
pacc-p_<condition>.tar one condition: pacc-p/<condition>/<pool>/<file_id>.flac, params.csv, README.txt, and attribution.csv where third-party audio is mixed in
pacc-p_sample.tar 14 fixed sources (5 VCTK utterances on both mics, 4 AMI clips) through every condition
pacc-p_room_cache.tar the 24 simulated room variants (impulse responses, geometry, measured RT60/DRR) with recreation and hash recipes

All audio is 16-bit mono FLAC. SHA256SUMS and TARBALL_INDEX.csv list every tarball; pacc-p_manifest.csv lists every file with its sha256. Tarballs are deterministic.

Per-clip records

Each condition's params.csv has one row per clip, logged when the clip was generated: input and output sample rate and length, source speech level and peak, every target and achieved SNR, layer gains, safety gain, the near-floor flag, codec and codec length error for the cafe conditions.

Known edge cases

See EDGE_CASES.txt: the cafe conditions' combined-SNR tolerance (three clips deviate by more than 0.5 dB), rows near the 16-bit quantisation floor, inherited AMI clipping, and the excluded sources.

Motivation

The design of this corpus was motivated by Delgado et al. (ICASSP 2026), who argue that deepfake detection must account for how audio is presented through real communication channels, and by Lee et al. (ICASSP 2026), whose noise-aware multi-LoRA framework highlights real-world noise conditions. Neither group was involved in producing this dataset.

  • H. Delgado, G. Ramondetti, E. Dalmasso, G. Karvitsky, D. Colibro, H. Talib. "On Deepfake Voice Detection - It's All in the Presentation." ICASSP 2026. arXiv:2509.26471.
  • W. Lee, H. Dinh-Xuan, T.-P. Doan, S. Jung. "Dynamic Noise-Aware Multi LoRA Framework Towards Real-World Audio Deepfake Detection." ICASSP 2026.

Licence and credits

PACC-P is released under CC BY 4.0 by Christopher Kleingertner (Moonscape Software). It is derived from:

  • CSTR VCTK Corpus 0.92. J. Yamagishi, C. Veaux, K. MacDonald. University of Edinburgh, CSTR, 2019. doi:10.7488/ds/2645. CC BY 4.0.
  • AMI Meeting Corpus. University of Edinburgh et al. https://groups.inf.ed.ac.uk/ami/corpus/. CC BY 4.0. J. Carletta (2006), "Announcing the AMI Meeting Corpus", ELRA Newsletter 11(1).
  • MUSAN. D. Snyder, G. Chen, D. Povey (2015), arXiv:1510.08484. https://www.openslr.org/17/. CC BY 4.0; only CC BY and public-domain files used.
  • FSD50K. E. Fonseca et al. (2022), IEEE/ACM TASLP 30. doi:10.5281/zenodo.4060432. CC BY; only CC0 and CC BY clips used.

Third-party clips keep their own licences; per-clip credits are in attribution.csv (4,696 clips: CC0, CC BY 3.0/4.0 and public domain). Please credit these sources alongside PACC-P.

Citation

@dataset{moonscape_pacc_p_2026,
  author    = {Kleingertner, Christopher and {Moonscape Software}},
  title     = {{PACC-P: Parallel Acoustic Confound Corpus, Presentation}},
  year      = {2026},
  version   = {1.0},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.23026395}
}