PACC-T
Parallel Acoustic Confound Corpus, Telecoms
The same 5,992 bona fide speech clips passed through 34 codecs, 12 tandem chains and 5 resample controls.
- Version
- 1.0
- Conditions
- 51
- Clips per condition
- 5,992
- Licence
- CC BY 4.0
- DOI
- 10.5281/zenodo.23026393
What this is
Speech is changed by the channel it travels through before anyone measures it. PACC-T isolates that effect for telecom channels. It takes one fixed set of 5,992 bona fide speech clips and passes every clip through 51 conditions: 34 codecs, 12 tandem chains (two codecs in sequence) and 5 resample-only controls.
Every condition holds the same clips, so any clip can be compared with itself across all 51 conditions. Whatever changes between two versions of the same clip is the channel, not the talker.
Read the controls first
A codec does not only compress. It also changes the sample rate and the available bandwidth, and those alone can move a measurement. The five resample controls carry the same clips through the same resampler with no codec at all. Compare each codec against the control at its output rate. If an effect appears in both, it is a rate effect, not a codec effect.
Get only what you need
Each condition is its own download, so a subset can be fetched without the full 20.7 GB. A small sample archive holds 14 fixed sources through every condition, which is enough to test a pipeline before committing to the whole corpus.
Reproducibility
Every clip has a record of the exact commands that produced it, with input and output SHA-256 hashes. Where commands were recovered after the fact, each was re-run on 60 clips per condition and reproduced the stored audio bit for bit in 2,700 of 2,700 cases. Tarballs are deterministic, and every file is listed with its SHA-256 in the manifest.
What it does not contain
No synthetic or spoofed speech. PACC-T is a bona fide reference: a baseline, a robustness benchmark and a condition-shift test set.
Datacard
6de5411767ceeaaf7c460e0c2b2dbd6d41861c99807ad0405ace0cf6bd33a211 ·
Download the exact file
PACC-T is 5,992 bona fide speech clips passed through 34 speech and audio codecs, 12 tandem codec chains and 5 resample-only controls. Every condition contains the same 5,992 clips, so any clip can be compared with itself across all 51 conditions. Every clip has a per-clip record of the exact commands that produced it.
It contains no synthetic or spoofed speech. It is a reference for what telecom channels do to real speech: a bona fide baseline, a robustness benchmark and a condition-shift test set.
| Conditions | 51 (34 codecs, 12 tandems, 5 controls) |
| Clips per condition | 5,992 (same clips in every condition) |
| Files | 305,592 FLAC, plus 5,992 sources in pacc_base.tar |
| Duration | about 7.2 h per condition, 366 h in total |
| Size | 20.7 GB of audio; one condition 0.18-0.75 GB; sources 0.89 GB |
| Download | per condition, so any subset can be fetched on its own |
Companion dataset: PACC-P (Presentation), the same 5,992 sources under 50 noise, reverberation, filtering and pitch/tempo conditions. doi:10.5281/zenodo.23026395
Before you measure
Sample rate and bandwidth are confounds, not side effects
Every narrowband codec decodes at 8 kHz, and many wideband codecs at 16 kHz. Any feature measured on a codec condition mixes what the codec did with what the lower sample rate and missing bandwidth did. Attribute an effect to a codec only after comparing against the resample control at the same output rate.
Resample controls
resample_ctrl_8k / 16k / 22k / 44k / 48k hold the same 5,992 sources, resampled with no codec. Each codec condition is paired with the control at its output rate (see the condition index). An effect that appears in both the codec and its control is a rate effect.
Output sample rate is not content bandwidth
A file's sample rate is an upper bound on its content, not a measure of it. evs_swb_48k_* is
stored at 48 kHz but, as super-wideband EVS, carries content only up to about 16 kHz. Tandem chains
are bounded by their narrowest hop. Estimate content bandwidth from the audio if a feature depends
on it.
Native bandwidth differs by pool
AMI sources are native 16 kHz: content stops at 8 kHz. In resample_ctrl_22k / 44k / 48k and in codecs decoded above 16 kHz, AMI clips are upsampled and carry no content above 8 kHz. VCTK sources are 48 kHz. Compare across rates within a pool, or account for this when pooling.
Formant measurement on band-limited conditions
Narrowband conditions (8 kHz output, or any condition whose content stops near 3.4-4 kHz) are a known failure case for LPC formant trackers configured for wideband speech. With a ceiling above the available content (Praat's default is 5500 Hz), the tracker fits spurious poles in the empty band. These "ghost formants" pull F2 and F3 low, often by hundreds of Hz, with no error flag. Set the formant ceiling from content bandwidth or speaker, check readings against the matching resample control, or treat F2 and above as unreliable in these conditions.
Tool defaults matter. For example, eGeMAPSv02 derives formants after resampling to 11 kHz with an 11th-order autocorrelation LPC, and some voice-quality measures (such as glottal-flow quotients) change with sample rate on identical content. Validate any feature on the resample controls before attributing an effect to a codec.
Sources
| Pool | Clips | Native format | Source corpus |
|---|---|---|---|
| AMI | 2,992 (4.45 h) | 16 kHz WAV, headset microphone | AMI Meeting Corpus: 139 meetings, 157 participants |
| VCTK_mic1 | 1,500 (1.36 h) | 48 kHz FLAC | CSTR VCTK Corpus 0.92, mic1: 109 speakers |
| VCTK_mic2 | 1,500 (1.36 h) | 48 kHz FLAC | CSTR VCTK Corpus 0.92, mic2: 108 speakers |
VCTK mic pairing: 1,485 utterances appear in both mic pools; 15 in each pool are unpaired. Joining mic1 to mic2 on utterance ID gives 1,485 pairs, not 1,500. All sources are mono 16-bit PCM. AMI file IDs encode meeting, channel, segment index and start/end time in seconds.
Per-source metadata ships as sources/metadata.csv in pacc_base.tar: speaker, gender and where
the label came from, VCTK age and utterance, AMI meeting, channel, participant and segment times.
Gender balance
| Pool | Female | Male | Label source |
|---|---|---|---|
| VCTK_mic1 | 750 | 750 | VCTK speaker metadata |
| VCTK_mic2 | 750 | 750 | VCTK speaker metadata |
| AMI | 1,126 (38%) | 1,866 (62%) | AMI participant metadata |
Sources were selected for gender balance. VCTK is balanced exactly. AMI is not: its clips were
selected using gender inferred from the audio (pitch and apparent vocal-tract length), and those
labels proved wrong for 17% of clips, mostly men labelled as women (441 clips, against 68 the other
way). Checked against AMI's own participant metadata, the AMI pool is 62% male. The gender
column gives the metadata label; the inferred label is kept, for transparency only, as
gender_inferred_not_recommended. For TS meetings, which have no entry in AMI's
participants.xml, sex is taken from the M/F prefix of the participant ID, which agrees with
participants.xml for all 189 participants listed there.
8 AMI clips with no usable speech are excluded from all of PACC (excluded_sources.csv). Earlier
internal builds also held RAVDESS and CREMA-D clips; they were removed so that PACC is CC BY only.
Conditions
Output sample rate is the codec's native decode rate. Compare each condition with the resample control at the same rate (last column) before attributing an effect to the codec.
Framing: block = frame-based codec; sample = sample-by-sample waveform coder with no frame structure. Tools that look for frame boundaries have nothing to find in sample-based conditions.
Codecs
| Condition | Family | Codec | Bitrate | Framing | Output rate | Rate control |
|---|---|---|---|---|---|---|
aac_32k |
media | AAC-LC (ffmpeg native) | 32 kbps | block | 22.05 kHz | resample_ctrl_22k |
aac_64k |
media | AAC-LC (ffmpeg native) | 64 kbps | block | 44.1 kHz | resample_ctrl_44k |
amr_nb_122 |
mobile | AMR-NB | 12.2 kbps | block | 8 kHz | resample_ctrl_8k |
amr_nb_475 |
mobile | AMR-NB | 4.75 kbps | block | 8 kHz | resample_ctrl_8k |
amr_wb |
mobile | AMR-WB | 23.85 kbps | block | 16 kHz | resample_ctrl_16k |
codec2_1300 |
low-rate | Codec 2 (1300 mode, see EDGE_CASES) | 1.3 kbps | block | 8 kHz | resample_ctrl_8k |
codec2_700 |
low-rate | Codec 2 700C | 0.7 kbps | block | 8 kHz | resample_ctrl_8k |
evs_24400_dtxadapt |
mobile (VoLTE) | EVS, DTX adaptive CNG | 24.4 kbps | block | 16 kHz | resample_ctrl_16k |
evs_24400_dtxfixed |
mobile (VoLTE) | EVS, DTX fixed 8-frame CNG | 24.4 kbps | block | 16 kHz | resample_ctrl_16k |
evs_24400_nodtx |
mobile (VoLTE) | EVS, DTX off | 24.4 kbps | block | 16 kHz | resample_ctrl_16k |
evs_9600_dtxadapt |
mobile (VoLTE) | EVS, DTX adaptive CNG | 9.6 kbps | block | 16 kHz | resample_ctrl_16k |
evs_9600_dtxfixed |
mobile (VoLTE) | EVS, DTX fixed 8-frame CNG | 9.6 kbps | block | 16 kHz | resample_ctrl_16k |
evs_9600_nodtx |
mobile (VoLTE) | EVS, DTX off | 9.6 kbps | block | 16 kHz | resample_ctrl_16k |
evs_swb_48k_dtxadapt |
mobile (VoLTE) | EVS super-wideband, DTX adaptive CNG | 24.4 kbps | block | 48 kHz | resample_ctrl_48k |
evs_swb_48k_dtxfixed |
mobile (VoLTE) | EVS super-wideband, DTX fixed 8-frame CNG | 24.4 kbps | block | 48 kHz | resample_ctrl_48k |
evs_swb_48k_nodtx |
mobile (VoLTE) | EVS super-wideband, DTX off | 24.4 kbps | block | 48 kHz | resample_ctrl_48k |
g711_alaw |
PSTN | G.711 A-law | 64 kbps | sample | 8 kHz | resample_ctrl_8k |
g711_ulaw |
PSTN | G.711 mu-law | 64 kbps | sample | 8 kHz | resample_ctrl_8k |
g722 |
PSTN wideband | G.722 (sub-band ADPCM) | 64 kbps | sample | 16 kHz | resample_ctrl_16k |
g726_16k |
PSTN | G.726 ADPCM | 16 kbps | sample | 8 kHz | resample_ctrl_8k |
g726_24k |
PSTN | G.726 ADPCM | 24 kbps | sample | 8 kHz | resample_ctrl_8k |
g726_32k |
PSTN | G.726 ADPCM | 32 kbps | sample | 8 kHz | resample_ctrl_8k |
gsm |
mobile | GSM 06.10 full rate | 13 kbps | block | 8 kHz | resample_ctrl_8k |
ilbc |
VoIP | iLBC | encoder default | block | 8 kHz | resample_ctrl_8k |
lc3 |
Bluetooth LE Audio | LC3 | encoder default | block | 16 kHz | resample_ctrl_16k |
mp3_128k |
media | MP3 (LAME) | 128 kbps | block | 44.1 kHz | resample_ctrl_44k |
mp3_32k |
media | MP3 (LAME) | 32 kbps | block | 22.05 kHz | resample_ctrl_22k |
opus_16k_auto |
VoIP | Opus, mode chosen by encoder | 16 kbps | block | 48 kHz | resample_ctrl_48k |
opus_16k_celt |
VoIP | Opus, CELT forced (lowdelay) | 16 kbps | block | 48 kHz | resample_ctrl_48k |
opus_32k_auto |
VoIP | Opus, mode chosen by encoder | 32 kbps | block | 48 kHz | resample_ctrl_48k |
opus_32k_celt |
VoIP | Opus, CELT forced (lowdelay) | 32 kbps | block | 48 kHz | resample_ctrl_48k |
opus_6k_auto |
VoIP | Opus, mode chosen by encoder | 6 kbps | block | 48 kHz | resample_ctrl_48k |
opus_6k_celt |
VoIP | Opus, CELT forced (lowdelay) | 6 kbps | block | 48 kHz | resample_ctrl_48k |
speex_8k |
VoIP | Speex narrowband | encoder default | block | 8 kHz | resample_ctrl_8k |
Tandem chains (two codecs in sequence)
The first codec's output is the second codec's input. Effective bandwidth is bounded by the
narrowest hop, and the output's frame structure and quantisation reflect the last codec: in
amr_wb_to_g711_ulaw the output carries G.711's sample-by-sample companding lattice at 8 kHz,
while the speech had already been through AMR-WB's frame-based coding.
| Condition | Chain | Framing | Output rate | Rate control |
|---|---|---|---|---|
amr_wb_to_g711_ulaw |
AMR-WB 23.85 kbps then G.711 mu-law 64 kbps | block then sample | 8 kHz | resample_ctrl_8k |
evs_24400_dtxfixed_to_amr_nb_475 |
EVS, DTX fixed 8-frame CNG 24.4 kbps then AMR-NB 4.75 kbps | block then block | 8 kHz | resample_ctrl_8k |
evs_24400_dtxfixed_to_amr_wb |
EVS, DTX fixed 8-frame CNG 24.4 kbps then AMR-WB 23.85 kbps | block then block | 16 kHz | resample_ctrl_16k |
evs_24400_dtxfixed_to_g711_ulaw |
EVS, DTX fixed 8-frame CNG 24.4 kbps then G.711 mu-law 64 kbps | block then sample | 8 kHz | resample_ctrl_8k |
evs_24400_nodtx_to_amr_nb_475 |
EVS, DTX off 24.4 kbps then AMR-NB 4.75 kbps | block then block | 8 kHz | resample_ctrl_8k |
evs_24400_nodtx_to_amr_wb |
EVS, DTX off 24.4 kbps then AMR-WB 23.85 kbps | block then block | 16 kHz | resample_ctrl_16k |
evs_24400_nodtx_to_g711_ulaw |
EVS, DTX off 24.4 kbps then G.711 mu-law 64 kbps | block then sample | 8 kHz | resample_ctrl_8k |
opus_32k_auto_to_amr_wb |
Opus, mode chosen by encoder 32 kbps then AMR-WB 23.85 kbps | block then block | 16 kHz | resample_ctrl_16k |
opus_32k_auto_to_evs_24400_dtxadapt |
Opus, mode chosen by encoder 32 kbps then EVS, DTX adaptive CNG 24.4 kbps | block then block | 16 kHz | resample_ctrl_16k |
opus_32k_auto_to_g711_ulaw |
Opus, mode chosen by encoder 32 kbps then G.711 mu-law 64 kbps | block then sample | 8 kHz | resample_ctrl_8k |
opus_32k_celt_to_amr_wb |
Opus, CELT forced (lowdelay) 32 kbps then AMR-WB 23.85 kbps | block then block | 16 kHz | resample_ctrl_16k |
opus_32k_celt_to_evs_24400_nodtx |
Opus, CELT forced (lowdelay) 32 kbps then EVS, DTX off 24.4 kbps | block then block | 16 kHz | resample_ctrl_16k |
Resample controls (no codec)
Each control is the source resampled with ffmpeg 8.1's default resampler (swresample, via -ar)
and written as 16-bit FLAC. The codec conditions use the same resampler for their own rate
conversion, so a control differs from its codec condition only by the codec.
resample_ctrl_48k is sample-identical to the VCTK sources (a no-op and determinism check); for
AMI it is an upsample from 16 kHz.
| Condition | Output rate |
|---|---|
resample_ctrl_8k |
8 kHz |
resample_ctrl_16k |
16 kHz |
resample_ctrl_22k |
22.05 kHz |
resample_ctrl_44k |
44.1 kHz |
resample_ctrl_48k |
48 kHz |
Files
| Tarball | Contents |
|---|---|
pacc_base.tar |
the 5,992 source clips (sources/<pool>/), identical in PACC-T and PACC-P |
pacc-t_<condition>.tar |
one condition: pacc-t/<condition>/<pool>/<file_id>.flac, params.csv, README.txt |
pacc-t_sample.tar |
14 fixed sources (5 VCTK utterances on both mics, 4 AMI clips) through every condition |
All audio is 16-bit mono FLAC. Extracting any set of tarballs builds one tree. SHA256SUMS and
TARBALL_INDEX.csv list every tarball; pacc-t_manifest.csv lists every file with its sha256.
Tarballs are deterministic: the same inputs rebuild byte-identical files.
Per-clip records and verification
Each condition's params.csv has one row per clip: input file and sha256, output sha256, sample
rate, frame count, and the full command sequence, with paths as placeholders.
- Logged (6 conditions: the EVS DTX-fixed conditions and their tandems): written at encode time, including the EVS encoder's own DTX status line.
- Recovered (45 conditions, encoded before per-clip logging existed): each recorded command was re-executed on 60 clips per condition (20 per pool) and reproduced the stored audio bit for bit, 2,700 of 2,700. Rows were then filled from the files.
The column params_origin says which applies. Before packaging, contract tests confirmed: every
condition holds exactly the 5,992 base sources once each; no excluded source ships; manifest,
params and file hashes agree.
Tools and versions
- ffmpeg 8.1 (full_build, www.gyan.dev, Windows static), with libopencore-amrnb, libvo-amrwbenc, libgsm, libilbc, libspeex, libopus, libmp3lame (LAME 3.100), liblc3 and libcodec2. The static build does not expose the other libraries' version strings.
- EVS: 3GPP TS 26.443 floating-point reference C code, banner "Version 12.7.0 / 13.3.0", mirror github.com/wanglihe/3gpp-evs at commit 519236cc07ca209cb3aa2cc32de6ca686269839b.
- Decoders: only
ilbcpins its decoder (-c:a libilbc); all others use ffmpeg's default decoder for the stream, as recorded in the params. A start-burst scan of every file (first 40 ms peak against the rest of the clip) flagged none.
Known edge cases
See EDGE_CASES.txt. In short: EVS DTX variants coincide on clips without pauses; LC3 passes two
very quiet clips through unchanged; codec2_1300 runs at 1300 bps despite a 700 bps argument;
350 AMI sources carry clipping from the original recordings.
Motivation
The design of this corpus was motivated by Delgado et al. (ICASSP 2026), who argue that deepfake detection must account for how audio is presented through real communication channels, and by Lee et al. (ICASSP 2026), whose noise-aware multi-LoRA framework highlights real-world conditions. Neither group was involved in producing this dataset.
- H. Delgado, G. Ramondetti, E. Dalmasso, G. Karvitsky, D. Colibro, H. Talib. "On Deepfake Voice Detection - It's All in the Presentation." ICASSP 2026. arXiv:2509.26471.
- W. Lee, H. Dinh-Xuan, T.-P. Doan, S. Jung. "Dynamic Noise-Aware Multi LoRA Framework Towards Real-World Audio Deepfake Detection." ICASSP 2026.
Licence and credits
PACC-T is released under CC BY 4.0 by Christopher Kleingertner (Moonscape Software). It is derived from:
- CSTR VCTK Corpus 0.92. J. Yamagishi, C. Veaux, K. MacDonald. University of Edinburgh, CSTR, 2019. doi:10.7488/ds/2645. CC BY 4.0.
- AMI Meeting Corpus. University of Edinburgh et al. https://groups.inf.ed.ac.uk/ami/corpus/. CC BY 4.0. J. Carletta (2006), "Announcing the AMI Meeting Corpus", ELRA Newsletter 11(1).
Please credit these sources alongside PACC-T.
Citation
@dataset{moonscape_pacc_t_2026,
author = {Kleingertner, Christopher and {Moonscape Software}},
title = {{PACC-T: Parallel Acoustic Confound Corpus, Telecoms}},
year = {2026},
version = {1.0},
publisher = {Zenodo},
doi = {10.5281/zenodo.23026393}
}