The same 5,992 bona fide speech clips passed through 34 codecs, 12 tandem chains and 5 resample controls.

Version
1.0
Conditions
51
Clips per condition
5,992
Licence
CC BY 4.0
DOI
10.5281/zenodo.23026393
Hugging Face: pending Zenodo: pending Companion: PACC-P

What this is

Speech is changed by the channel it travels through before anyone measures it. PACC-T isolates that effect for telecom channels. It takes one fixed set of 5,992 bona fide speech clips and passes every clip through 51 conditions: 34 codecs, 12 tandem chains (two codecs in sequence) and 5 resample-only controls.

Every condition holds the same clips, so any clip can be compared with itself across all 51 conditions. Whatever changes between two versions of the same clip is the channel, not the talker.

Read the controls first

A codec does not only compress. It also changes the sample rate and the available bandwidth, and those alone can move a measurement. The five resample controls carry the same clips through the same resampler with no codec at all. Compare each codec against the control at its output rate. If an effect appears in both, it is a rate effect, not a codec effect.

Get only what you need

Each condition is its own download, so a subset can be fetched without the full 20.7 GB. A small sample archive holds 14 fixed sources through every condition, which is enough to test a pipeline before committing to the whole corpus.

Reproducibility

Every clip has a record of the exact commands that produced it, with input and output SHA-256 hashes. Where commands were recovered after the fact, each was re-run on 60 clips per condition and reproduced the stored audio bit for bit in 2,700 of 2,700 cases. Tarballs are deterministic, and every file is listed with its SHA-256 in the manifest.

What it does not contain

No synthetic or spoofed speech. PACC-T is a bona fide reference: a baseline, a robustness benchmark and a condition-shift test set.

Datacard

Snapshot: PACC-T v1.0 datacard · SHA-256: 6de5411767ceeaaf7c460e0c2b2dbd6d41861c99807ad0405ace0cf6bd33a211 · Download the exact file

PACC-T is 5,992 bona fide speech clips passed through 34 speech and audio codecs, 12 tandem codec chains and 5 resample-only controls. Every condition contains the same 5,992 clips, so any clip can be compared with itself across all 51 conditions. Every clip has a per-clip record of the exact commands that produced it.

It contains no synthetic or spoofed speech. It is a reference for what telecom channels do to real speech: a bona fide baseline, a robustness benchmark and a condition-shift test set.

Conditions 51 (34 codecs, 12 tandems, 5 controls)
Clips per condition 5,992 (same clips in every condition)
Files 305,592 FLAC, plus 5,992 sources in pacc_base.tar
Duration about 7.2 h per condition, 366 h in total
Size 20.7 GB of audio; one condition 0.18-0.75 GB; sources 0.89 GB
Download per condition, so any subset can be fetched on its own

Companion dataset: PACC-P (Presentation), the same 5,992 sources under 50 noise, reverberation, filtering and pitch/tempo conditions. doi:10.5281/zenodo.23026395

Before you measure

Sample rate and bandwidth are confounds, not side effects

Every narrowband codec decodes at 8 kHz, and many wideband codecs at 16 kHz. Any feature measured on a codec condition mixes what the codec did with what the lower sample rate and missing bandwidth did. Attribute an effect to a codec only after comparing against the resample control at the same output rate.

Resample controls

resample_ctrl_8k / 16k / 22k / 44k / 48k hold the same 5,992 sources, resampled with no codec. Each codec condition is paired with the control at its output rate (see the condition index). An effect that appears in both the codec and its control is a rate effect.

Output sample rate is not content bandwidth

A file's sample rate is an upper bound on its content, not a measure of it. evs_swb_48k_* is stored at 48 kHz but, as super-wideband EVS, carries content only up to about 16 kHz. Tandem chains are bounded by their narrowest hop. Estimate content bandwidth from the audio if a feature depends on it.

Native bandwidth differs by pool

AMI sources are native 16 kHz: content stops at 8 kHz. In resample_ctrl_22k / 44k / 48k and in codecs decoded above 16 kHz, AMI clips are upsampled and carry no content above 8 kHz. VCTK sources are 48 kHz. Compare across rates within a pool, or account for this when pooling.

Formant measurement on band-limited conditions

Narrowband conditions (8 kHz output, or any condition whose content stops near 3.4-4 kHz) are a known failure case for LPC formant trackers configured for wideband speech. With a ceiling above the available content (Praat's default is 5500 Hz), the tracker fits spurious poles in the empty band. These "ghost formants" pull F2 and F3 low, often by hundreds of Hz, with no error flag. Set the formant ceiling from content bandwidth or speaker, check readings against the matching resample control, or treat F2 and above as unreliable in these conditions.

Tool defaults matter. For example, eGeMAPSv02 derives formants after resampling to 11 kHz with an 11th-order autocorrelation LPC, and some voice-quality measures (such as glottal-flow quotients) change with sample rate on identical content. Validate any feature on the resample controls before attributing an effect to a codec.

Sources

Pool Clips Native format Source corpus
AMI 2,992 (4.45 h) 16 kHz WAV, headset microphone AMI Meeting Corpus: 139 meetings, 157 participants
VCTK_mic1 1,500 (1.36 h) 48 kHz FLAC CSTR VCTK Corpus 0.92, mic1: 109 speakers
VCTK_mic2 1,500 (1.36 h) 48 kHz FLAC CSTR VCTK Corpus 0.92, mic2: 108 speakers

VCTK mic pairing: 1,485 utterances appear in both mic pools; 15 in each pool are unpaired. Joining mic1 to mic2 on utterance ID gives 1,485 pairs, not 1,500. All sources are mono 16-bit PCM. AMI file IDs encode meeting, channel, segment index and start/end time in seconds.

Per-source metadata ships as sources/metadata.csv in pacc_base.tar: speaker, gender and where the label came from, VCTK age and utterance, AMI meeting, channel, participant and segment times.

Gender balance

Pool Female Male Label source
VCTK_mic1 750 750 VCTK speaker metadata
VCTK_mic2 750 750 VCTK speaker metadata
AMI 1,126 (38%) 1,866 (62%) AMI participant metadata

Sources were selected for gender balance. VCTK is balanced exactly. AMI is not: its clips were selected using gender inferred from the audio (pitch and apparent vocal-tract length), and those labels proved wrong for 17% of clips, mostly men labelled as women (441 clips, against 68 the other way). Checked against AMI's own participant metadata, the AMI pool is 62% male. The gender column gives the metadata label; the inferred label is kept, for transparency only, as gender_inferred_not_recommended. For TS meetings, which have no entry in AMI's participants.xml, sex is taken from the M/F prefix of the participant ID, which agrees with participants.xml for all 189 participants listed there.

8 AMI clips with no usable speech are excluded from all of PACC (excluded_sources.csv). Earlier internal builds also held RAVDESS and CREMA-D clips; they were removed so that PACC is CC BY only.

Conditions

Output sample rate is the codec's native decode rate. Compare each condition with the resample control at the same rate (last column) before attributing an effect to the codec.

Framing: block = frame-based codec; sample = sample-by-sample waveform coder with no frame structure. Tools that look for frame boundaries have nothing to find in sample-based conditions.

Codecs

Condition Family Codec Bitrate Framing Output rate Rate control
aac_32k media AAC-LC (ffmpeg native) 32 kbps block 22.05 kHz resample_ctrl_22k
aac_64k media AAC-LC (ffmpeg native) 64 kbps block 44.1 kHz resample_ctrl_44k
amr_nb_122 mobile AMR-NB 12.2 kbps block 8 kHz resample_ctrl_8k
amr_nb_475 mobile AMR-NB 4.75 kbps block 8 kHz resample_ctrl_8k
amr_wb mobile AMR-WB 23.85 kbps block 16 kHz resample_ctrl_16k
codec2_1300 low-rate Codec 2 (1300 mode, see EDGE_CASES) 1.3 kbps block 8 kHz resample_ctrl_8k
codec2_700 low-rate Codec 2 700C 0.7 kbps block 8 kHz resample_ctrl_8k
evs_24400_dtxadapt mobile (VoLTE) EVS, DTX adaptive CNG 24.4 kbps block 16 kHz resample_ctrl_16k
evs_24400_dtxfixed mobile (VoLTE) EVS, DTX fixed 8-frame CNG 24.4 kbps block 16 kHz resample_ctrl_16k
evs_24400_nodtx mobile (VoLTE) EVS, DTX off 24.4 kbps block 16 kHz resample_ctrl_16k
evs_9600_dtxadapt mobile (VoLTE) EVS, DTX adaptive CNG 9.6 kbps block 16 kHz resample_ctrl_16k
evs_9600_dtxfixed mobile (VoLTE) EVS, DTX fixed 8-frame CNG 9.6 kbps block 16 kHz resample_ctrl_16k
evs_9600_nodtx mobile (VoLTE) EVS, DTX off 9.6 kbps block 16 kHz resample_ctrl_16k
evs_swb_48k_dtxadapt mobile (VoLTE) EVS super-wideband, DTX adaptive CNG 24.4 kbps block 48 kHz resample_ctrl_48k
evs_swb_48k_dtxfixed mobile (VoLTE) EVS super-wideband, DTX fixed 8-frame CNG 24.4 kbps block 48 kHz resample_ctrl_48k
evs_swb_48k_nodtx mobile (VoLTE) EVS super-wideband, DTX off 24.4 kbps block 48 kHz resample_ctrl_48k
g711_alaw PSTN G.711 A-law 64 kbps sample 8 kHz resample_ctrl_8k
g711_ulaw PSTN G.711 mu-law 64 kbps sample 8 kHz resample_ctrl_8k
g722 PSTN wideband G.722 (sub-band ADPCM) 64 kbps sample 16 kHz resample_ctrl_16k
g726_16k PSTN G.726 ADPCM 16 kbps sample 8 kHz resample_ctrl_8k
g726_24k PSTN G.726 ADPCM 24 kbps sample 8 kHz resample_ctrl_8k
g726_32k PSTN G.726 ADPCM 32 kbps sample 8 kHz resample_ctrl_8k
gsm mobile GSM 06.10 full rate 13 kbps block 8 kHz resample_ctrl_8k
ilbc VoIP iLBC encoder default block 8 kHz resample_ctrl_8k
lc3 Bluetooth LE Audio LC3 encoder default block 16 kHz resample_ctrl_16k
mp3_128k media MP3 (LAME) 128 kbps block 44.1 kHz resample_ctrl_44k
mp3_32k media MP3 (LAME) 32 kbps block 22.05 kHz resample_ctrl_22k
opus_16k_auto VoIP Opus, mode chosen by encoder 16 kbps block 48 kHz resample_ctrl_48k
opus_16k_celt VoIP Opus, CELT forced (lowdelay) 16 kbps block 48 kHz resample_ctrl_48k
opus_32k_auto VoIP Opus, mode chosen by encoder 32 kbps block 48 kHz resample_ctrl_48k
opus_32k_celt VoIP Opus, CELT forced (lowdelay) 32 kbps block 48 kHz resample_ctrl_48k
opus_6k_auto VoIP Opus, mode chosen by encoder 6 kbps block 48 kHz resample_ctrl_48k
opus_6k_celt VoIP Opus, CELT forced (lowdelay) 6 kbps block 48 kHz resample_ctrl_48k
speex_8k VoIP Speex narrowband encoder default block 8 kHz resample_ctrl_8k

Tandem chains (two codecs in sequence)

The first codec's output is the second codec's input. Effective bandwidth is bounded by the narrowest hop, and the output's frame structure and quantisation reflect the last codec: in amr_wb_to_g711_ulaw the output carries G.711's sample-by-sample companding lattice at 8 kHz, while the speech had already been through AMR-WB's frame-based coding.

Condition Chain Framing Output rate Rate control
amr_wb_to_g711_ulaw AMR-WB 23.85 kbps then G.711 mu-law 64 kbps block then sample 8 kHz resample_ctrl_8k
evs_24400_dtxfixed_to_amr_nb_475 EVS, DTX fixed 8-frame CNG 24.4 kbps then AMR-NB 4.75 kbps block then block 8 kHz resample_ctrl_8k
evs_24400_dtxfixed_to_amr_wb EVS, DTX fixed 8-frame CNG 24.4 kbps then AMR-WB 23.85 kbps block then block 16 kHz resample_ctrl_16k
evs_24400_dtxfixed_to_g711_ulaw EVS, DTX fixed 8-frame CNG 24.4 kbps then G.711 mu-law 64 kbps block then sample 8 kHz resample_ctrl_8k
evs_24400_nodtx_to_amr_nb_475 EVS, DTX off 24.4 kbps then AMR-NB 4.75 kbps block then block 8 kHz resample_ctrl_8k
evs_24400_nodtx_to_amr_wb EVS, DTX off 24.4 kbps then AMR-WB 23.85 kbps block then block 16 kHz resample_ctrl_16k
evs_24400_nodtx_to_g711_ulaw EVS, DTX off 24.4 kbps then G.711 mu-law 64 kbps block then sample 8 kHz resample_ctrl_8k
opus_32k_auto_to_amr_wb Opus, mode chosen by encoder 32 kbps then AMR-WB 23.85 kbps block then block 16 kHz resample_ctrl_16k
opus_32k_auto_to_evs_24400_dtxadapt Opus, mode chosen by encoder 32 kbps then EVS, DTX adaptive CNG 24.4 kbps block then block 16 kHz resample_ctrl_16k
opus_32k_auto_to_g711_ulaw Opus, mode chosen by encoder 32 kbps then G.711 mu-law 64 kbps block then sample 8 kHz resample_ctrl_8k
opus_32k_celt_to_amr_wb Opus, CELT forced (lowdelay) 32 kbps then AMR-WB 23.85 kbps block then block 16 kHz resample_ctrl_16k
opus_32k_celt_to_evs_24400_nodtx Opus, CELT forced (lowdelay) 32 kbps then EVS, DTX off 24.4 kbps block then block 16 kHz resample_ctrl_16k

Resample controls (no codec)

Each control is the source resampled with ffmpeg 8.1's default resampler (swresample, via -ar) and written as 16-bit FLAC. The codec conditions use the same resampler for their own rate conversion, so a control differs from its codec condition only by the codec. resample_ctrl_48k is sample-identical to the VCTK sources (a no-op and determinism check); for AMI it is an upsample from 16 kHz.

Condition Output rate
resample_ctrl_8k 8 kHz
resample_ctrl_16k 16 kHz
resample_ctrl_22k 22.05 kHz
resample_ctrl_44k 44.1 kHz
resample_ctrl_48k 48 kHz

Files

Tarball Contents
pacc_base.tar the 5,992 source clips (sources/<pool>/), identical in PACC-T and PACC-P
pacc-t_<condition>.tar one condition: pacc-t/<condition>/<pool>/<file_id>.flac, params.csv, README.txt
pacc-t_sample.tar 14 fixed sources (5 VCTK utterances on both mics, 4 AMI clips) through every condition

All audio is 16-bit mono FLAC. Extracting any set of tarballs builds one tree. SHA256SUMS and TARBALL_INDEX.csv list every tarball; pacc-t_manifest.csv lists every file with its sha256. Tarballs are deterministic: the same inputs rebuild byte-identical files.

Per-clip records and verification

Each condition's params.csv has one row per clip: input file and sha256, output sha256, sample rate, frame count, and the full command sequence, with paths as placeholders.

  • Logged (6 conditions: the EVS DTX-fixed conditions and their tandems): written at encode time, including the EVS encoder's own DTX status line.
  • Recovered (45 conditions, encoded before per-clip logging existed): each recorded command was re-executed on 60 clips per condition (20 per pool) and reproduced the stored audio bit for bit, 2,700 of 2,700. Rows were then filled from the files.

The column params_origin says which applies. Before packaging, contract tests confirmed: every condition holds exactly the 5,992 base sources once each; no excluded source ships; manifest, params and file hashes agree.

Tools and versions

  • ffmpeg 8.1 (full_build, www.gyan.dev, Windows static), with libopencore-amrnb, libvo-amrwbenc, libgsm, libilbc, libspeex, libopus, libmp3lame (LAME 3.100), liblc3 and libcodec2. The static build does not expose the other libraries' version strings.
  • EVS: 3GPP TS 26.443 floating-point reference C code, banner "Version 12.7.0 / 13.3.0", mirror github.com/wanglihe/3gpp-evs at commit 519236cc07ca209cb3aa2cc32de6ca686269839b.
  • Decoders: only ilbc pins its decoder (-c:a libilbc); all others use ffmpeg's default decoder for the stream, as recorded in the params. A start-burst scan of every file (first 40 ms peak against the rest of the clip) flagged none.

Known edge cases

See EDGE_CASES.txt. In short: EVS DTX variants coincide on clips without pauses; LC3 passes two very quiet clips through unchanged; codec2_1300 runs at 1300 bps despite a 700 bps argument; 350 AMI sources carry clipping from the original recordings.

Motivation

The design of this corpus was motivated by Delgado et al. (ICASSP 2026), who argue that deepfake detection must account for how audio is presented through real communication channels, and by Lee et al. (ICASSP 2026), whose noise-aware multi-LoRA framework highlights real-world conditions. Neither group was involved in producing this dataset.

  • H. Delgado, G. Ramondetti, E. Dalmasso, G. Karvitsky, D. Colibro, H. Talib. "On Deepfake Voice Detection - It's All in the Presentation." ICASSP 2026. arXiv:2509.26471.
  • W. Lee, H. Dinh-Xuan, T.-P. Doan, S. Jung. "Dynamic Noise-Aware Multi LoRA Framework Towards Real-World Audio Deepfake Detection." ICASSP 2026.

Licence and credits

PACC-T is released under CC BY 4.0 by Christopher Kleingertner (Moonscape Software). It is derived from:

  • CSTR VCTK Corpus 0.92. J. Yamagishi, C. Veaux, K. MacDonald. University of Edinburgh, CSTR, 2019. doi:10.7488/ds/2645. CC BY 4.0.
  • AMI Meeting Corpus. University of Edinburgh et al. https://groups.inf.ed.ac.uk/ami/corpus/. CC BY 4.0. J. Carletta (2006), "Announcing the AMI Meeting Corpus", ELRA Newsletter 11(1).

Please credit these sources alongside PACC-T.

Citation

@dataset{moonscape_pacc_t_2026,
  author    = {Kleingertner, Christopher and {Moonscape Software}},
  title     = {{PACC-T: Parallel Acoustic Confound Corpus, Telecoms}},
  year      = {2026},
  version   = {1.0},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.23026393}
}