---
license: cc-by-4.0
pretty_name: "PACC-T: Parallel Acoustic Confound Corpus, Telecoms"
language:
  - en
size_categories:
  - 100K<n<1M
task_categories:
  - audio-classification
tags:
  - speech
  - codec
  - telephony
  - bona-fide
  - robustness
  - anti-spoofing
  - parallel-corpus
---

# PACC-T: Parallel Acoustic Confound Corpus, Telecoms

PACC-T is 5,992 bona fide speech clips passed through 34 speech and audio codecs,
12 tandem codec chains and 5 resample-only controls. Every condition contains
the same 5,992 clips, so any clip can be compared with itself across all 51 conditions.
Every clip has a per-clip record of the exact commands that produced it.

It contains no synthetic or spoofed speech. It is a reference for what telecom channels do to real
speech: a bona fide baseline, a robustness benchmark and a condition-shift test set.

| | |
|---|---|
| Conditions | 51 (34 codecs, 12 tandems, 5 controls) |
| Clips per condition | 5,992 (same clips in every condition) |
| Files | 305,592 FLAC, plus 5,992 sources in `pacc_base.tar` |
| Duration | about 7.2 h per condition, 366 h in total |
| Size | 20.7 GB of audio; one condition 0.18-0.75 GB; sources 0.89 GB |
| Download | per condition, so any subset can be fetched on its own |

Companion dataset: **PACC-P** (Presentation), the same 5,992 sources under 50 noise, reverberation,
filtering and pitch/tempo conditions. doi:10.5281/zenodo.23026395

## Before you measure

### Sample rate and bandwidth are confounds, not side effects
Every narrowband codec decodes at 8 kHz, and many wideband codecs at 16 kHz.
Any feature measured on a codec condition mixes what the codec did with what
the lower sample rate and missing bandwidth did. Attribute an effect to a codec
only after comparing against the resample control at the same output rate.

### Resample controls
resample_ctrl_8k / 16k / 22k / 44k / 48k hold the same 5,992 sources,
resampled with no codec. Each codec condition is paired with the control at its
output rate (see the condition index). An effect that appears in both the codec
and its control is a rate effect.

### Output sample rate is not content bandwidth
A file's sample rate is an upper bound on its content, not a measure of it. `evs_swb_48k_*` is
stored at 48 kHz but, as super-wideband EVS, carries content only up to about 16 kHz. Tandem chains
are bounded by their narrowest hop. Estimate content bandwidth from the audio if a feature depends
on it.

### Native bandwidth differs by pool
AMI sources are native 16 kHz: content stops at 8 kHz. In resample_ctrl_22k /
44k / 48k and in codecs decoded above 16 kHz, AMI clips are upsampled and carry
no content above 8 kHz. VCTK sources are 48 kHz. Compare across rates within a
pool, or account for this when pooling.

### Formant measurement on band-limited conditions
Narrowband conditions (8 kHz output, or any condition whose content stops near
3.4-4 kHz) are a known failure case for LPC formant trackers configured for
wideband speech. With a ceiling above the available content (Praat's default is
5500 Hz), the tracker fits spurious poles in the empty band. These "ghost
formants" pull F2 and F3 low, often by hundreds of Hz, with no error flag. Set
the formant ceiling from content bandwidth or speaker, check readings against
the matching resample control, or treat F2 and above as unreliable in these
conditions.

Tool defaults matter. For example, eGeMAPSv02 derives formants after resampling
to 11 kHz with an 11th-order autocorrelation LPC, and some voice-quality
measures (such as glottal-flow quotients) change with sample rate on identical
content. Validate any feature on the resample controls before attributing an
effect to a codec.

## Sources

| Pool | Clips | Native format | Source corpus |
|---|---|---|---|
| AMI | 2,992 (4.45 h) | 16 kHz WAV, headset microphone | AMI Meeting Corpus: 139 meetings, 157 participants |
| VCTK_mic1 | 1,500 (1.36 h) | 48 kHz FLAC | CSTR VCTK Corpus 0.92, mic1: 109 speakers |
| VCTK_mic2 | 1,500 (1.36 h) | 48 kHz FLAC | CSTR VCTK Corpus 0.92, mic2: 108 speakers |

VCTK mic pairing: 1,485 utterances appear in both mic pools; 15 in each pool are unpaired. Joining
mic1 to mic2 on utterance ID gives 1,485 pairs, not 1,500. All sources are mono 16-bit PCM.
AMI file IDs encode meeting, channel, segment index and start/end time in seconds.

Per-source metadata ships as `sources/metadata.csv` in `pacc_base.tar`: speaker, gender and where
the label came from, VCTK age and utterance, AMI meeting, channel, participant and segment times.

### Gender balance

| Pool | Female | Male | Label source |
|---|---|---|---|
| VCTK_mic1 | 750 | 750 | VCTK speaker metadata |
| VCTK_mic2 | 750 | 750 | VCTK speaker metadata |
| AMI | 1,126 (38%) | 1,866 (62%) | AMI participant metadata |

Sources were selected for gender balance. VCTK is balanced exactly. AMI is not: its clips were
selected using gender inferred from the audio (pitch and apparent vocal-tract length), and those
labels proved wrong for 17% of clips, mostly men labelled as women (441 clips, against 68 the other
way). Checked against AMI's own participant metadata, the AMI pool is 62% male. The `gender`
column gives the metadata label; the inferred label is kept, for transparency only, as
`gender_inferred_not_recommended`. For TS meetings, which have no entry in AMI's
`participants.xml`, sex is taken from the M/F prefix of the participant ID, which agrees with
`participants.xml` for all 189 participants listed there.

8 AMI clips with no usable speech are excluded from all of PACC (`excluded_sources.csv`). Earlier
internal builds also held RAVDESS and CREMA-D clips; they were removed so that PACC is CC BY only.

## Conditions

Output sample rate is the codec's native decode rate. Compare each condition with the resample
control at the same rate (last column) before attributing an effect to the codec.

Framing: **block** = frame-based codec; **sample** = sample-by-sample waveform coder with no frame
structure. Tools that look for frame boundaries have nothing to find in sample-based conditions.

### Codecs

| Condition | Family | Codec | Bitrate | Framing | Output rate | Rate control |
|---|---|---|---|---|---|---|
| `aac_32k` | media | AAC-LC (ffmpeg native) | 32 kbps | block | 22.05 kHz | `resample_ctrl_22k` |
| `aac_64k` | media | AAC-LC (ffmpeg native) | 64 kbps | block | 44.1 kHz | `resample_ctrl_44k` |
| `amr_nb_122` | mobile | AMR-NB | 12.2 kbps | block | 8 kHz | `resample_ctrl_8k` |
| `amr_nb_475` | mobile | AMR-NB | 4.75 kbps | block | 8 kHz | `resample_ctrl_8k` |
| `amr_wb` | mobile | AMR-WB | 23.85 kbps | block | 16 kHz | `resample_ctrl_16k` |
| `codec2_1300` | low-rate | Codec 2 (1300 mode, see EDGE_CASES) | 1.3 kbps | block | 8 kHz | `resample_ctrl_8k` |
| `codec2_700` | low-rate | Codec 2 700C | 0.7 kbps | block | 8 kHz | `resample_ctrl_8k` |
| `evs_24400_dtxadapt` | mobile (VoLTE) | EVS, DTX adaptive CNG | 24.4 kbps | block | 16 kHz | `resample_ctrl_16k` |
| `evs_24400_dtxfixed` | mobile (VoLTE) | EVS, DTX fixed 8-frame CNG | 24.4 kbps | block | 16 kHz | `resample_ctrl_16k` |
| `evs_24400_nodtx` | mobile (VoLTE) | EVS, DTX off | 24.4 kbps | block | 16 kHz | `resample_ctrl_16k` |
| `evs_9600_dtxadapt` | mobile (VoLTE) | EVS, DTX adaptive CNG | 9.6 kbps | block | 16 kHz | `resample_ctrl_16k` |
| `evs_9600_dtxfixed` | mobile (VoLTE) | EVS, DTX fixed 8-frame CNG | 9.6 kbps | block | 16 kHz | `resample_ctrl_16k` |
| `evs_9600_nodtx` | mobile (VoLTE) | EVS, DTX off | 9.6 kbps | block | 16 kHz | `resample_ctrl_16k` |
| `evs_swb_48k_dtxadapt` | mobile (VoLTE) | EVS super-wideband, DTX adaptive CNG | 24.4 kbps | block | 48 kHz | `resample_ctrl_48k` |
| `evs_swb_48k_dtxfixed` | mobile (VoLTE) | EVS super-wideband, DTX fixed 8-frame CNG | 24.4 kbps | block | 48 kHz | `resample_ctrl_48k` |
| `evs_swb_48k_nodtx` | mobile (VoLTE) | EVS super-wideband, DTX off | 24.4 kbps | block | 48 kHz | `resample_ctrl_48k` |
| `g711_alaw` | PSTN | G.711 A-law | 64 kbps | sample | 8 kHz | `resample_ctrl_8k` |
| `g711_ulaw` | PSTN | G.711 mu-law | 64 kbps | sample | 8 kHz | `resample_ctrl_8k` |
| `g722` | PSTN wideband | G.722 (sub-band ADPCM) | 64 kbps | sample | 16 kHz | `resample_ctrl_16k` |
| `g726_16k` | PSTN | G.726 ADPCM | 16 kbps | sample | 8 kHz | `resample_ctrl_8k` |
| `g726_24k` | PSTN | G.726 ADPCM | 24 kbps | sample | 8 kHz | `resample_ctrl_8k` |
| `g726_32k` | PSTN | G.726 ADPCM | 32 kbps | sample | 8 kHz | `resample_ctrl_8k` |
| `gsm` | mobile | GSM 06.10 full rate | 13 kbps | block | 8 kHz | `resample_ctrl_8k` |
| `ilbc` | VoIP | iLBC | encoder default | block | 8 kHz | `resample_ctrl_8k` |
| `lc3` | Bluetooth LE Audio | LC3 | encoder default | block | 16 kHz | `resample_ctrl_16k` |
| `mp3_128k` | media | MP3 (LAME) | 128 kbps | block | 44.1 kHz | `resample_ctrl_44k` |
| `mp3_32k` | media | MP3 (LAME) | 32 kbps | block | 22.05 kHz | `resample_ctrl_22k` |
| `opus_16k_auto` | VoIP | Opus, mode chosen by encoder | 16 kbps | block | 48 kHz | `resample_ctrl_48k` |
| `opus_16k_celt` | VoIP | Opus, CELT forced (lowdelay) | 16 kbps | block | 48 kHz | `resample_ctrl_48k` |
| `opus_32k_auto` | VoIP | Opus, mode chosen by encoder | 32 kbps | block | 48 kHz | `resample_ctrl_48k` |
| `opus_32k_celt` | VoIP | Opus, CELT forced (lowdelay) | 32 kbps | block | 48 kHz | `resample_ctrl_48k` |
| `opus_6k_auto` | VoIP | Opus, mode chosen by encoder | 6 kbps | block | 48 kHz | `resample_ctrl_48k` |
| `opus_6k_celt` | VoIP | Opus, CELT forced (lowdelay) | 6 kbps | block | 48 kHz | `resample_ctrl_48k` |
| `speex_8k` | VoIP | Speex narrowband | encoder default | block | 8 kHz | `resample_ctrl_8k` |

### Tandem chains (two codecs in sequence)

The first codec's output is the second codec's input. Effective bandwidth is bounded by the
narrowest hop, and the output's frame structure and quantisation reflect the last codec: in
`amr_wb_to_g711_ulaw` the output carries G.711's sample-by-sample companding lattice at 8 kHz,
while the speech had already been through AMR-WB's frame-based coding.

| Condition | Chain | Framing | Output rate | Rate control |
|---|---|---|---|---|
| `amr_wb_to_g711_ulaw` | AMR-WB 23.85 kbps then G.711 mu-law 64 kbps | block then sample | 8 kHz | `resample_ctrl_8k` |
| `evs_24400_dtxfixed_to_amr_nb_475` | EVS, DTX fixed 8-frame CNG 24.4 kbps then AMR-NB 4.75 kbps | block then block | 8 kHz | `resample_ctrl_8k` |
| `evs_24400_dtxfixed_to_amr_wb` | EVS, DTX fixed 8-frame CNG 24.4 kbps then AMR-WB 23.85 kbps | block then block | 16 kHz | `resample_ctrl_16k` |
| `evs_24400_dtxfixed_to_g711_ulaw` | EVS, DTX fixed 8-frame CNG 24.4 kbps then G.711 mu-law 64 kbps | block then sample | 8 kHz | `resample_ctrl_8k` |
| `evs_24400_nodtx_to_amr_nb_475` | EVS, DTX off 24.4 kbps then AMR-NB 4.75 kbps | block then block | 8 kHz | `resample_ctrl_8k` |
| `evs_24400_nodtx_to_amr_wb` | EVS, DTX off 24.4 kbps then AMR-WB 23.85 kbps | block then block | 16 kHz | `resample_ctrl_16k` |
| `evs_24400_nodtx_to_g711_ulaw` | EVS, DTX off 24.4 kbps then G.711 mu-law 64 kbps | block then sample | 8 kHz | `resample_ctrl_8k` |
| `opus_32k_auto_to_amr_wb` | Opus, mode chosen by encoder 32 kbps then AMR-WB 23.85 kbps | block then block | 16 kHz | `resample_ctrl_16k` |
| `opus_32k_auto_to_evs_24400_dtxadapt` | Opus, mode chosen by encoder 32 kbps then EVS, DTX adaptive CNG 24.4 kbps | block then block | 16 kHz | `resample_ctrl_16k` |
| `opus_32k_auto_to_g711_ulaw` | Opus, mode chosen by encoder 32 kbps then G.711 mu-law 64 kbps | block then sample | 8 kHz | `resample_ctrl_8k` |
| `opus_32k_celt_to_amr_wb` | Opus, CELT forced (lowdelay) 32 kbps then AMR-WB 23.85 kbps | block then block | 16 kHz | `resample_ctrl_16k` |
| `opus_32k_celt_to_evs_24400_nodtx` | Opus, CELT forced (lowdelay) 32 kbps then EVS, DTX off 24.4 kbps | block then block | 16 kHz | `resample_ctrl_16k` |

### Resample controls (no codec)

Each control is the source resampled with ffmpeg 8.1's default resampler (swresample, via `-ar`)
and written as 16-bit FLAC. The codec conditions use the same resampler for their own rate
conversion, so a control differs from its codec condition only by the codec.
`resample_ctrl_48k` is sample-identical to the VCTK sources (a no-op and determinism check); for
AMI it is an upsample from 16 kHz.

| Condition | Output rate |
|---|---|
| `resample_ctrl_8k` | 8 kHz |
| `resample_ctrl_16k` | 16 kHz |
| `resample_ctrl_22k` | 22.05 kHz |
| `resample_ctrl_44k` | 44.1 kHz |
| `resample_ctrl_48k` | 48 kHz |

## Files

| Tarball | Contents |
|---|---|
| `pacc_base.tar` | the 5,992 source clips (`sources/<pool>/`), identical in PACC-T and PACC-P |
| `pacc-t_<condition>.tar` | one condition: `pacc-t/<condition>/<pool>/<file_id>.flac`, `params.csv`, `README.txt` |
| `pacc-t_sample.tar` | 14 fixed sources (5 VCTK utterances on both mics, 4 AMI clips) through every condition |

All audio is 16-bit mono FLAC. Extracting any set of tarballs builds one tree. `SHA256SUMS` and
`TARBALL_INDEX.csv` list every tarball; `pacc-t_manifest.csv` lists every file with its sha256.
Tarballs are deterministic: the same inputs rebuild byte-identical files.

## Per-clip records and verification

Each condition's `params.csv` has one row per clip: input file and sha256, output sha256, sample
rate, frame count, and the full command sequence, with paths as placeholders.

- **Logged** (6 conditions: the EVS DTX-fixed conditions and their tandems): written at encode time,
  including the EVS encoder's own DTX status line.
- **Recovered** (45 conditions, encoded before per-clip logging existed): each recorded command was
  re-executed on 60 clips per condition (20 per pool) and reproduced the stored audio bit for bit,
  2,700 of 2,700. Rows were then filled from the files.

The column `params_origin` says which applies. Before packaging, contract tests confirmed: every
condition holds exactly the 5,992 base sources once each; no excluded source ships; manifest,
params and file hashes agree.

## Tools and versions

- ffmpeg 8.1 (full_build, www.gyan.dev, Windows static), with libopencore-amrnb, libvo-amrwbenc,
  libgsm, libilbc, libspeex, libopus, libmp3lame (LAME 3.100), liblc3 and libcodec2. The static
  build does not expose the other libraries' version strings.
- EVS: 3GPP TS 26.443 floating-point reference C code, banner "Version 12.7.0 / 13.3.0",
  mirror github.com/wanglihe/3gpp-evs at commit 519236cc07ca209cb3aa2cc32de6ca686269839b.
- Decoders: only `ilbc` pins its decoder (`-c:a libilbc`); all others use ffmpeg's default decoder
  for the stream, as recorded in the params. A start-burst scan of every file (first 40 ms peak
  against the rest of the clip) flagged none.

## Known edge cases

See `EDGE_CASES.txt`. In short: EVS DTX variants coincide on clips without pauses; LC3 passes two
very quiet clips through unchanged; `codec2_1300` runs at 1300 bps despite a 700 bps argument;
350 AMI sources carry clipping from the original recordings.

## Motivation

The design of this corpus was motivated by Delgado et al. (ICASSP 2026), who argue that deepfake
detection must account for how audio is presented through real communication channels, and by Lee
et al. (ICASSP 2026), whose noise-aware multi-LoRA framework highlights real-world conditions.
Neither group was involved in producing this dataset.

- H. Delgado, G. Ramondetti, E. Dalmasso, G. Karvitsky, D. Colibro, H. Talib. "On Deepfake Voice
  Detection - It's All in the Presentation." ICASSP 2026. arXiv:2509.26471.
- W. Lee, H. Dinh-Xuan, T.-P. Doan, S. Jung. "Dynamic Noise-Aware Multi LoRA Framework Towards
  Real-World Audio Deepfake Detection." ICASSP 2026.

## Licence and credits

PACC-T is released under **CC BY 4.0** by Christopher Kleingertner (Moonscape Software). It is derived from:

- **CSTR VCTK Corpus 0.92.** J. Yamagishi, C. Veaux, K. MacDonald. University of Edinburgh, CSTR,
  2019. doi:10.7488/ds/2645. CC BY 4.0.
- **AMI Meeting Corpus.** University of Edinburgh et al. https://groups.inf.ed.ac.uk/ami/corpus/.
  CC BY 4.0. J. Carletta (2006), "Announcing the AMI Meeting Corpus", ELRA Newsletter 11(1).

Please credit these sources alongside PACC-T.

## Citation

```bibtex
@dataset{moonscape_pacc_t_2026,
  author    = {Kleingertner, Christopher and {Moonscape Software}},
  title     = {{PACC-T: Parallel Acoustic Confound Corpus, Telecoms}},
  year      = {2026},
  version   = {1.0},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.23026393}
}
```
