MARE
Moonscape Acoustic Research Engine
In development. MARE (Moonscape Acoustic Research Engine) is still in development. It is currently undergoing calibration and validation testing, as well as benchmarking against existing toolsets, DisVoice and eGeMAPS.
What it is
MARE is Moonscape's proprietary speech measurement instrument. It takes a recording and reports measurements of it in physical units: hertz, decibels and simple ratios. The measurements shown on this page come from classical signal processing, and MARE is designed so that the same audio always produces the same numbers.
Why it runs on PACC
The Parallel Acoustic Confound Corpus holds one fixed set of 5,992 bona fide speech clips under 101 conditions: 51 telecom conditions in PACC-T and 50 presentation conditions in PACC-P. MARE has been run on every condition. Its output files are not part of the PACC downloads. Because every condition holds the same clips, the difference between a clip's measurements in two conditions is the effect of the condition, not of a different talker.
Example: one clip, three steps
The clip is labelled clip-0001 here, a 3.501-second utterance from the CSTR VCTK Corpus 0.92 (CC BY 4.0). The label is a stand-in; the source file name is not published. The table shows MARE's measurements of that one clip at three steps:
- Base (
sample_pool): the clip as it is stored in the corpus, at 48 kHz, with no condition applied. - Resample (
resample_ctrl_16k): the clip resampled to 16 kHz with no codec. It shows what the change of sample rate does on its own. - Codec (
evs_9600_dtxadapt): the EVS codec at 9.6 kbps with adaptive DTX, which outputs at 16 kHz, the same rate as the resample step. Its matched control in PACC-T is the 16 kHz resample above.
The two Change columns are measured against the base.
| Measurement | Unit | Base | Resample | Change | EVS 9.6 kbps | Change |
|---|---|---|---|---|---|---|
| Mean pitch (F0)Mean fundamental frequency over voiced frames only. | Hz | 182.4 | 182.4 | 0.0 | 183.3 | +0.8 |
| Pitch variabilityStandard deviation of the fundamental frequency over voiced frames. | Hz | 36.7 | 36.8 | 0.0 | 37.1 | +0.3 |
| Spectral centroidPower-weighted mean frequency of the spectrum, its centre of mass. | Hz | 4543.8 | 1250.3 | −3293.5 | 1237.5 | −3306.3 |
| Harmonics-to-noise ratioHow much of the voiced signal is periodic (harmonic) compared with noise. | dB † | 16.4 | 16.7 | +0.3 | 17.1 | +0.7 |
| Cepstral peak prominence, smoothedStrength of periodicity in the voice. A sharp cepstral peak means strongly periodic voice; a flat or absent peak means noisy or aperiodic. | dB | 10.9 | 10.9 | 0.0 | 10.3 | −0.6 |
| Shimmer (local)Pulse-to-pulse variation in amplitude, relative to the mean amplitude. | fraction (0.05 = 5%) † | 0.059 | 0.059 | 0.000 | 0.065 | +0.006 |
† The unit follows the standard definition of this measure. Values are displayed rounded; the download below holds them at the precision stored in the file.
What the table shows
For spectral centroid, nearly all of the change happens at the resample step, before any codec is involved. The centroid is 4544 Hz in the 48 kHz original, 1250 Hz after resampling to 16 kHz (a change of 3293 Hz), and 1238 Hz after the codec, which moves it a further 13 Hz. A 16 kHz file cannot hold anything above 8 kHz, so the resample removes the high-frequency content that the original carried.
That is why a codec has to be read against the resample control at the same output rate and not against the original. Cepstral peak prominence and shimmer behave the other way round for this clip: the resample leaves them almost where they were and the codec moves them. This is one clip; it shows how to read the columns, not how large either effect is across the corpus.
What the output looks like
MARE writes one JSON record per clip to a file called master.jsonl. In this run a full record has 104 fields. Below are the three records above, cut down to the six fields in the table. In the file each record is a single line; here it is spread over several lines for reading.
sample_pool
{
"clip": "clip-0001",
"pitch_mean": 182.442249,
"pitch_std": 36.74946,
"spectral_centroid_mean": 4543.79641,
"hnr_mean": 16.421658,
"cpps": 10.853513,
"shimmer_local": 0.058919
}
resample_ctrl_16k
{
"clip": "clip-0001",
"pitch_mean": 182.443304,
"pitch_std": 36.753268,
"spectral_centroid_mean": 1250.343481,
"hnr_mean": 16.70839,
"cpps": 10.850298,
"shimmer_local": 0.059106
}
evs_9600_dtxadapt
{
"clip": "clip-0001",
"pitch_mean": 183.259097,
"pitch_std": 37.062463,
"spectral_centroid_mean": 1237.529861,
"hnr_mean": 17.128357,
"cpps": 10.281997,
"shimmer_local": 0.065254
}
Download the example (JSON)
Provenance of the example
master_CODEC_STUDY_codec_study_sample_pool_eng.jsonl, SHA-256 abd1b5df4512be20e3fb244277cebc46eb3fb2af5d4f30a86b86e5a63abc4db2 ·
Resample master: master_CODEC_STUDY_codec_study_resample_ctrl_16k_eng.jsonl, SHA-256 373094ff36b17add9104bb8baa67828f94554334b4426f4e8033f97f627910c1 ·
EVS master: master_CODEC_STUDY_codec_study_evs_9600_dtxadapt_eng.jsonl, SHA-256 6fcd7bce29c2ea3a29b7723a9b176dfd3ae6526d1ed75e45e9af0898016b72fc
The example was extracted from those three files by a script in the site repository (scripts/extract_mare_example.py), which refuses a clip that is excluded from the release or that has a failed measurement pass. Each master file holds 6,000 rows, which includes the 8 sources that were excluded from PACC. This clip is not one of them.
Reading the example
- One clip is an illustration, not a result. It shows what MARE reports. It does not show how large or how consistent an effect is across the corpus.
- MARE is under calibration and validation, so treat every figure as provisional.