PROVENANCE MANIFEST

Data Lineage & Regulatory Compliance (In accordance with GDPR, AIDA, PIPEDA, BC PIPA, and Quebec Law 25)

Transparency of origin is a core MoonScape value. The following tables document the completed upstream source corpora that comprise the Human Speech Atlas and the Synthetic Speech Atlas feature databases, separated by commercial viability and original distribution licenses.

If we have made an error in our licensing attributions please contact us immediately.

Human Speech Atlas (HSA)

All extracted telemetry in the Human Speech Atlas is derived from public domain (CC0-1.0) upstream sources, rendering the resulting structured datasets fully cleared for commercial and academic application.

NOTE:All personally identifiable data in the HSA has been removed, including File IDs linking the data to the Mozilla source data is anonymized via a non-reversible de-identification process to meet requirements for biometric data under various national and international regulatory frameworks like GDPR, AIDA, and Quebec Law 25.

For a complete list of languages assessed, click here.

Synthetic Speech Atlas (SSA) — Commercial

The following completed synthetic datasets were ingested from upstream sources with permissive or commercially-viable open data licenses. The audits on these datasets are licenseable for commercial use. Please contact us for more information.

Dataset NameRelease YearOrganization / OriginUpstream License
ASVspoof 2019 LA 2019 ASVspoof Consortium ODC-BY
ASVspoof 2021 LA & DF 2021 ASVspoof Consortium ODC-BY
ASVspoof 5 (Train, Dev, Eval) 2024 ASVspoof Consortium ODC-BY
FakeOrReal 2024 FakeOrReal Team GNU LGPL v3.0
In-The-Wild 2022 Various (Real-world Deepfakes) Apache 2.0
SONAR 2023 Meta CC-BY 4.0

Synthetic Speech Atlas (SSA) — Non-Commercial Public Releases

The following completed synthetic datasets were ingested from upstream sources featuring Non-Commercial (NC) and or Share-Alike (SA) clauses.

NOTE:These datasets are strictly provided free under relevant Share ALike licenses. No data from these datasets has been incorporated or used in any of our commercial data in any way.

Dataset NameRelease YearOrganization / OriginUpstream License
WaveFake v1.2 2021 WaveFake Team CC-BY-SA 4.0
LibriSeVoc 2022 LibriSeVoc Team CC-BY-SA 4.0

Synthetic Speech Atlas (SSA) — Internal Research Only

The following completed synthetic datasets were ingested from upstream sources featuring Non-Commercial (NC) and/or No-Derivatives (ND) clauses.

NOTE:These datasets are strictly used for internal research purposes, they are not commercially available. No data from these datasets has been incorporated or used in any of our commercial data in any way.

Dataset NameRelease YearOrganization / OriginUpstream License
ADD 2022 (T1 & T3) 2022 ADD Challenge CC-BY-NC-ND 4.0
ADD 2023 (R1 & R2) 2023 ADD Challenge CC-BY-NC-ND 4.0
AnimeVox 2025 Taresh Rajput CC-BY-NC-SA 4.0
CodecFake 2024 EmoFake Team CC-BY-NC-ND 4.0
EmoFake 2024 CC-BY-NC-ND 4.0
MLAAD 2023 MLAAD Team CC-BY-NC-ND 4.0

Baseline Human Data

Part or whole of following datasets are used in the establishment of our Commercial baseline scores used to normalize the wider corpus.

Dataset NameRelease YearOrganization / OriginUpstream License
AMI Meeting Corpus 2006 University of Edinburgh CC-BY 4.0
Crowd-sourced Emotional Multimodal Actors Dataset: CREMA-D v7 2014 H. Cao et al. ODC-BY
Human Screaming Detection Dataset 2024 Ren-Di Wu MIT
Telecommunications & Signal Processing Laboratory 48kHz Dataset 2018 McGill University Simplified BSD
Voice Cloning Tool Kit: VCTK 2019 University of Edinburgh CC-BY 4.0

Baseline Human Data — Internal Research Only

The following datasets were ingested from upstream sources featuring Non-Commercial (NC) and/or No-Derivatives (ND) clauses.

NOTE:These datasets are strictly used for internal research purposes, they are not commercially available. No data from these datasets has been incorporated or used in any of our commercial data in any way.

Dataset NameRelease YearOrganization / OriginUpstream License
Expressive Anecholic Recordings of Speech: EARS 2024 Julius Richter et al. CC-BY-NC 4.0
Ryerson Audio-Visual Database of Emotional Speech and Song: RAVDESS 2018 Ryerson University CC-BY-NC-SA 4.0
Toronto Emotional Speech Set 2010 University of Toronto CC-BY-NC-ND 4.0
Variably Intense Vocalizations of Affect and Emotion: VIVAE 2020 Holz, N et al. CC-BY-NC 4.0

Biomedical Research Data — Internal Research Only

The following datasets were ingested from upstream sources and relate to degenerative medical conditions related to vocal systems, and are used for research purposes only.

NOTE:These datasets are strictly used for internal research purposes, they are not commercially available. No data from these datasets has been incorporated or used in any of our commercial data in any way. Moonscape makes no claims about any diagnostic or clinical capacity to our research findings, any and all of which should be considered unverified until reviewed and replicated via a peer reviewed system.

Dataset NameRelease YearOrganization / OriginUpstream License
EasyCall 2021 Italian Institute of Techonology et al. CC-BY-NC 4.0
Saarbruecken Voice Database 2008 Pützer, Manfred et al. CC-BY 4.0
TORGO 2012 University of Toronto LDC User Agreement for Non-members
VOC-ALS 2024 Dubbioso, R., Spisto et al. Apache 2.0