Transparency of origin is a core MoonScape value. The following tables document the completed upstream source corpora that comprise the Human Speech Atlas and the Synthetic Speech Atlas feature databases, separated by commercial viability and original distribution licenses.
If we have made an error in our licensing attributions please contact us immediately.
All extracted telemetry in the Human Speech Atlas is derived from public domain (CC0-1.0) upstream sources, rendering the resulting structured datasets fully cleared for commercial and academic application.
NOTE:All personally identifiable data in the HSA has been removed, including File IDs linking the data to the Mozilla source data is anonymized via a non-reversible de-identification process to meet requirements for biometric data under various national and international regulatory frameworks like GDPR, AIDA, and Quebec Law 25.
For a complete list of languages assessed, click here.
The following completed synthetic datasets were ingested from upstream sources with permissive or commercially-viable open data licenses. The audits on these datasets are licenseable for commercial use. Please contact us for more information.
| Dataset Name | Release Year | Organization / Origin | Upstream License |
|---|---|---|---|
| ASVspoof 2019 LA | 2019 | ASVspoof Consortium | ODC-BY |
| ASVspoof 2021 LA & DF | 2021 | ASVspoof Consortium | ODC-BY |
| ASVspoof 5 (Train, Dev, Eval) | 2024 | ASVspoof Consortium | ODC-BY |
| FakeOrReal | 2024 | FakeOrReal Team | GNU LGPL v3.0 |
| In-The-Wild | 2022 | Various (Real-world Deepfakes) | Apache 2.0 |
| SONAR | 2023 | Meta | CC-BY 4.0 |
The following completed synthetic datasets were ingested from upstream sources featuring Non-Commercial (NC) and or Share-Alike (SA) clauses.
NOTE:These datasets are strictly provided free under relevant Share ALike licenses. No data from these datasets has been incorporated or used in any of our commercial data in any way.
| Dataset Name | Release Year | Organization / Origin | Upstream License |
|---|---|---|---|
| WaveFake v1.2 | 2021 | WaveFake Team | CC-BY-SA 4.0 |
| LibriSeVoc | 2022 | LibriSeVoc Team | CC-BY-SA 4.0 |
The following completed synthetic datasets were ingested from upstream sources featuring Non-Commercial (NC) and/or No-Derivatives (ND) clauses.
NOTE:These datasets are strictly used for internal research purposes, they are not commercially available. No data from these datasets has been incorporated or used in any of our commercial data in any way.
| Dataset Name | Release Year | Organization / Origin | Upstream License |
|---|---|---|---|
| ADD 2022 (T1 & T3) | 2022 | ADD Challenge | CC-BY-NC-ND 4.0 |
| ADD 2023 (R1 & R2) | 2023 | ADD Challenge | CC-BY-NC-ND 4.0 |
| AnimeVox | 2025 | Taresh Rajput | CC-BY-NC-SA 4.0 |
| CodecFake | 2024 | EmoFake Team | CC-BY-NC-ND 4.0 |
| EmoFake | 2024 | CC-BY-NC-ND 4.0 | |
| MLAAD | 2023 | MLAAD Team | CC-BY-NC-ND 4.0 |
Part or whole of following datasets are used in the establishment of our Commercial baseline scores used to normalize the wider corpus.
| Dataset Name | Release Year | Organization / Origin | Upstream License |
|---|---|---|---|
| AMI Meeting Corpus | 2006 | University of Edinburgh | CC-BY 4.0 |
| Crowd-sourced Emotional Multimodal Actors Dataset: CREMA-D v7 | 2014 | H. Cao et al. | ODC-BY |
| Human Screaming Detection Dataset | 2024 | Ren-Di Wu | MIT |
| Telecommunications & Signal Processing Laboratory 48kHz Dataset | 2018 | McGill University | Simplified BSD |
| Voice Cloning Tool Kit: VCTK | 2019 | University of Edinburgh | CC-BY 4.0 |
The following datasets were ingested from upstream sources featuring Non-Commercial (NC) and/or No-Derivatives (ND) clauses.
NOTE:These datasets are strictly used for internal research purposes, they are not commercially available. No data from these datasets has been incorporated or used in any of our commercial data in any way.
| Dataset Name | Release Year | Organization / Origin | Upstream License |
|---|---|---|---|
| Expressive Anecholic Recordings of Speech: EARS | 2024 | Julius Richter et al. | CC-BY-NC 4.0 |
| Ryerson Audio-Visual Database of Emotional Speech and Song: RAVDESS | 2018 | Ryerson University | CC-BY-NC-SA 4.0 |
| Toronto Emotional Speech Set | 2010 | University of Toronto | CC-BY-NC-ND 4.0 |
| Variably Intense Vocalizations of Affect and Emotion: VIVAE | 2020 | Holz, N et al. | CC-BY-NC 4.0 |
The following datasets were ingested from upstream sources and relate to degenerative medical conditions related to vocal systems, and are used for research purposes only.
NOTE:These datasets are strictly used for internal research purposes, they are not commercially available. No data from these datasets has been incorporated or used in any of our commercial data in any way. Moonscape makes no claims about any diagnostic or clinical capacity to our research findings, any and all of which should be considered unverified until reviewed and replicated via a peer reviewed system.
| Dataset Name | Release Year | Organization / Origin | Upstream License |
|---|---|---|---|
| EasyCall | 2021 | Italian Institute of Techonology et al. | CC-BY-NC 4.0 |
| Saarbruecken Voice Database | 2008 | Pützer, Manfred et al. | CC-BY 4.0 |
| TORGO | 2012 | University of Toronto | LDC User Agreement for Non-members |
| VOC-ALS | 2024 | Dubbioso, R., Spisto et al. | Apache 2.0 |