Gladia's speech-to-text models are built with the help of openly licensed datasets. We gratefully acknowledge the providers whose work makes our research possible.
The corpora below are distributed under the Creative Commons Attribution 4.0 International (CC-BY 4.0) license and are used to train and evaluate Gladia's speech-to-text models.
Few-shot Learning Evaluation of Universal Representations of Speech — a parallel speech corpus spanning 102 languages, designed for multilingual speech recognition and language identification.
A spoken language identification dataset covering 107 languages, totalling roughly 6,600 hours of speech automatically collected from YouTube.
Around 100 hours of multi-modal meeting recordings with rich manual annotations, widely used for speech recognition and speaker diarization research.
These datasets are used under the terms of the Creative Commons Attribution 4.0 International License. Attribution does not imply that the dataset providers endorse Gladia or its products.