Raw Speech Corpus

Raw Corpus for Speech

Bodo Raw Speech Corpus

SKU: LDCIL-187

176:53:28  hours of 113 GB | 456 Speakers | 77443 Audio segments | 48 kHz | 16 bit wav Bodo, one of the scheduled language of India,...

View Dataset →

Dogri Raw Speech Corpus

SKU: LDCIL-215

17:10:26 Hours | 11 GB speech data | 61 Speakers | 12,036 Audio segments | 48 kHz | 16 bit wav.     Dogri, the language...

View Dataset →

Odia Raw Speech Corpus

SKU: LDCIL-161

138:06:18 hours  |   89 GB | 474 Speakers | 73,418 Audio segments | 48 kHz | 16 bit wav. Odia is an Indo-Aryan language;...

View Dataset →

Urdu Raw Speech Corpus

SKU: LDCIL-165

99:18:21 Hours | 64.2 GB | 499 Speakers | 88,708 Audio Segments | 48 kHz | 16 bit wav.   Urdu is one of the Modern Indo-Aryan...

View Dataset →