Assamese Raw Speech Corpus
SKU: LDCIL-68
54:21:12 Hours | 32.5 GB | 304 Speakers | 37,570 Audio Segments | 48 kHz | 16 bit wav. Assamese is the official language of...
View Dataset →Raw Corpus for Speech
SKU: LDCIL-68
54:21:12 Hours | 32.5 GB | 304 Speakers | 37,570 Audio Segments | 48 kHz | 16 bit wav. Assamese is the official language of...
View Dataset →SKU: LDCIL-147
128:46:59 Hours | 81.2 GB | 476 Speakers | 73,470 Audio Segments | 48 kHz | 16 bit wav. Bengali is the official language of West...
View Dataset →SKU: LDCIL-187
176:53:28 hours of 113 GB | 456 Speakers | 77443 Audio segments | 48 kHz | 16 bit wav Bodo, one of the scheduled language of India,...
View Dataset →SKU: LDCIL-247
Dataset Description: 138:09:27 Hours | 88.9 GB | 140 Speakers | 359 Audio Segments | 48 kHz | 16 bit wav LDC-IL has taken a...
View Dataset →SKU: LDCIL-215
17:10:26 Hours | 11 GB speech data | 61 Speakers | 12,036 Audio segments | 48 kHz | 16 bit wav. Dogri, the language...
View Dataset →SKU: LDCIL-231
57:17:08 Hours | 37 GB | 204 Speakers| 25,712 Audio Segments | 48 kHz | 16 bit wav . Gujarati is one of the major...
View Dataset →SKU: LDCIL-229
64:44:02 Hours | 7.1 GB | 233 Speakers| 26,223 Audio Segments | 16 kHz | 16 bit wav . Gujarati is one of the major literary...
View Dataset →SKU: LDCIL-153
121:00:06 Hours | 76.6 GB | 488 Speakers | 70686 Audio Segments | 48 kHz | 16 bit wav. Hindi is a Major, Indo-Aryan language, a...
View Dataset →SKU: LDCIL-226
25:47:11 Hours | 15.5 GB | 53 Speakers| 16,044 Audio Segments | 48 kHz | 16 bit wav . English language is a blend of Anglo-Saxon which is...
View Dataset →SKU: LDCIL-230
23:43:04 Hours | 15.3 GB | 56 Speakers| 14,455 Audio Segments | 48 kHz | 16 bit wav . English language is a blend of Anglo-Saxon...
View Dataset →SKU: LDCIL-154
179:32:52 hours of 115 GB | 656 Speakers | 99109 Audio segments | 48 k H z | 16 bit wav Kannada is one of the Ancient Indian languages...
View Dataset →SKU: LDCIL-224
28:10:07 Hours | 18 GB speech data | 150 Speakers | 16,380 Audio segments | 48 kHz | 16 bit wav. Kashmiri Language belongs...
View Dataset →SKU: LDCIL-155
156:37:51 Hours | 100 GB | 504 Speakers | 72,938 Audio Segments | 48 kHz | 16 bit wav. Konkani belongs to the Indo-European...
View Dataset →SKU: LDCIL-156
78:45:33 Hours | 49.2 GB | 306 Speakers | 45,198 Audio Segments | 48 kHz | 16 bit wav Maithili...
View Dataset →SKU: LDCIL-254
109:09:50 hours | 206 Audio Segments | 122 Speakers The LDC-IL Maithili Raw Speech dataset Vol.II comprises audio files in wav format,...
View Dataset →SKU: LDCIL-258
41:54:30 hours | 26 GB | 21,412 Audio Segments | 300 speakers The LDC-IL Maithili Sentence Aligned Speech Corpus(Tirhuta Script)...
View Dataset →SKU: LDCIL-157
164:01:02 Hours | 105 GB | 458Speakers| 43670 Audio Segments |48 kHz | 16 bit wav. Malayalam is the official language of Kerala and...
View Dataset →SKU: LDCIL-158
156:28:32 hours | 100 GB | 620 Speakers | 66,231 Audio segments | 48 khz | 16 bit wav Manipuri is the Administrative Language...
View Dataset →SKU: LDCIL-166
89:17:25 Hours | 58 GB speech data | 307 Speakers | 58544 Audio segments | 48 kHz | 16 bit wav. The Marathi language is an...
View Dataset →SKU: LDCIL-214
97:43:54 Hours | 62.2 GB speech data | 1916 Speakers | 1,916 Audio segments | 48 kHz | 16 bit wav. The LDC-IL...
View Dataset →SKU: LDCIL-SP-MULTI-RAW-001
A multi-language speech corpus package covering several Indian language recording sets.
View Dataset →SKU: LDCIL-160
87:14:44 Hours | 56.5GB | 350 Speakers | 48975 Audio Segments | 48 kHz | 16 bit wav. Nepali belongs to the Indo-Aryan language family ....
View Dataset →SKU: LDCIL-161
138:06:18 hours | 89 GB | 474 Speakers | 73,418 Audio segments | 48 kHz | 16 bit wav. Odia is an Indo-Aryan language;...
View Dataset →SKU: LDCIL-162
101:09:28 Hours | 65.5 GB | 467 Speakers | 76,230 Audio Segments | 48 kHz | 16 bit wav. Punjabi is one of the Indo-Aryan...
View Dataset →SKU: LDCIL-419
50:58:13 Hours | 15.08 GB | 60 Speakers | 11844 Audio Segments The Santali Raw Speech Corpus speech data is recorded using the SDCP (Speech...
View Dataset →SKU: LDCIL-217
139:11:41 Hours | 86 GB speech data | 452 Speakers | 60,287 Audio segments | 48 kHz | 16 bit wav. Tamil is one of the...
View Dataset →SKU: LDCIL-197
22:43:59 Hours | 15 GB | 80 Speakers | 10,510 Audio Segments | 48 kHz | 16 bit wav. Telugu is the official language of...
View Dataset →SKU: LDCIL-165
99:18:21 Hours | 64.2 GB | 499 Speakers | 88,708 Audio Segments | 48 kHz | 16 bit wav. Urdu is one of the Modern Indo-Aryan...
View Dataset →