Assamese Raw Speech Corpus
SKU: LDCIL-68
54:21:12 Hours | 32.5 GB | 304 Speakers | 37,570 Audio Segments | 48 kHz | 16 bit wav. Assamese is the official language of...
View Dataset →Speech type resource
SKU: LDCIL-68
54:21:12 Hours | 32.5 GB | 304 Speakers | 37,570 Audio Segments | 48 kHz | 16 bit wav. Assamese is the official language of...
View Dataset →SKU: LDCIL-233
800x600 Normal 0 false false false EN-US X-NONE HI MicrosoftInternetExplorer4 /* Style Definitions */ table.MsoNormalTable...
View Dataset →SKU: LDCIL-256
Assamese Text to Speech Corpus 44:49:34 hours | 28.85 GB | 32,594 Audio Segments | 2 Speakers The LDC-IL Assamese Text to...
View Dataset →SKU: LDCIL-147
128:46:59 Hours | 81.2 GB | 476 Speakers | 73,470 Audio Segments | 48 kHz | 16 bit wav. Bengali is the official language of West...
View Dataset →SKU: LDCIL-234
Dataset Description: 69:10:03 hours | 43.3 GB | 40,240 Audio Segments | 450 speakers The annotated speech corpus gives wide range of...
View Dataset →SKU: LDCIL-187
176:53:28 hours of 113 GB | 456 Speakers | 77443 Audio segments | 48 kHz | 16 bit wav Bodo, one of the scheduled language of India,...
View Dataset →SKU: LDCIL-247
Dataset Description: 138:09:27 Hours | 88.9 GB | 140 Speakers | 359 Audio Segments | 48 kHz | 16 bit wav LDC-IL has taken a...
View Dataset →SKU: LDCIL-215
17:10:26 Hours | 11 GB speech data | 61 Speakers | 12,036 Audio segments | 48 kHz | 16 bit wav. Dogri, the language...
View Dataset →SKU: LDCIL-249
08:32:54 hours | 5.6 GB | 5,039 Audio Segments | 61 Speakers The LDC-IL Dogri Sentence Aligned Speech dataset comprises audio files in wav...
View Dataset →SKU: LDCIL-231
57:17:08 Hours | 37 GB | 204 Speakers| 25,712 Audio Segments | 48 kHz | 16 bit wav . Gujarati is one of the major...
View Dataset →SKU: LDCIL-229
64:44:02 Hours | 7.1 GB | 233 Speakers| 26,223 Audio Segments | 16 kHz | 16 bit wav . Gujarati is one of the major literary...
View Dataset →SKU: LDCIL-153
121:00:06 Hours | 76.6 GB | 488 Speakers | 70686 Audio Segments | 48 kHz | 16 bit wav. Hindi is a Major, Indo-Aryan language, a...
View Dataset →SKU: LDCIL-235
Dataset Description: 72:34:52 hours | 45.9 GB | 42,275 Audio Segments | 473 speakers The annotated speech corpus gives wide range of...
View Dataset →SKU: LDCIL-226
25:47:11 Hours | 15.5 GB | 53 Speakers| 16,044 Audio Segments | 48 kHz | 16 bit wav . English language is a blend of Anglo-Saxon which is...
View Dataset →SKU: LDCIL-230
23:43:04 Hours | 15.3 GB | 56 Speakers| 14,455 Audio Segments | 48 kHz | 16 bit wav . English language is a blend of Anglo-Saxon...
View Dataset →SKU: LDCIL-244
Dataset Description: 09:21:08 hours | 5.53 GB | 5,676 Audio Segments | 52 speakers The annotated speech corpus gives wide range of...
View Dataset →SKU: LDCIL-245
Dataset Description: 11:17:40 hours | 7.27 GB | 6,166 Audio Segments | 53 speakers The annotated speech corpus gives wide range of...
View Dataset →SKU: LDCIL-154
179:32:52 hours of 115 GB | 656 Speakers | 99109 Audio segments | 48 k H z | 16 bit wav Kannada is one of the Ancient Indian languages...
View Dataset →SKU: LDCIL-236
Dataset Description: 107:48:50 hours | 69.4 GB | 65,533 Audio Segments | 600 speakers The annotated speech corpus gives wide range of...
View Dataset →SKU: LDCIL-224
28:10:07 Hours | 18 GB speech data | 150 Speakers | 16,380 Audio segments | 48 kHz | 16 bit wav. Kashmiri Language belongs...
View Dataset →SKU: LDCIL-155
156:37:51 Hours | 100 GB | 504 Speakers | 72,938 Audio Segments | 48 kHz | 16 bit wav. Konkani belongs to the Indo-European...
View Dataset →SKU: LDCIL-237
Dataset Description: 83:19:42 hours | 53.5 GB | 34,091 Audio Segments | 487 speakers The annotated speech corpus gives wide range of...
View Dataset →SKU: LDCIL-156
78:45:33 Hours | 49.2 GB | 306 Speakers | 45,198 Audio Segments | 48 kHz | 16 bit wav Maithili...
View Dataset →SKU: LDCIL-254
109:09:50 hours | 206 Audio Segments | 122 Speakers The LDC-IL Maithili Raw Speech dataset Vol.II comprises audio files in wav format,...
View Dataset →SKU: LDCIL-238
Dataset Description: 41:54:30 hours | 26 GB | 21,412 Audio Segments | 300 speakers The annotated speech corpus gives wide range of...
View Dataset →SKU: LDCIL-258
41:54:30 hours | 26 GB | 21,412 Audio Segments | 300 speakers The LDC-IL Maithili Sentence Aligned Speech Corpus(Tirhuta Script)...
View Dataset →SKU: LDCIL-263
30:59:20 hours | 19.56 GB | 32260 Audio Segments | 2 Speakers The LDC-IL Maithili Text to Speech dataset comprises audio files in wav...
View Dataset →SKU: LDCIL-157
164:01:02 Hours | 105 GB | 458Speakers| 43670 Audio Segments |48 kHz | 16 bit wav. Malayalam is the official language of Kerala and...
View Dataset →SKU: LDCIL-239
Dataset Description: 123:29:55 hours | 79.6 GB | 89,269 Audio Segments | 451 speakers The annotated speech corpus gives wide range of...
View Dataset →SKU: LDCIL-158
156:28:32 hours | 100 GB | 620 Speakers | 66,231 Audio segments | 48 khz | 16 bit wav Manipuri is the Administrative Language...
View Dataset →