A Gold Standard Assamese Raw Text Corpus
SKU: LDCIL-191
1,01,27,030 Words | 1,084 Tittles | XML format | 6 domains Assamese or Oxomiya is the language spoken by the natives of the state of Assam...
View Dataset →Raw Corpus for Text
SKU: LDCIL-191
1,01,27,030 Words | 1,084 Tittles | XML format | 6 domains Assamese or Oxomiya is the language spoken by the natives of the state of Assam...
View Dataset →SKU: LDCIL-196
42,37,440 Words | 1,460 Tittles | XML format | 3 domains Bengali is the official language of West Bengal and Tripura. It belongs to the...
View Dataset →SKU: LDCIL-180
29,15,544 Words | 80 Tittles | XML format | 5 domains Bodo is a major tribal language that belongs to the Tibeto-Burman language family....
View Dataset →SKU: LDCIL-251
22,19,592 Words | 55 Titles | XML format | 4 Domains | 28 Sub-categories Chhattisgarhi, a tongue of approximately 17 million people,...
View Dataset →SKU: LDCIL-179
8,01,771 Words | 183 Tittles | XML format | 05 Text Domains Dogri is an Indo-Aryan language spoken by about five...
View Dataset →SKU: LDCIL-181
28, 62,413 Words | 1,364 Tittles | XML format | 06 Text Domains Gujarati is a major Indo-Aryan language and the...
View Dataset →SKU: LDCIL-64
1,03,17,177 Words | 1,223 Tittles | XML format | 4 domains Hindi is a Major, Indo-Aryan language,...
View Dataset →SKU: LDCIL-TXT-HI-RAW-001
A curated raw text corpus for Hindi language research, NLP experimentation, and linguistic analysis.
View Dataset →SKU: LDCIL-184
77,63,124 words | 1772 Titles | Data and Metadata in XML format | 6 text domains Kannada is one of the Ancient Indian language which...
View Dataset →SKU: LDCIL-152
4,66,054 Words | 108 Tittles | XML format | 2 domains Kashmiri language is one of the 22 scheduled languages of India and is the part...
View Dataset →SKU: LDCIL-257
10, 13,658 words | 123 Titles | XML format | 6 domains |59 sub-categories A Gold Standard Kashmiri Raw Text Corpus Vol. II is a...
View Dataset →SKU: LDCIL-186
39,95,611 Words | 282 Tittles | XML format | 4 domains Konkani is the principal and administrative language of Goa. Konkani is an...
View Dataset →SKU: LDCIL-188
53,16,552 Words | 499 Tittles | XML format | 5 domains Maithili is an Indio-Aryan language, a direct descendant of Sanskrit. Which is...
View Dataset →SKU: LDCIL-253
8,11,680 Words | 54 Titles | XML format | 3 Domains | 21 Sub-categories The Maithili Raw Text Corpus endows an unrivaled window in...
View Dataset →SKU: LDCIL-185
63, 70,954 Words | 1,119 Titles | XML format | 6 domains Malayalam is a highly agglutinative and morphologically rich language. The actual...
View Dataset →SKU: LDCIL-183
61,45,278 words | 4,31,27,842 characters | 6 Domains Manipuri Text Corpus is encoded in a machine-readable form and...
View Dataset →SKU: LDCIL-164
21,57,109 Words | 678 Tittles | XML format | 5 domains Marathi is an Indo-Aryan language. It is the official language of Maharashtra state...
View Dataset →SKU: LDCIL-133
70,57,524 Words | 1,347 Tittles | XML format | 6 domains Nepali is one of the official language of West Bengal and Sikkim state. It...
View Dataset →SKU: LDCIL-132
15, 88, 287 Words | 206 Titles | XML format | 05 Text Domains Odia ( formerly Oriya) is a major Indo-Aryan language, which is...
View Dataset →SKU: LDCIL-131
1,01,25,770 Words | 2,470 Tittles | XML format | 5 domains Punjabi is the principal and administrative language of Punjab....
View Dataset →SKU: LDCIL-250
11,99,502 Words | 74 Titles | XML format | 3 Domains | 27 Sub-categories Rajasthani is a broad linguistic category that encompasses a...
View Dataset →SKU: LDCIL-189
1,09,31,902 Words | 1,963 Titles | XML format | 6 text domains Tamil is one of the longest-surviving...
View Dataset →SKU: LDCIL-169
30,10,993 Words | 859 Titles | XML format | 6 Domains Telugu is a highly agglutinative and morphologically rich language. The actual...
View Dataset →SKU: LDCIL-252
30,13,530 Words | 160 Titles | XML format | 6 Domains | 29 Sub-categories Telugu is a highly agglutinative and morphologically rich...
View Dataset →SKU: LDCIL-176
5161927 Words | 739 Titles | XML format | 5 domains. Urdu is one of the prominent language used in the Indian sub-continent. It...
View Dataset →SKU: LDCIL-433
8,16,073 | Words |619,0666 characters | 55 Titles Tulu has about 18.4 lakh (1.84 million) speakers. Although it has a rich literary and...
View Dataset →