A Gold Standard Chhattisgarhi Raw Text Corpus
OverView
Dataset Description: 14,74,496 Words | 51 Titles | XML format | Aesthetics and Mass Media DomainChhattisgarhi, a tongue of approximately 17 million people, carries profound cultural and historical significance within ...Your request cart is empty!
Dataset Description
Dataset Description:
14,74,496 Words | 51 Titles | XML format | Aesthetics and Mass Media Domain
Chhattisgarhi, a tongue of approximately 17 million people, carries profound cultural and historical significance within the region of Chhattisgarh. The Chhattisgarhi Raw Text Corpus endows an unrivaled window in documenting the colloquialisms, idioms, regional vocabularies, and grammar that are essential to establishing frameworks for linguistic processing. The Chhattisgarhi Raw Text Corpus is an extensive repository encapsulating the viable linguistic elements of Chhattisgarhi textual materials.
The corpus of Chhattisgarhi text can be broadly classified as literary and non-literary texts. Data has been collected from books, magazines, newspapers and websites and it is verified to be true to the original texts and then warehoused. Chhattisgarhi Text Corpus encoded in a machine-readable form and stored in a standard format. The major encoding being used is Unicode and stored in XML format. The data is embedded with metadata information. The corpus has been created from the contemporary text in typed and crawled methods.
The available Text Corpus details: Domains Words Percentage of Total Corpus Aesthetics 14,35,667 (97.04 %) Mass Media 38,829 (2.6 %). A detailed explanation of the Chhattisgarhi Text Corpus will be available in the Chhattisgarhi Raw Text Corpus Documentation.
For any research-based citations, please use the following citations:
1. Ankita Tiwari, Satyaendra Kumar Awasthi, Narayan Kumar Choudhary 2023. A Gold Standard Chhattisgarhi Raw Text Corpus. Central Institute of Indian Languages, Mysore.
2. Rejitha K. S. and Narayan Kumar Choudhary. (ed.). 2023. Compendium of LDC-IL Sentence Aligned Speech Corpus. Central Institute of Indian Languages, Mysore. ISBN: 978-81-19411-34-4.
3. Choudhary, Narayan & L. Ramamoorthy. 2019. "LDC-IL Raw Text Corpora: An Overview" in Linguistic Resources for AI/NLP in Indian Languages, Central Institute of Indian Languages, Mysore. pp. 1-10.
Item specifics
- Authors Ankita Tiwari, Satyaendra Kumar Awasthi, Rajesha N., Manasa G., Srikanth D., Narayan Kumar Choudhary, Shailendra Mohan
- Corpus Type Raw Text Corpus
- Catalogue Number 1437
- ISBN 978-81-19411-64-1
- Data Source Typed + Cleaned
- Character Count 6918702
- Word Count 14,74,496
- Release Date 08-Jan-24
- Terms and Conditions General instructions for use of the resources provided by LDC-IL.