A Gold Standard Chhattisgarhi Raw Text Corpus

0 reviews requests (16)

Owner Central Institute of Indian Languages

Catalogue Number: 1437

Stock In Stock

OverView

Dataset Description: 14,74,496 Words | 51 Titles | XML format | Aesthetics and Mass Media DomainChhattisgarhi, a tongue of approximately 17 million people, carries profound cultural and historical significance within ...

Please Login to see the price

Tags: Chhattisgarhi Raw text Corpus

Categories Cart Account Search Recent View Go to Top

Dataset Description

Dataset Description:

14,74,496 Words | 51 Titles | XML format | Aesthetics and Mass Media Domain

Chhattisgarhi, a tongue of approximately 17 million people, carries profound cultural and historical significance within the region of Chhattisgarh. The Chhattisgarhi Raw Text Corpus endows an unrivaled window in documenting the colloquialisms, idioms, regional vocabularies, and grammar that are essential to establishing frameworks for linguistic processing. The Chhattisgarhi Raw Text Corpus is an extensive repository encapsulating the viable linguistic elements of Chhattisgarhi textual materials.

The corpus of Chhattisgarhi text can be broadly classified as literary and non-literary texts. Data has been collected from books, magazines, newspapers and websites and it is verified to be true to the original texts and then warehoused. Chhattisgarhi Text Corpus encoded in a machine-readable form and stored in a standard format. The major encoding being used is Unicode and stored in XML format. The data is embedded with metadata information. The corpus has been created from the contemporary text in typed and crawled methods.

The available Text Corpus details: Domains Words Percentage of Total Corpus Aesthetics 14,35,667 (97.04 %) Mass Media 38,829 (2.6 %). A detailed explanation of the Chhattisgarhi Text Corpus will be available in the Chhattisgarhi Raw Text Corpus Documentation.

For any research-based citations, please use the following citations:

1. Ankita Tiwari, Satyaendra Kumar Awasthi, Narayan Kumar Choudhary 2023. A Gold Standard Chhattisgarhi Raw Text Corpus. Central Institute of Indian Languages, Mysore.

2. Rejitha K. S. and Narayan Kumar Choudhary. (ed.). 2023. Compendium of LDC-IL Sentence Aligned Speech Corpus. Central Institute of Indian Languages, Mysore. ISBN: 978-81-19411-34-4.

3. Choudhary, Narayan & L. Ramamoorthy. 2019. "LDC-IL Raw Text Corpora: An Overview" in Linguistic Resources for AI/NLP in Indian Languages, Central Institute of Indian Languages, Mysore. pp. 1-10.

Item specifics

Authors Ankita Tiwari, Satyaendra Kumar Awasthi, Rajesha N., Manasa G., Srikanth D., Narayan Kumar Choudhary, Shailendra Mohan
Corpus Type Raw Text Corpus
Catalogue Number 1437
ISBN 978-81-19411-64-1
Data Source Typed + Cleaned
Character Count 6918702
Word Count 14,74,496
Release Date 08-Jan-24
Terms and Conditions General instructions for use of the resources provided by LDC-IL.

A Gold Standard Chhattisgarhi Raw Text Corpus

OverView

A Gold Standard Chhattisgarhi Raw Text Corpus

Dataset Description

Item specifics

Write a review