Konkani Parts of Speech Annotated Corpus
SKU: LDCIL-423 | Model: 1700 | ISBN: 978-81-69175-91-3
Description & Overview
569408 Tags| 468967 Words | 53161 Sentences
The Linguistic Data Consortium for Indian Languages (LDC-IL) is developed Parts-of-Speech annotated corpus for Scheduled Indian languages. The corpus is annotated with Part-of-Speech (PoS) tags based on the Bureau of Indian Standards (BIS) PoS Tagset. This data is a significant resource for natural language processing and linguistic research. LDC-IL developed annotated text corpora for Konkani. The Konkani PoS annotated corpus is automatically tagged and then verified by linguistic experts to ensure accuracy and consistency.
Konkani PoS annotated Corpus contains 569408 Part-of-Speech tags.
For any research-based citations, please use the following citations:
1. Mr. Saurabh Varik, Dr. Narayan Choudhary. 2026. Konkani Parts of Speech Annotated Corpus. Central Institute of Indian Languages, Mysore. 978-81-69175-91-3.
2. Rejitha K. S. and Narayan Kumar Choudhary. (ed.). 2026. LDC-IL Parts of Speech Annotated Corpus Based on BIS Framework. Central Institute of Indian Languages, Mysore. 978-81-69175-60-9.
Dataset Specifications
| Authors | Mr. Saurabh Varik, Dr. Narayan Choudhary |
|---|---|
| Catalogue Number | 1700 |
| ISBN | 978-81-69175-91-3 |
| Data Source | Annotated |
| Word Count | 468967 |
| Release Date | 15/09/2026 |
| Terms and Conditions | General instructions for use of the resources provided by LDC-IL. |
| Tag Count | 569408 |