Maithili Parts of Speech Annotated Corpus

SKU: LDCIL-424 | Model: 1699 | ISBN: 978-81-69175-29-6

Description & Overview


427438 Tags| 371983 Words | 25431 Sentences


The Linguistic Data Consortium for Indian Languages (LDC-IL) is developed Parts-of-Speech annotated corpus for Scheduled Indian languages. The corpus is annotated with Part-of-Speech (PoS) tags based on the Bureau of Indian Standards (BIS) PoS Tagset. This data is a significant resource for natural language processing and linguistic research. LDC-IL developed annotated text corpora for Maithili. The Maithili PoS annotated corpus is automatically tagged and then verified by linguistic experts to ensure accuracy and consistency.
Maithili PoS annotated Corpus contains 427438 Part-of-Speech tags.

For any research-based citations, please use the following citations:

1. Dinesh Mishra, Dr. Narayan Choudhary, Rajesha N., Manasa G. 2026. Maithili Parts of Speech Annotated Corpus. Central Institute of Indian Languages, Mysore- 978-81-69175-29-6.

2. Rejitha K. S. and Narayan Kumar Choudhary. (ed.). 2026. LDC-IL Parts of Speech Annotated Corpus Based on BIS Framework. Central Institute of Indian Languages, Mysore. 978-93-48633-33-0.

Dataset Specifications

Authors Dinesh Mishra, Dr. Narayan Choudhary
Corpus Type Parts of Speech Annotated Text Corpus
Catalogue Number 1699
ISBN 978-81-69175-29-6.
Data Source Annotated
Word Count 371983
Release Date 15/09/2026
Terms and Conditions General instructions for use of the resources provided by LDC-IL.
Tag Count 427438

Documentation & Sample Files