Manipuri Sentence Aligned Speech Corpus (Bengali Script)

SKU: LDCIL-261 | Model: 1503 | ISBN: 978-93-48633-68-2

Description & Overview

116:34:24 hours | 75.9 GB  | 60,819 Audio Segments | 589 speakers

The LDC-Manipuri Sentence Aligned Speech dataset comprises audio files in wav format, accompanied by a corresponding textual layer containing orthographically normalized annotation in Bengali script. This dataset spans a duration of 116:34:24 (hh:mm:ss), consisting of read speech with continuous text, representative sentences, and date formats. The data is derived from 295 female and 294 male native Manipuri speakers, encompassing diverse age groups and regions. A comprehensive explanation of dataset can be found in the Manipuri Sentence Aligned Speech Documentation. 


For any research-based citations, please use the following citations:

  1. Amom Nandaraj Meetei, Yumnam ,Premila Chanu, Rajesha N, Manasa,G, Stephen Fernandes, Nithin S, Roopashri M.R ,Dr. Narayan Kumar Choudhary, Prof. Shailendra Mohan. 2025. Manipuri Sentence Aligned Speech Corpus. Central Institute of Indian Languages, Mysore. 978-93-48633-68-2


Dataset Specifications

Authors Amom Nandaraj Meetei., Yumnam Premila Chanu., Rajesha N., Manasa G., Nithin S., Dr. Narayan Kumar Choudhary., Prof. Shailendra Mohan
Corpus Type Sentence Aligned Speech Corpus
Catalogue Number 1503
ISBN 978-93-48633-68-2
Data Source On Field
Duration 123:29:55 (hh:mm:ss)
# of Audio Segments 60819
Release Date 2025/03/20
Terms and Conditions General instructions for use of the resources provided by LDC-IL.

Documentation & Sample Files