Kannada Sentence Aligned Speech Corpus

0 reviews requests (8)

Owner Central Institute of Indian Languages

Catalogue Number: 1423

Stock In Stock

OverView

Dataset Description: 107:48:50 hours | 69.4 GB | 65,533 Audio Segments | 600 speakersThe annotated speech corpus gives wide range of linguistic information especially useful to analyse phonetics. The LDC-IL Kannada Se...

Please Login to see the price

Tags: Kannada Sentence Aligned Speech Corpus

Categories Cart Account Search Recent View Go to Top

Dataset Description

Dataset Description:

107:48:50 hours | 69.4 GB | 65,533 Audio Segments | 600 speakers

The annotated speech corpus gives wide range of linguistic information especially useful to analyse phonetics. The LDC-IL Kannada Sentence Aligned Speech dataset comprises audio files in wav format, accompanied by a corresponding textual layer containing phonetically normalized and orthographically normalized annotations in Kannada script. This dataset spans a duration of 107:48:50 (hh:mm:ss), consisting of read speech with continuous text, representative sentences, and date formats. The data is derived from 300 female and 300 male native Kannada speakers, encompassing diverse age groups and regions. A comprehensive explanation of dataset can be found in the Kannada Sentence Aligned Speech Documentation.

For any research-based citations, please use the following citations:

1. Vijayalaxmi F. Patil, Chetan Baji, Kavitha Lenin, Reshma S., Rajesha N., Manasa G., Srikanth D., Stephen Fernandes, Nithin S., Narayan Kumar Choudhary, Shailendra Mohan. 2023. Kannada Sentence Aligned Speech Corpus Central Institute of Indian Languages, Mysore. 978-81-19411-19-1.

2. Rejitha K. S. and Narayan Kumar Choudhary. (ed.). 2023. Compendium of LDC-IL Sentence Aligned Speech Corpus. Central Institute of Indian Languages, Mysore. ISBN: 978-81-19411-34-4.

3. Choudhary, N. 2021. LDC-IL: The Indian Repository of Resources for Language Technology. Language Resources & Evaluation. Springer, Vol. 55, Issue 1. doi: https://doi.org/10.1007/s10579-020-09523-3

Item specifics

Authors Vijayalaxmi F. Patil, Chetan Baji, Kavitha Lenin, Reshma S., Rajesha N., Manasa G., Srikanth D., Stephen Fernandes, Nithin S., Narayan Kumar Choudhary, Shailendra Mohan
Corpus Type Sentence Annotated Corpus
Catalogue Number 1423
ISBN 978-81-19411-19-1
Data Source On Field
Duration 107:48:50
# of Audio Segments 65533
Release Date 08-01-2024
Terms and Conditions General instructions for use of the resources provided by LDC-IL.

Kannada Sentence Aligned Speech Corpus

OverView

Kannada Sentence Aligned Speech Corpus

Dataset Description

Item specifics

Write a review