Datasets

Dataset

IndicVoices-ST

by ai4bharat

BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages Overview BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English. This reposito…

CC-BY-4.0other 210 211M < n < 10M rows

published 10 Nov 2024

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

CC-BY-4.0audiotextCite

Tags

multilinguality:multilinguallanguage:aslanguage:bnlanguage:gulanguage:hilanguage:knlanguage:mllanguage:mrlanguage:nelanguage:orlanguage:palanguage:talanguage:telanguage:urlicense:cc-by-4.0size_categories:1M<n<10Mformat:parquetmodality:audiomodality:textlibrary:datasetslibrary:dasklibrary:mlcroissantlibrary:polarsarxiv:2411.04699region:us