Datasets
Dataset
IndicVoices-ST
by ai4bharat
BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages Overview BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English. This reposito…
CC-BY-4.0other 210 211M < n < 10M rows
published 10 Nov 2024
Preview
The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.
Cite this
CC-BY-4.0audiotextCite
Tags
multilinguality:multilinguallanguage:aslanguage:bnlanguage:gulanguage:hilanguage:knlanguage:mllanguage:mrlanguage:nelanguage:orlanguage:palanguage:talanguage:telanguage:urlicense:cc-by-4.0size_categories:1M<n<10Mformat:parquetmodality:audiomodality:textlibrary:datasetslibrary:dasklibrary:mlcroissantlibrary:polarsarxiv:2411.04699region:us