Datasets
Dataset
Spoken-Tutorial
by ai4bharat
BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages Overview BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English. This reposito…
CC-BY-4.0other 58 2100K < n < 1M rows
published 13 Nov 2024
Preview
The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.
Cite this
CC-BY-4.0audiotextCite
Tags
multilinguality:multilinguallanguage:aslanguage:bnlanguage:gulanguage:hilanguage:knlanguage:mllanguage:mrlanguage:nelanguage:salanguage:orlanguage:palanguage:talanguage:telicense:cc-by-4.0size_categories:100K<n<1Mformat:parquetmodality:audiomodality:textlibrary:datasetslibrary:pandaslibrary:mlcroissantlibrary:polarsarxiv:2411.04699region:us