Datasets

Dataset

Vanipedia

by ai4bharat

BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages Overview BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English. This reposito…

CC-BY-4.0other 28 210K < n < 100K rows

published 13 Nov 2024

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

CC-BY-4.0audiotextCite

Tags

multilinguality:multilinguallanguage:aslanguage:bnlanguage:gulanguage:hilanguage:knlanguage:mllanguage:mrlanguage:nelanguage:orlanguage:palanguage:talanguage:telanguage:sdlanguage:urlanguage:enlicense:cc-by-4.0size_categories:10K<n<100Kformat:parquetmodality:audiomodality:textlibrary:datasetslibrary:pandaslibrary:mlcroissantlibrary:polarsarxiv:2411.04699region:us