Datasets

Dataset

vaani-non-null

by adjaysagar

Vaani by Language — Non-Null Transcript Subset Processed output from ARTPARK-IISc/Vaani: transcript-not-null filtered, reorganized by language. Hours computed from WAV duration in each parquet (run computevaanihours.

CC-BY-4.0automatic-speech-recognition 18 01M < n < 10M rows

published 9 Mar 2026

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

CC-BY-4.0textCite

Tags

task_categories:automatic-speech-recognitiontask_categories:text-to-speechlanguage:hilanguage:bnlanguage:telanguage:knlanguage:mrlanguage:orlanguage:talanguage:mllanguage:nelanguage:enlanguage:aslanguage:palanguage:urlanguage:gulanguage:multilinguallicense:cc-by-4.0size_categories:1M<n<10Mformat:parquetmodality:textlibrary:datasetslibrary:dasklibrary:polarslibrary:mlcroissantregion:usspeechasrmultilingualvaaniindia