Datasets

Dataset

SpeechArenaBench

by ai4bharat

SpeechArenaBench SpeechArenaBench is a large-scale human-preference dataset for evaluating multilingual Text-to-Speech (TTS) systems across 10 Indian languages. It accompanies the paper "Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation a…

MITother 256 8100K < n < 1M rows

published 30 Apr 2026

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

MITaudiotextCite

Tags

language:bnlanguage:gulanguage:hilanguage:knlanguage:mllanguage:mrlanguage:orlanguage:talanguage:telanguage:urlicense:mitsize_categories:100K<n<1Mformat:parquetformat:optimized-parquetmodality:audiomodality:textlibrary:datasetslibrary:dasklibrary:polarslibrary:mlcroissantarxiv:2604.21481doi:10.57967/hf/9219region:ustext-to-speechttsspeech-synthesisevaluationhuman-preferencebradley-terryindic-languagesmultilingualcode-mixingaudio