Datasets
Dataset
indic-multilingual-tts-v2
by kapilkarda
Indic Multilingual TTS v2 Combined dataset for training multilingual Indian-language TTS models, specifically prepared for Spark-TTS BiCodec and LLM fine-tuning. Dataset Summary Metric Value Total samples 556,524 (550K train + 6.
CC-BY-4.0text-to-speech 51 1100K < n < 1M rows
published 19 Feb 2026
Preview
The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.
Cite this
CC-BY-4.0Cite
Tags
task_categories:text-to-speechtask_categories:audio-classificationlanguage:aslanguage:bnlanguage:enlanguage:gulanguage:hilanguage:knlanguage:mllanguage:mrlanguage:nelanguage:orlanguage:palanguage:talanguage:telicense:cc-by-4.0size_categories:100K<n<1Mregion:usttsspeech-synthesisindian-languagesindicmultilingualemotionspark-ttsbicodec