Datasets

Dataset

indic-multilingual-tts-v2

by kapilkarda

Indic Multilingual TTS v2 Combined dataset for training multilingual Indian-language TTS models, specifically prepared for Spark-TTS BiCodec and LLM fine-tuning. Dataset Summary Metric Value Total samples 556,524 (550K train + 6.

CC-BY-4.0text-to-speech 51 1100K < n < 1M rows

published 19 Feb 2026

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

CC-BY-4.0Cite

Tags

task_categories:text-to-speechtask_categories:audio-classificationlanguage:aslanguage:bnlanguage:enlanguage:gulanguage:hilanguage:knlanguage:mllanguage:mrlanguage:nelanguage:orlanguage:palanguage:talanguage:telicense:cc-by-4.0size_categories:100K<n<1Mregion:usttsspeech-synthesisindian-languagesindicmultilingualemotionspark-ttsbicodec