Datasets

Dataset

Language_Indentification_v2

by Process-Venue

Dataset Card for Language Identification Dataset Sample Notebook: Kaggle Dataset link: Dataset Summary A comprehensive dataset for Indian language identification and text classification.

CC-BY-SA-4.0text-classification 47 2100K < n < 1M rows

published 24 Feb 2025

Preview

First 5 rows of default / train, from the Hugging Face dataset viewer.

HeadlineLanguage
ବ୍ୟାଙ୍କ୍‌ ପୁନର୍ପୁଞ୍ଜିକରଣରେ ଭାଗ ନେଇପାରେ ଭାରତୀୟ ଜୀବନ ବୀମା ନିଗମ Odia
सरकारले केही गरी ल्याउन सकेन चिनीNepali
फ्लैट पर बुलाता था कोच, मेडल व सरकारी नौकरी का झांसा देकर किया रेप: राजस्थान में महिला शूटर्सHindi
अशे स्थितींत घट्ट हातांनी आनी चुकीच्या पद्दतीन कातीचेर चणाचें पिठ वापरप टाळचें.Konkani
انگريزيءَ ۾ ڇو ڳالهايان ۽ جي هُن گهر ۾ سنڌي نه سکي ته پوءِ هُو ڪٿان سنڌي سکندو؟Sindhi

Cite this

CC-BY-SA-4.0textCite

Tags

task_categories:text-classificationlanguage:hilanguage:enlanguage:mrlanguage:palanguage:nelanguage:sdlanguage:aslanguage:gulanguage:talanguage:telanguage:urlanguage:orlicense:cc-by-sa-4.0size_categories:100K<n<1Mformat:parquetmodality:textlibrary:datasetslibrary:pandaslibrary:mlcroissantlibrary:polarsregion:usmultilinguallanguage-identificationtext-classificationindian