Datasets

Dataset

lid_all_scripts_dataset

by blue-machines

blue-machines/lidallscriptsdataset All-scripts language identification corpus derived from unifieddataset. csv.

No licensetext-classification 34 01M < n < 10M rows

published 27 Jul 2026

Preview

First 5 rows of default / train, from the Hugging Face dataset viewer.

idtexttypelanguage
hi_314194Yimchunger boli ko Google ka samarthan milta hairomanized_textHindi
hi_314194यिमचुंगर डायलेक्ट को गूगल का सपोर्ट हैnative_scriptHindi
hi_314194Yimchunger dialect को Google का support हैcode_mixedHindi
hi_314195Indira Gandhi Rashtriya Vridha Pension Yojna online portal ke madhyam se financial sahayata pradan karti hairomanized_textHindi
hi_314195इंदिरा गांधी नेशनल ओल्ड एज पेंशन स्कीम ऑनलाइन पोर्टल्स के ज़रिए फाइनेंशियल सपोर्ट देती है।native_scriptHindi

Cite this

No licensetextCite

Tags

task_categories:text-classificationlanguage:enlanguage:hilanguage:bnlanguage:gulanguage:knlanguage:mllanguage:mrlanguage:orlanguage:talanguage:telicense:othersize_categories:1M<n<10Mformat:parquetmodality:textlibrary:datasetslibrary:dasklibrary:polarslibrary:mlcroissantregion:uslanguage-identificationindiccode-mixedromanizednative-script