Datasets
Dataset
lid_all_scripts_dataset
by blue-machines
blue-machines/lidallscriptsdataset All-scripts language identification corpus derived from unifieddataset. csv.
No licensetext-classification 34 01M < n < 10M rows
published 27 Jul 2026
Preview
First 5 rows of default / train, from the Hugging Face dataset viewer.
| id | text | type | language |
|---|---|---|---|
| hi_314194 | Yimchunger boli ko Google ka samarthan milta hai | romanized_text | Hindi |
| hi_314194 | यिमचुंगर डायलेक्ट को गूगल का सपोर्ट है | native_script | Hindi |
| hi_314194 | Yimchunger dialect को Google का support है | code_mixed | Hindi |
| hi_314195 | Indira Gandhi Rashtriya Vridha Pension Yojna online portal ke madhyam se financial sahayata pradan karti hai | romanized_text | Hindi |
| hi_314195 | इंदिरा गांधी नेशनल ओल्ड एज पेंशन स्कीम ऑनलाइन पोर्टल्स के ज़रिए फाइनेंशियल सपोर्ट देती है। | native_script | Hindi |
Cite this
No licensetextCite
Tags
task_categories:text-classificationlanguage:enlanguage:hilanguage:bnlanguage:gulanguage:knlanguage:mllanguage:mrlanguage:orlanguage:talanguage:telicense:othersize_categories:1M<n<10Mformat:parquetmodality:textlibrary:datasetslibrary:dasklibrary:polarslibrary:mlcroissantregion:uslanguage-identificationindiccode-mixedromanizednative-script