Datasets
Dataset
lsd_all_scripts_dataset
by blue-machines
blue-machines/lsdallscriptsdataset Language-switch detection (LSD) training corpus (trainingstage2lsd. csv).
No licensetext-classification 29 01M < n < 10M rows
published 27 Jul 2026
Preview
First 5 rows of default / train, from the Hugging Face dataset viewer.
| language_switch | text | script_type | language |
|---|---|---|---|
| no_switch | काय चाललंय? | native | Marathi |
| no_switch | ठीक आहे. | native | Marathi |
| no_switch | आज काय करायचं? | native | Marathi |
| no_switch | 1. हे बघ. | native | Marathi |
| no_switch | घर आहे का? | native | Marathi |
Cite this
No licensetextCite
Tags
task_categories:text-classificationlanguage:enlanguage:hilanguage:bnlanguage:gulanguage:knlanguage:mllanguage:mrlanguage:orlanguage:talanguage:telicense:othersize_categories:1M<n<10Mformat:parquetmodality:textlibrary:datasetslibrary:pandaslibrary:polarslibrary:mlcroissantregion:uslanguage-switch-detectionlanguage-identificationindiccode-mixedromanized