Datasets
Dataset
sft_saaras_flow
by sarvam
sftsaarasflow Synthetic SFT data for ASR transcription correction, generated to distill a large model (Gemini 3. 1 Pro) into a smaller correction model.
Apache-2.0text-generation 34 01K < n < 10K rows
published 11 Jun 2026
Preview
First 5 rows of default / train, from the Hugging Face dataset viewer.
| language | user_transcript | corrected_transcript | mode | style_preset | language_hint | app_category | detected_patterns | faithfulness_enabled | high_polish | english_full_style | memory_hints | persona_bias | persona | variation | detected_languages | correction_prompt | step1_meta | step3_meta | error | messages |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| english | hey guys i looked at the odd bug um it seems like the j w t token is expiring too early lets extend it to twenty four hours wait no twelve hours should be enough ill push a fix to staging in like ten minutes | Hey guys, I looked at the Auth bug. It seems like the JWT token is expiring too early. Let's extend it to 12 hours. I'll push a fix to staging in 10 minutes. | transcribe | casual | en-US | chat_messaging | ["technical_content"] | true | false | true | [{"type":"correction","wrong_form":"odd bug","preferred_form":"auth bug"}] | User is a backend developer; uses standard tech acronyms. | {"role":"Backend Software Engineer","device":"laptop","geography":"San Francisco, USA","scenario":"Dictating a quick status update to a team Slack channel about a bug fix","speech_style":"Casual technical, fast-paced"} | {"length_class":"medium","has_self_correction":true,"speaker_errors":["filler (um)","filler (like)"],"asr_errors":["odd bug (auth bug)","j w t (JWT)"],"notes":"Standard tech worker Slack update with a domain-specific homophone error and self-correction."} | ["en-US"] | ABSOLUTE TOP-PRIORITY RULE — strict Latin-script output (user selected: English): The user has explicitly selected English as their primary language. - The final output MUST be entirely in Latin/Roman script. ZERO Devanagari, Tamil, Telugu, Bengali, Gujarati, Kannada, Malayalam, Punjabi, or Odia ch… | {"usage":{"prompt_tokens":1998,"candidates_tokens":395,"cached_tokens":null,"thoughts_tokens":784},"warnings":[],"round":1} | {"usage":{"prompt_tokens":9691,"candidates_tokens":54,"cached_tokens":null,"thoughts_tokens":1594}} | [{"role":"system","content":"ABSOLUTE TOP-PRIORITY RULE — strict Latin-script output (user selected: English):\nThe user has explicitly selected English as their primary language.\n- The final output MUST be entirely in Latin/Roman script. ZERO Devanagari, Tamil, Telugu, Bengali, Gujarati, Kannada,… | |
| english | hey sarah i just looked at your pr for the off service um it looks mostly good but we need to actually wait can you add a few more unit tests around the token refresh logic let me know when its done and ill approve | Hey Sarah, I just looked at your PR for the Auth service, it looks mostly good but can you add a few more unit tests around the token refresh logic? Let me know when it's done and I'll approve | transcribe | casual | en-US | chat_messaging | ["technical_content"] | true | false | true | [{"type":"correction","wrong_form":"off service","preferred_form":"auth service"}] | User is a backend engineer; prefer standard tech abbreviations. | {"role":"Backend Software Engineer","device":"laptop","geography":"San Francisco, USA","scenario":"Dictating a quick PR review update via Slack to a coworker.","speech_style":"Casual, technical register, slightly run-on with mid-thought pivots."} | {"length_class":"medium","has_self_correction":true,"speaker_errors":["filler ('um')","false start/pivot ('but we need to actually wait')"],"asr_errors":["homophone error ('off' instead of 'auth')","lack of punctuation and capitalization"],"notes":"Medium length Slack message dictation containing t… | ["en-IN"] | ABSOLUTE TOP-PRIORITY RULE — strict Latin-script output (user selected: English): The user has explicitly selected English as their primary language. - The final output MUST be entirely in Latin/Roman script. ZERO Devanagari, Tamil, Telugu, Bengali, Gujarati, Kannada, Malayalam, Punjabi, or Odia ch… | {"usage":{"prompt_tokens":1998,"candidates_tokens":418,"cached_tokens":null,"thoughts_tokens":735},"warnings":[],"round":1} | {"usage":{"prompt_tokens":9691,"candidates_tokens":54,"cached_tokens":null,"thoughts_tokens":738}} | [{"role":"system","content":"ABSOLUTE TOP-PRIORITY RULE — strict Latin-script output (user selected: English):\nThe user has explicitly selected English as their primary language.\n- The final output MUST be entirely in Latin/Roman script. ZERO Devanagari, Tamil, Telugu, Bengali, Gujarati, Kannada,… | |
| english | "hey team uh the queue three dashboard role out is delayed until thursday actually wait let's make i | "Hey team, the Q3 dashboard rollout is delayed until Friday, we found a bug in the API response cach | transcribe | casual | en-US | chat_messaging | ["technical_content"] | true | false | true | [{"type":"entity","wrong_form":"zero","preferred_form":"Jira"}] | User is a tech PM; prefers standard agile terminology. | {"role":"Product Manager","device":"laptop","geography":"New York, USA","scenario":"Dictating a stat | {"length_class":"medium","has_self_correction":true,"speaker_errors":["filler words (uh, um)","self- | ["en-IN"] | "ABSOLUTE TOP-PRIORITY RULE — strict Latin-script output (user selected: English):\nThe user has e | {"usage":{"prompt_tokens":1998,"candidates_tokens":420,"cached_tokens":null,"thoughts_tokens":726}," | {"usage":{"prompt_tokens":9687,"candidates_tokens":39,"cached_tokens":null,"thoughts_tokens":841}} | [{"role":"system","content":"ABSOLUTE TOP-PRIORITY RULE — strict Latin-script output (user selecte | |
| english | "hey sara just wanted to confirm the viewing for 23 albert marl street is at uh 3 pm today wait actu | "Hey Sarah, just wanted to confirm the viewing for 23 Albemarle Street is at 3:30 PM today. I got he | transcribe | casual | en-IN | chat_messaging | [] | true | false | true | [{"type":"correction","wrong_form":"albert marl","preferred_form":"Albemarle"},{"type":"entity","wro | User is a real estate agent in London | {"role":"Real Estate Agent","device":"mobile","geography":"London, UK","scenario":"Dictating a quick | {"length_class":"medium","has_self_correction":true,"speaker_errors":["uh","wait actually make that | ["en-IN"] | "ABSOLUTE TOP-PRIORITY RULE — strict Latin-script output (user selected: English):\nThe user has e | {"usage":{"prompt_tokens":1998,"candidates_tokens":467,"cached_tokens":null,"thoughts_tokens":738}," | {"usage":{"prompt_tokens":9697,"candidates_tokens":65,"cached_tokens":null,"thoughts_tokens":1490}} | [{"role":"system","content":"ABSOLUTE TOP-PRIORITY RULE — strict Latin-script output (user selecte | |
| english | "hey team just wanted to give a quick update on the q 3 road map actually wait let's make it the q f | "Team, I wanted to provide a quick update on the Q4 roadmap. We are pushing the Auth refactor to the | transcribe | formal | en-US | ["email_structure","dictated_request"] | true | true | true | [{"type":"correction","wrong_form":"art refactor","preferred_form":"auth refactor"}] | User is a tech product manager; prefer standard Agile/tech terminology. | {"role":"Product Manager","device":"laptop","geography":"San Francisco, USA","scenario":"Dictating a | {"length_class":"medium","has_self_correction":true,"speaker_errors":["filler (ah)","self-correction | ["en-IN"] | "ABSOLUTE TOP-PRIORITY RULE — strict Latin-script output (user selected: English):\nThe user has e | {"usage":{"prompt_tokens":1998,"candidates_tokens":427,"cached_tokens":null,"thoughts_tokens":696}," | {"usage":{"prompt_tokens":9895,"candidates_tokens":48,"cached_tokens":null,"thoughts_tokens":693}} | [{"role":"system","content":"ABSOLUTE TOP-PRIORITY RULE — strict Latin-script output (user selecte |
Cite this
Apache-2.0textCite
Tags
task_categories:text-generationlanguage:enlanguage:hilanguage:mrlanguage:gulanguage:knlanguage:talanguage:telanguage:mllanguage:orlanguage:bnlanguage:palicense:apache-2.0size_categories:1K<n<10Kformat:parquetformat:optimized-parquetmodality:textlibrary:datasetslibrary:pandaslibrary:polarslibrary:mlcroissantregion:ussyntheticasrtranscription-correctionsftdistillation