These are the languages Sonexis works in.
Sonexis focuses first on Indian English, Hindi, Hinglish, Punjabi and Marwadi, including code-switched speech. Exact availability, volume, rights and technical specifications are confirmed against each requirement.
How availability works
Availability is requirement-specific. Sonexis does not operate an open public dataset catalogue: reviewed supply, sourcing routes and collection feasibility are confirmed privately after scope review.
What each status means
- Reviewed supply available
- A specific dataset has cleared review for rights, provenance, consent, metadata and quality, and is available now.
- Current focus
- A language Sonexis actively works in and actively develops sourcing and collection routes for. Availability, volume and specification are confirmed against the brief.
- May be sourced
- Data may exist outside what Sonexis has reviewed. We would locate and review it against your brief.
- Custom collection capability
- Can be collected to your brief once the requirement is funded, when existing data cannot meet it.
Languages we actively work in
These five are where Sonexis focuses first, including code-switched speech. Each can be sourced against a brief or collected to a funded one.
Indian English
EN-INUse cases
ASR evaluation, voice agents, support conversations, multilingual benchmarks.
Speech behaviour
Regional variation, Indian phrasing, mixed vocabulary, informal flow.
Hindi
HIUse cases
ASR training, customer support, onboarding, conversational AI.
Speech behaviour
Regional accent variation, formal and informal switching, short replies.
Hinglish
HI-ENUse cases
Voice agents, support, product discovery, real-world multilingual testing.
Speech behaviour
Natural Hindi-English switching, mixed sentence structure, informal phrasing.
Punjabi
PAUse cases
Regional speech testing, multilingual voice datasets, support scenarios.
Speech behaviour
Accent variation, mixed Hindi or English flow, informal speech.
Marwadi
MWRUse cases
Regional voice data, underrepresented language testing, local conversation flows.
Speech behaviour
Regional vocabulary, Hindi-adjacent expressions, and longer contextual phrasing less common in standard training data.
Additional languages that may be sourced or supported through funded collection
Data may exist outside what Sonexis has reviewed. We would locate and review it against your brief, or collect to specification once the requirement is funded.
Languages
Other Indian and regional languages — including Telugu, Kannada, Malayalam, Bengali and Gujarati — are assessed against the brief. Send the requirement and we will assess sourcing and collection routes for it.
Code-switched combinations
Code switching is natural in Indian speech. Models and evaluation sets that treat languages in isolation can fail on within-utterance switches, which are common in real customer conversations.
Availability is confirmed against your requirement, not published as a catalogue.
Sonexis does not operate an open public dataset catalogue. Where supply is located and clears review, it can be shared privately against the brief, with evidence and controlled sample access. Where suitable supply is not found, Sonexis runs managed sourcing or scopes buyer-funded collection.
Where a dataset clears provenance, rights, consent, metadata and quality review and is genuinely available, it moves to Reviewed supply available and is published with a Dataset Passport.
See what a Dataset Passport records →India is not a single-language market.
Real conversations often shift between languages within a single utterance, not just across turns. Users interrupt, correct themselves, reply briefly, speak in noisy environments, and mix language and register depending on context and speaker. A model or evaluation set built on clean single-language speech can miss these patterns entirely.
Language choice in a voice dataset is a design decision. It should reflect how the target user actually speaks, including accent variation, informal phrasing, regional vocabulary, and natural conversational behaviour, not what is easiest to collect.
Request language coverage.
Tell us the languages you need, the use case, and the volume. We will confirm whether coverage is available or can be built.
Submit a requirement