Language coverage

These are the languages Sonexis works in.

Sonexis focuses first on Indian English, Hindi, Hinglish, Punjabi and Marwadi, including code-switched speech. Exact availability, volume, rights and technical specifications are confirmed against each requirement.

How availability works

Availability is requirement-specific. Sonexis does not operate an open public dataset catalogue: reviewed supply, sourcing routes and collection feasibility are confirmed privately after scope review.

What each status means

Reviewed supply available
A specific dataset has cleared review for rights, provenance, consent, metadata and quality, and is available now.
Current focus
A language Sonexis actively works in and actively develops sourcing and collection routes for. Availability, volume and specification are confirmed against the brief.
May be sourced
Data may exist outside what Sonexis has reviewed. We would locate and review it against your brief.
Custom collection capability
Can be collected to your brief once the requirement is funded, when existing data cannot meet it.
Current focus

Languages we actively work in

These five are where Sonexis focuses first, including code-switched speech. Each can be sourced against a brief or collected to a funded one.

Indian English

EN-IN

Use cases

ASR evaluation, voice agents, support conversations, multilingual benchmarks.

Speech behaviour

Regional variation, Indian phrasing, mixed vocabulary, informal flow.

Hindi

HI

Use cases

ASR training, customer support, onboarding, conversational AI.

Speech behaviour

Regional accent variation, formal and informal switching, short replies.

Hinglish

HI-EN

Use cases

Voice agents, support, product discovery, real-world multilingual testing.

Speech behaviour

Natural Hindi-English switching, mixed sentence structure, informal phrasing.

Punjabi

PA

Use cases

Regional speech testing, multilingual voice datasets, support scenarios.

Speech behaviour

Accent variation, mixed Hindi or English flow, informal speech.

Marwadi

MWR

Use cases

Regional voice data, underrepresented language testing, local conversation flows.

Speech behaviour

Regional vocabulary, Hindi-adjacent expressions, and longer contextual phrasing less common in standard training data.

May be sourced · Custom collection capability

Additional languages that may be sourced or supported through funded collection

Data may exist outside what Sonexis has reviewed. We would locate and review it against your brief, or collect to specification once the requirement is funded.

Languages

Tamil Marathi

Other Indian and regional languages — including Telugu, Kannada, Malayalam, Bengali and Gujarati — are assessed against the brief. Send the requirement and we will assess sourcing and collection routes for it.

Code-switched combinations

Code switching is natural in Indian speech. Models and evaluation sets that treat languages in isolation can fail on within-utterance switches, which are common in real customer conversations.

Hindi-English Hindi-Marwadi Tamil-English Hindi-Punjabi
How availability is confirmed

Availability is confirmed against your requirement, not published as a catalogue.

Sonexis does not operate an open public dataset catalogue. Where supply is located and clears review, it can be shared privately against the brief, with evidence and controlled sample access. Where suitable supply is not found, Sonexis runs managed sourcing or scopes buyer-funded collection.

Where a dataset clears provenance, rights, consent, metadata and quality review and is genuinely available, it moves to Reviewed supply available and is published with a Dataset Passport.

See what a Dataset Passport records →
Why language design matters

India is not a single-language market.

Real conversations often shift between languages within a single utterance, not just across turns. Users interrupt, correct themselves, reply briefly, speak in noisy environments, and mix language and register depending on context and speaker. A model or evaluation set built on clean single-language speech can miss these patterns entirely.

Language choice in a voice dataset is a design decision. It should reflect how the target user actually speaks, including accent variation, informal phrasing, regional vocabulary, and natural conversational behaviour, not what is easiest to collect.

Request language coverage.

Tell us the languages you need, the use case, and the volume. We will confirm whether coverage is available or can be built.

Submit a requirement