Real users do not speak like scripts.
They interrupt, self-correct, mix languages inside a single sentence, use regional accents, give short replies, and speak in imperfect environments.
A model trained or evaluated on clean, single-language, studio-recorded speech meets little of that in production. A scripted demo rarely exercises the gap. Live traffic does.
"Bhai, delivery kab tak aayegi? Customer wait kar raha hai."
"Panch minute mein pahunch jayega. Traffic mein thoda delay ho gaya tha."
Illustrative specimen — control recorded internally
Speech structure matters
One-speaker and two-speaker audio are Sonexis’s current speech focus. The required structure is matched to the model and deployment scenario rather than treated as interchangeable.
One speaker
Isolated speech and speaker-specific evaluation
Isolated prompts, intent phrases, read speech where the use case needs it, spontaneous monologue, and pronunciation or acoustic evaluation against a single voice.
Two speakers
Dialogue, voice agents and real interaction
Natural dialogue, customer and support interaction, interviewer and respondent flows, voice-agent evaluation, turn-taking, interruptions, corrections, code-switching, short answers and context changes.
Three or more
Assessed against the deployment need
For three-or-more-speaker requirements, Sonexis can assess existing supply, sourcing or buyer-funded collection feasibility where the deployment need and demand justify the additional participant, diarisation and QA complexity.
Three routes, in this order
Every requirement is checked against them in sequence. Collection is the fallback, not the starting point.
Route 1
Reviewed existing supply
Where a dataset already exists and has cleared review for rights, provenance, metadata and quality, that is the fastest and cheapest route for a buyer.
Route 2
Managed sourcing
Where suitable data may exist outside anything Sonexis has reviewed, we locate a supplier who holds it and review their identity, rights and provenance before anything is licensed.
Route 3
Buyer-funded custom collection
When suitable existing data does not meet the requirement, Sonexis scopes and manages a funded collection directly or through approved collection partners, with defined consent, QA, metadata and delivery requirements.
How availability works
Sonexis does not operate an open public dataset catalogue. Availability is confirmed against each buyer requirement. Where supply is located and clears review, it can be shared privately against the brief, with evidence and controlled sample access. Every requirement is matched against reviewed supply and sourcing options before custom collection is considered.
What we check, and what stays declared.
A dataset is not usable because someone says it is good. It is usable when its origin, rights, consent, structure and defects are known — and when the parts that remain unknown are written down instead of omitted.
Six dimensions are reviewed separately. Each produces two answers: what was established, and what was not. There is no single quality score.
- Provenance
- Rights
- Consent
- Metadata
- Quality control
- Use-case fit
Where Sonexis collects, consent is task-specific and captured against a versioned instrument. The scopes granted are snapshotted onto the recording at submission, so a later edit to the consent text cannot retroactively widen what a past contributor agreed to. Export is gated against those snapshotted scopes and fails closed.
How verification works →Illustrative structure — not a live record
{ "consent_version": "recorded per contributor", "granted_scopes": [ "AI_TRAINING", "EVALUATION" ], "scopes_snapshotted_at_submission": true, "delivery_blocked_if_scope_missing": true, "withdrawal_request": "tracked to the recording" }
Every dataset arrives with its limitations in writing
A delivery contains audio, transcripts where scoped, structured metadata, consent references, QA records and a manifest — and a Dataset Passport recording what was reviewed, what was established, and what was not.
- Audio integrity
- Independently re-verified 2026-08-09 — 6 of 6 checksums matched the manifest
- Ownership
- Internal record — not independently verified
- Collection-time QA record
- Not established
Sixteen limitations are published with the specimen, not disclosed on request. Commercial status and the full evidence record are on the Dataset Passport page.
Starting with how India actually speaks
Sonexis focuses first on Indian English, Hindi, Hinglish, Punjabi and Marwadi, including code-switched speech. Exact availability, volume, rights and technical specifications are confirmed against each requirement.
- Current focusIndian English
- Current focusHindi
- Current focusHinglish
- Current focusPunjabi
- Current focusMarwadi
Beyond that focus, Sonexis reviews and sources against requirements in other languages and regions. A language we have not listed is not a language we refuse: it is one where the supplier, rights position and quality evidence still have to be established before anything is offered.
Availability is requirement-specific. Sonexis reviews suitable existing supply and sourcing routes against the brief before sharing controlled options.
Common questions
- What is Sonexis?
- Sonexis helps AI teams source, verify, evaluate, license and, where required, build realistic human datasets for multilingual production systems, starting with Indian and code-switched speech.
- What speech structures does Sonexis support?
- Sonexis focuses on one-speaker and two-speaker audio. Three-or-more-speaker requirements are assessed case by case, where the deployment need justifies the additional participant, attribution, diarisation and QA complexity.
- How does Sonexis verify data?
- Sonexis reviews provenance, rights, consent evidence, metadata, technical quality and use-case fit separately, and records the scope and the limitations in a Dataset Passport. Each dimension states what was established and what was not, rather than resolving to a single score.
- Does Sonexis only collect custom data?
- No. Sonexis checks suitable existing supply first, uses managed sourcing where needed, and scopes buyer-funded collection only when existing data cannot meet the requirement.
Tell us what your model needs
Describe the requirement. We check whether suitable existing data exists and can be licensed, then run managed sourcing, and scope funded collection only where existing data cannot meet the brief.
Hold data you can lawfully license? Apply for supplier review. Applying is a request for review, not verification.