Use cases

Voice data designed around real AI workflows.

Sonexis helps AI teams obtain human data for training, fine-tuning, evaluation and benchmarking across voice and conversational AI. Every use case below follows the same order: check suitable existing supply, locate and review potential suppliers if needed, and scope buyer-funded collection only where existing data cannot meet the brief.

How Sonexis supports AI teams

Speaker structure comes before format

One-speaker and two-speaker audio are Sonexis’s current speech focus. The required structure is matched to the model and deployment scenario rather than treated as interchangeable: a set that suits a single-speaker recogniser is often the wrong shape for a two-party voice agent, and the difference is structural, not acoustic.

One speaker — isolated prompts, intent phrases, read speech where the use case needs it, spontaneous monologue, and speaker-specific pronunciation or acoustic evaluation.

Two speakers — natural dialogue, customer and support interaction, interviewer and respondent flows, voice-agent evaluation, turn-taking, interruptions, corrections, code-switching, short answers and context changes.

For three-or-more-speaker requirements, Sonexis can assess existing supply, sourcing or buyer-funded collection feasibility where the deployment need and demand justify the additional participant, diarisation and QA complexity.

Voice agents

Voice agent evaluation, support and onboarding flows

Two-speaker. This is where a deployed system meets the behaviour no scripted set contains.

The problem

Agents perform well in demos and fail in live traffic: a caller barges in over the prompt, self-corrects mid-sentence, replies in three words, changes intent halfway, spells out an order number or an email address, switches language inside one sentence, or has to be recovered after the agent has already got it wrong — on a noisy line. Those behaviours are common in support, onboarding and KYC traffic, and rare in a scripted corpus.

A fitting dataset needs

Two-speaker scenario conversations that actually contain barge-in, interruption, self-correction, short replies, intent changes, entity spelling, failed recovery, code-switching, turn-taking and realistic call noise — with the turn structure and speaker attribution needed to score the agent’s side separately from the caller’s.

ASR

ASR training and evaluation

One-speaker and two-speaker are different briefs here, and mixing them hides which one the model fails.

The problem

ASR systems fail on accents, noise, code switching and informal speech. Data that is too clean overestimates real-world performance, so the number that ships is not the number that holds. A model that scores well on isolated utterances can still degrade sharply on two-party audio, because turn boundaries, overlap and channel structure are a separate failure surface.

A fitting dataset needs

One-speaker recognition and evaluation: isolated utterances and monologue across your accent range, with language tags, transcripts to the required convention, and stated recording conditions.

Two-speaker conversational ASR: the same, plus speaker attribution, turn structure, whether overlap occurs and is marked, and the channel structure of the delivered files — mixed or separated — stated rather than assumed.

Conversational AI

Conversational AI, sales and product discovery

The problem

Conversational systems have to survive messy dialogue, ambiguity, incomplete information and changing context. Sales and product-discovery traffic is the hardest form of it: users compare, explore, ask unclear questions, shift intent and challenge the answer. These conditions are difficult to simulate synthetically and rarely appear in standard sets.

A fitting dataset needs

Multi-turn conversations built around real user behaviour, including topic shifts, partial information and correction loops, across the buyer personas and product-comparison scenarios your system actually meets.

TTS

TTS and voice synthesis

Usually one-speaker, and the requirements are tighter than for recognition.

The problem

Conversational speech collected for recognition is not automatically suitable for synthesis. The properties that make two-party audio realistic — overlap, variable distance, background noise, mixed channels, many voices — are the properties that make it unusable for building a voice.

A fitting dataset needs

Tighter speaker identity, consent that explicitly covers voice synthesis, consistent recording conditions across sessions, clean signal quality, deliberate phonetic coverage, and controlled style or delivery. Suitability for synthesis is assessed separately from suitability for recognition, never inherited from it.

Evaluation

Targeted evaluation and multilingual benchmarks

Built around the failures you have actually seen, in the structure they occur in.

The problem

A general benchmark tells you a model is good on average. It does not tell you why it drops a caller who switches to Hindi mid-sentence, or mis-hears a spelled-out reference number on a noisy line. Standard sets routinely miss code switching and local speech behaviour — the things that decide real-world quality — so the failure that costs you traffic is the one the score does not measure.

A fitting dataset needs

Evaluation sets scoped around named deployment failures rather than around coverage: the accents, code-switch points, environments and speaker structures where your system already breaks, sized for repeatable measurement and held stable so a later run is comparable to an earlier one.

Deliveries can include structured metadata, QA records, consent references and delivery in agreed formats, depending on the agreed scope and on which route meets the requirement. Scope is confirmed with the buyer before any work begins.

Discuss your use case.

Tell us what you are building and what data would make the most impact. We check existing supply first, then managed sourcing, and scope funded collection only where nothing suitable exists.

Submit a requirement