Dataset Passport — illustrative specimen

Sonexis multilingual conversational sample set

This is an illustrative capability specimen. It shows the structure and evidence standard of a Sonexis Dataset Passport, using material recorded internally as Sonexis-recorded and Sonexis-controlled. Those origin, ownership and control statements are internal records and were not independently established for this release.

— It shows the evidence standard, not a catalogue entry. Availability is confirmed against each buyer requirement.

— It is an illustrative capability specimen. Not offered as licensable production supply. This is demonstration-scale material.

— It is limited to the evidence actually reviewed. Fields that were not established are shown as not established rather than omitted.

Origin and collection context

Origin categoryRecorded internally as Sonexis Direct Collection. Origin and ownership are internal records, not independently verified statements.
Collection routeRecorded in internal documentation as direct Sonexis collection. Not independently re-verified for this website release.
Collection purposeDemonstration and controlled sample evaluation.
Set contentsSix samples across three languages — Hindi, Indian English and Hinglish — each in a conversational and a single-speaker form.
Total duration934.64 seconds (15 minutes 35 seconds) across the six samples.
Speech typeNaturally spontaneous conversation. Not scripted, no fixed wording, not read and not acted. Interruptions, repetitions, fillers, self-corrections and code-switching are preserved rather than cleaned away.
Single-speaker formSelected-speaker extracts from the same spontaneous conversations. These demonstrate single-speaker packaging and are not naturally recorded monologues.

Evidence classes differ across this table and are stated separately below. File integrity and technical format were independently re-verified against the delivered files on 2026-08-09. Origin, collector, ownership and licensing authority were not independently re-verified: they are recorded in internal Sonexis documentation. Inspecting a delivered file can establish what the file is. It cannot establish who recorded it or who holds rights in it.

Technical specification

Measured directly from the delivered files, not transcribed from a supplier claim.

Container / codecWAV, PCM signed 16-bit little-endian (pcm_s16le)
Sample rate16 000 Hz
Bit depth16-bit
Channels1 (mono)
Per-sample duration184.805 s · 119.300 s · 181.000 s · 119.690 s · 210.245 s · 119.600 s
IntegritySHA-256 recorded per file. Source and output hashes match for all six files.
Independent verificationAll six files re-hashed and re-probed on 2026-08-09: 6 of 6 checksums matched the manifest, and codec, sample rate, bit depth, channel count and duration matched the recorded values in every case.
ProcessingAudio was not altered, trimmed, resampled, denoised, normalised or re-encoded during packaging.

16 kHz, 16-bit, mono PCM WAV does not by itself establish telephony capture method, codec history, or suitability for a specific telephony model.

Rights and consent

OwnershipRecorded internally as Sonexis-owned direct collection. This is an internal assertion by the accountable Sonexis representative, not independently established for public licensing, and the underlying rights instrument was not reviewed for this website release.
Licensing controlRecorded internally as Sonexis. Not independently established for public licensing.
Consent statusInternally confirmed by Sonexis against retained records.
Uses contributors authorised Sonexis to makeCommercial AI data use; commercialisation and licensing of the data; technical and commercial evaluation; controlled sample sharing with prospective customers.
Contributor identitiesNot included in the package.
Consent documentsNot included in the package. Supporting records are maintained internally by Sonexis.
External legal reviewNot carried out and not claimed.
Licence offered hereNone. Commercial status: illustrative capability specimen, not offered as licensable production supply. This specimen is published for illustration. It conveys no evaluation, training, fine-tuning, redistribution or production right.

Underlying contributor consent documents are not included in this public specimen and were not independently re-reviewed for this website release. The permitted uses above describe what contributors authorised Sonexis to do. They are not a grant of rights to any reader of this page. Production or model-training use of any Sonexis material requires a separate written agreement.

Speakers and recording conditions

Speaker genderMale. All speakers.
Speaker ageApproximately 22–25 years. Exact individual ages are not recorded.
Speaker originRajasthan, India.
Broad regional labelNorth India, resting on confirmed speaker origin only.
Accent classificationNo independent phonetic or accent-classification study has been completed.
Speaker countTwo per conversational sample; the demographic description applies to both.
Conversation structureTwo-speaker for the conversational samples. The single-speaker form is a selected-speaker extract from the same two-speaker recording, not a natively single-speaker capture. No three-or-more-speaker material is present in this set.
Channel structureSingle mixed channel. Both speakers share one mono channel — measured, see the technical specification above. The reviewed files are not channel-separated.
Speaker attributionPresent in the transcript as manually assigned speaker labels with timestamps, reviewed against the audio. No automated diarisation result is claimed.
Overlap markingNot established. The transcript carries no overlap annotation.
EnvironmentQuiet indoor environment with low background noise.
Acoustic measurementNo measured ambient-noise level, signal-to-noise ratio or room acoustic measurement is available.
Recording deviceNot recorded.
Studio / soundproofingNo claim is made in either direction.
DomainGeneral conversational speech. No specific industry or operational domain is assigned.

Transcripts, QA and PII

Transcript formatTimestamped speaker transcript, plain text. Hindi in Devanagari with spoken English terms preserved; Indian English in Latin script; Hinglish in Devanagari plus Latin-script English code-switching.
Transcript reviewManually reviewed against the matching audio for controlled sample sharing. Recorded complete by the accountable owner on 2026-07-23.
Transcript qualificationThis does not constitute production acceptance or confirmation against any buyer specification.
Manual audio reviewEnd-to-end listening completed on all six files. Playability confirmed.
Language labelsManually confirmed.
Speaker structureManually reviewed. Single-speaker extracts were checked for residual opposite-speaker speech; none was identified in the final shared excerpts.
Manual PII listeningNo obvious personally identifiable information was heard. This is not an absolute or forensic guarantee.
Automated PII scanSonexis review dated 2026-08-09. All six transcripts scanned for phone numbers, email addresses, URLs, long digit strings, government-ID patterns, monetary amounts and dates: zero matches.
Automated scan limitationThe scan covers Latin-script and numeric patterns only. It would not detect a personal name written in Devanagari. It supplements manual listening and does not replace it.
QA basisThis is a current Sonexis review, dated as above. It is not collection-time QA evidence, and no collection-time QA record is claimed for this material.

Known limitations

Sixteen. Published with the specimen rather than disclosed on request.

01Short evaluation excerpts. Not confirmed production inventory.

02Demonstration scale. 15 minutes 35 seconds in total, across six samples and three languages.

03Not assessed against any buyer’s acceptance criteria.

04Exact individual speaker ages are not recorded.

05No independent phonetic or accent-classification study has been completed.

06Recording-device details are not recorded.

07No measured ambient-noise, signal-to-noise or room acoustic values are available.

08No studio or soundproofing claim is made in either direction.

09No specific industry or operational domain is assigned.

10Any broad topic guidance used before recording is not recorded in the metadata.

11Manual listening heard no obvious PII; this is not an absolute or forensic guarantee.

12All speakers are male and from one state. This set is not demographically representative.

1316 kHz mono does not establish telephony capture method or codec history.

14No production volume, delivery timeline or commercial term is committed.

15No external legal review of the rights position has been carried out.

16Underlying contributor consent documents are not included in this public specimen and were not independently re-reviewed for this website release.

Verification scope

The distinction between what was confirmed, what was independently verified, and what was never established.

OwnershipInternal record — not independently verified
ConsentInternally confirmed by Sonexis against retained records — not independently re-reviewed for this release
Direct collectionInternal record — not independently verified
Speaker demographicsInternal record — not independently verified
Speaker regionInternal record — not independently verified
Recording environmentInternal record — not independently verified
Audio technical propertiesIndependently re-measured 2026-08-09 — establishes format only
Audio integrityIndependently re-verified 2026-08-09, 6/6 — establishes file integrity only
Transcript vs audioManually reviewed by accountable owner, 2026-07-23
Transcript PII (automated)Scanned by Sonexis 2026-08-09 — Latin/numeric patterns only
Accent classificationNot independently classified
Acoustic measurementsNot recorded
Recording device detailsNot recorded
Collection-time QA recordNot established
External legal reviewNot carried out

Delivery structure

sample_pack/
├── 00_README_FIRST.txt
├── 01_SAMPLE_MANIFEST.csv          one row per sample, with SHA-256
├── 02_KNOWN_LIMITATIONS.txt
├── 03_RIGHTS_AND_USAGE_NOTICE.txt
├── 01_Hindi/
│   ├── conversational/  audio/ transcript/ metadata/ qa_summary/
│   └── single_speaker/  audio/ transcript/ metadata/ qa_summary/
├── 02_Indian_English/  (same structure)
└── 03_Hinglish/        (same structure)

Every sample carries its own metadata file, QA summary and transcript. The manifest records the SHA-256 of each audio file so a recipient can verify integrity independently rather than trusting the packaging.

If this is the standard you need

Tell us the requirement. We check whether suitable existing data exists and can be licensed, then run managed sourcing, and scope funded collection only where existing data cannot meet the brief.