Dataset Passport — illustrative specimen
Sonexis multilingual conversational sample set
This is an illustrative capability specimen. It shows the structure and evidence standard of a Sonexis Dataset Passport, using material recorded internally as Sonexis-recorded and Sonexis-controlled. Those origin, ownership and control statements are internal records and were not independently established for this release.
— It shows the evidence standard, not a catalogue entry. Availability is confirmed against each buyer requirement.
— It is an illustrative capability specimen. Not offered as licensable production supply. This is demonstration-scale material.
— It is limited to the evidence actually reviewed. Fields that were not established are shown as not established rather than omitted.
Origin and collection context
| Origin category | Recorded internally as Sonexis Direct Collection. Origin and ownership are internal records, not independently verified statements. |
| Collection route | Recorded in internal documentation as direct Sonexis collection. Not independently re-verified for this website release. |
| Collection purpose | Demonstration and controlled sample evaluation. |
| Set contents | Six samples across three languages — Hindi, Indian English and Hinglish — each in a conversational and a single-speaker form. |
| Total duration | 934.64 seconds (15 minutes 35 seconds) across the six samples. |
| Speech type | Naturally spontaneous conversation. Not scripted, no fixed wording, not read and not acted. Interruptions, repetitions, fillers, self-corrections and code-switching are preserved rather than cleaned away. |
| Single-speaker form | Selected-speaker extracts from the same spontaneous conversations. These demonstrate single-speaker packaging and are not naturally recorded monologues. |
Evidence classes differ across this table and are stated separately below. File integrity and technical format were independently re-verified against the delivered files on 2026-08-09. Origin, collector, ownership and licensing authority were not independently re-verified: they are recorded in internal Sonexis documentation. Inspecting a delivered file can establish what the file is. It cannot establish who recorded it or who holds rights in it.
Technical specification
Measured directly from the delivered files, not transcribed from a supplier claim.
| Container / codec | WAV, PCM signed 16-bit little-endian (pcm_s16le) |
| Sample rate | 16 000 Hz |
| Bit depth | 16-bit |
| Channels | 1 (mono) |
| Per-sample duration | 184.805 s · 119.300 s · 181.000 s · 119.690 s · 210.245 s · 119.600 s |
| Integrity | SHA-256 recorded per file. Source and output hashes match for all six files. |
| Independent verification | All six files re-hashed and re-probed on 2026-08-09: 6 of 6 checksums matched the manifest, and codec, sample rate, bit depth, channel count and duration matched the recorded values in every case. |
| Processing | Audio was not altered, trimmed, resampled, denoised, normalised or re-encoded during packaging. |
16 kHz, 16-bit, mono PCM WAV does not by itself establish telephony capture method, codec history, or suitability for a specific telephony model.
Rights and consent
| Ownership | Recorded internally as Sonexis-owned direct collection. This is an internal assertion by the accountable Sonexis representative, not independently established for public licensing, and the underlying rights instrument was not reviewed for this website release. |
| Licensing control | Recorded internally as Sonexis. Not independently established for public licensing. |
| Consent status | Internally confirmed by Sonexis against retained records. |
| Uses contributors authorised Sonexis to make | Commercial AI data use; commercialisation and licensing of the data; technical and commercial evaluation; controlled sample sharing with prospective customers. |
| Contributor identities | Not included in the package. |
| Consent documents | Not included in the package. Supporting records are maintained internally by Sonexis. |
| External legal review | Not carried out and not claimed. |
| Licence offered here | None. Commercial status: illustrative capability specimen, not offered as licensable production supply. This specimen is published for illustration. It conveys no evaluation, training, fine-tuning, redistribution or production right. |
Underlying contributor consent documents are not included in this public specimen and were not independently re-reviewed for this website release. The permitted uses above describe what contributors authorised Sonexis to do. They are not a grant of rights to any reader of this page. Production or model-training use of any Sonexis material requires a separate written agreement.
Speakers and recording conditions
| Speaker gender | Male. All speakers. |
| Speaker age | Approximately 22–25 years. Exact individual ages are not recorded. |
| Speaker origin | Rajasthan, India. |
| Broad regional label | North India, resting on confirmed speaker origin only. |
| Accent classification | No independent phonetic or accent-classification study has been completed. |
| Speaker count | Two per conversational sample; the demographic description applies to both. |
| Conversation structure | Two-speaker for the conversational samples. The single-speaker form is a selected-speaker extract from the same two-speaker recording, not a natively single-speaker capture. No three-or-more-speaker material is present in this set. |
| Channel structure | Single mixed channel. Both speakers share one mono channel — measured, see the technical specification above. The reviewed files are not channel-separated. |
| Speaker attribution | Present in the transcript as manually assigned speaker labels with timestamps, reviewed against the audio. No automated diarisation result is claimed. |
| Overlap marking | Not established. The transcript carries no overlap annotation. |
| Environment | Quiet indoor environment with low background noise. |
| Acoustic measurement | No measured ambient-noise level, signal-to-noise ratio or room acoustic measurement is available. |
| Recording device | Not recorded. |
| Studio / soundproofing | No claim is made in either direction. |
| Domain | General conversational speech. No specific industry or operational domain is assigned. |
Transcripts, QA and PII
| Transcript format | Timestamped speaker transcript, plain text. Hindi in Devanagari with spoken English terms preserved; Indian English in Latin script; Hinglish in Devanagari plus Latin-script English code-switching. |
| Transcript review | Manually reviewed against the matching audio for controlled sample sharing. Recorded complete by the accountable owner on 2026-07-23. |
| Transcript qualification | This does not constitute production acceptance or confirmation against any buyer specification. |
| Manual audio review | End-to-end listening completed on all six files. Playability confirmed. |
| Language labels | Manually confirmed. |
| Speaker structure | Manually reviewed. Single-speaker extracts were checked for residual opposite-speaker speech; none was identified in the final shared excerpts. |
| Manual PII listening | No obvious personally identifiable information was heard. This is not an absolute or forensic guarantee. |
| Automated PII scan | Sonexis review dated 2026-08-09. All six transcripts scanned for phone numbers, email addresses, URLs, long digit strings, government-ID patterns, monetary amounts and dates: zero matches. |
| Automated scan limitation | The scan covers Latin-script and numeric patterns only. It would not detect a personal name written in Devanagari. It supplements manual listening and does not replace it. |
| QA basis | This is a current Sonexis review, dated as above. It is not collection-time QA evidence, and no collection-time QA record is claimed for this material. |
Known limitations
Sixteen. Published with the specimen rather than disclosed on request.
01Short evaluation excerpts. Not confirmed production inventory.
02Demonstration scale. 15 minutes 35 seconds in total, across six samples and three languages.
03Not assessed against any buyer’s acceptance criteria.
04Exact individual speaker ages are not recorded.
05No independent phonetic or accent-classification study has been completed.
06Recording-device details are not recorded.
07No measured ambient-noise, signal-to-noise or room acoustic values are available.
08No studio or soundproofing claim is made in either direction.
09No specific industry or operational domain is assigned.
10Any broad topic guidance used before recording is not recorded in the metadata.
11Manual listening heard no obvious PII; this is not an absolute or forensic guarantee.
12All speakers are male and from one state. This set is not demographically representative.
1316 kHz mono does not establish telephony capture method or codec history.
14No production volume, delivery timeline or commercial term is committed.
15No external legal review of the rights position has been carried out.
16Underlying contributor consent documents are not included in this public specimen and were not independently re-reviewed for this website release.
Verification scope
The distinction between what was confirmed, what was independently verified, and what was never established.
| Ownership | Internal record — not independently verified |
| Consent | Internally confirmed by Sonexis against retained records — not independently re-reviewed for this release |
| Direct collection | Internal record — not independently verified |
| Speaker demographics | Internal record — not independently verified |
| Speaker region | Internal record — not independently verified |
| Recording environment | Internal record — not independently verified |
| Audio technical properties | Independently re-measured 2026-08-09 — establishes format only |
| Audio integrity | Independently re-verified 2026-08-09, 6/6 — establishes file integrity only |
| Transcript vs audio | Manually reviewed by accountable owner, 2026-07-23 |
| Transcript PII (automated) | Scanned by Sonexis 2026-08-09 — Latin/numeric patterns only |
| Accent classification | Not independently classified |
| Acoustic measurements | Not recorded |
| Recording device details | Not recorded |
| Collection-time QA record | Not established |
| External legal review | Not carried out |
Delivery structure
sample_pack/ ├── 00_README_FIRST.txt ├── 01_SAMPLE_MANIFEST.csv one row per sample, with SHA-256 ├── 02_KNOWN_LIMITATIONS.txt ├── 03_RIGHTS_AND_USAGE_NOTICE.txt ├── 01_Hindi/ │ ├── conversational/ audio/ transcript/ metadata/ qa_summary/ │ └── single_speaker/ audio/ transcript/ metadata/ qa_summary/ ├── 02_Indian_English/ (same structure) └── 03_Hinglish/ (same structure)
Every sample carries its own metadata file, QA summary and transcript. The manifest records the SHA-256 of each audio file so a recipient can verify integrity independently rather than trusting the packaging.
If this is the standard you need
Tell us the requirement. We check whether suitable existing data exists and can be licensed, then run managed sourcing, and scope funded collection only where existing data cannot meet the brief.