License one numeral recording —
it does the work of six datasets.
A native speaker says a number in Bantu, then crosses to the English value — in one take. Your model can't count in Bantu; this is the structured proof, and the fix.
Hear one take do the work.
The highlight swaps at the stored boundary. Play the Bantu numeral and the English value separately — that's the segmentation a number-ASR model needs. { } machine record shows it's real structured data.
{
"iso": "tsn",
"n_value": null,
"kind": "bantu_english",
"switch_ms": 53556,
"duration_ms": 68140,
"segments": [
{
"label": "bantu",
"start_ms": 0,
"end_ms": 53556
},
{
"label": "english",
"start_ms": 53556,
"end_ms": 68140
}
],
"speaker": "Speaker BEM-B",
"consent": "granted",
"deidentified": true
}
What one numeral take is worth.
The exact Bantu words for a value — recorded, native-verified. Your model stops hallucinating Bantu numbers.
Bantu→English at the value, with the stored switch point — training + eval for how people actually say numbers.
The English half, native-accented — accent-robust number recognition for Bantu-accented speech.
Each Bantu numeral decomposes into its language's Full Syllable Inventory — the units an aligner needs.
Consented native pronunciations keyed to value — a clean basis for number TTS and voice.
The English value is isolated and machine-transcribable against a gold answer — a self-grading eval set.
The benchmark grades itself.
We isolated the English value and gave it to a frontier ASR (Google Cloud Speech-to-Text). It drops the marker, mangles the value, or normalizes it wrong. The gap is the work — gold answer key built in.
Part of the BantuNomics family.
One consented recording = six datasets — for numbers, tone, and clinical health alike. License the program, not a dataset.
Talk to us →