BantuNomics Sign in
B BantuNomics Numbers · BTS-S100
Private access Sign out
BantuNomics Numbers The Bantu numeral substrate for AI systems

Does our data improve your model? We measured it.

0.0% baseline · 0.0% scrambled-control · 75.8% fine-tuned — exact-match on unseen tasks.

We took models that score essentially zero on Bantu numeral tasks, fine-tuned them on BantuNomics-generated data (~$5 of compute for the headline run), and graded them on a frozen, fingerprinted held-out exam under all-or-nothing exact-match scoring. A control model trained on the same records with shuffled answers learned the style and still scored 0% — the entire gain is the content being correct.

Per-model results (all runs, nothing omitted)

ModelRoundBaselineControlFine-tunedNote
Qwen3.5-4BR2 (enriched data)0.0%0.0%75.8%numerals 95% · arithmetic 87% on unseen items
Qwen3.5-4BR13.6%56.8%arithmetic 24/24; data enrichment → R2
Gemma 4 E4B-itR12.9%18.7%model choice matters enormously
AfriqueQwen-8BR20.0%16.5%African-language pretraining ≠ Bantu numeracy; best calendar score of any model
What this teaches about improving your model:
  1. The data is the dial. Enriching coverage moved 57% → 76% on a harder exam; coverage holes become zeros.
  2. Content beats fluency. The control proves style-learning scores nothing; only verified forms move the metric.
  3. Model choice dominates. Same lessons, same budget: 18.7% vs 75.8%.
  4. Measure before and after. The loop below is free at evaluation tier.

Run the loop on your own model

1. GET /api/tools/benchmark?iso=<your language> # the eval set 2. POST /api/tools/benchmark/score {"iso": "...", "submissions": {"88": "your model's answer", ...}} 3. Train on the generated corpora (Full Subscription: /api/tools/corpus.jsonl|asr|chatml…) 4. POST score again → GET /api/tools/benchmark/runs shows your delta. MCP: calc_benchmark · calc_score · calc_benchmark_runs (headless).

The interactive dashboard (all 9 runs)

Full reports

Technical report (PDF) Plain-language report (PDF)

Every number, method, command and failure — we publish the incidents too, because for a serious buyer the honesty is the credibility.

Access

Start where you are.

Prove it free, validate it on your own data, or license the whole program.

01 · Free
Evaluation

Score your model against the foundational layer — the Alphabet Test and the L26 Lite suite, with saved results. Self-serve, no cost.

Start evaluating →
Most labs start here
02 · 75 days
Validation Pilot

Three languages you choose — and everything we hold for them. Measured on your own held-out data.

Scope a pilot →
03 · Program
Full Annual Subscription

Every product, every language, the full consented corpus — and everything curated while you're subscribed.

Start a Full Annual Subscription →