Tunisian Arabic (Derja) Speech Corpus
Audio recordings by paid native speakers of Tunisian Derja, each with a transcription in Arabic script that a human produced or corrected. Built for training and evaluating speech recognition on a dialect that published models get wrong 30–50% of the time.
- Audio format
- M4A, AAC 192 kbit/s, 44.1 kHz, mono. WAV on request.
- Transcription
- UTF-8, Arabic script, human-produced or human-corrected
- Languages
- ar-TN live; ar-DZ, ar-LY, ar-MA and four Tamazight varieties on order
- Regional coverage
- 24 mapped regions in Tunisia, 21 dialect zones across four countries
- Delivery
- JSON, JSONL or CSV manifest plus audio files, watermarked
- Licence
- Commercial, evaluation and white-label terms — agreed per project
What one delivered record looks like
These are the actual columns of our export view, not a proposed schema. You know before the first email whether the fields fit your training pipeline.
- id
- audio_url
- text
- duration_seconds
- quality_score
- country_code
- language_code
- submitted_at
- region
- gender
- age
- tier
{
"id": "sub_8f3a1c07",
"audio_url": "audio/tn/2026/08/rec_8f3a1c07.m4a",
"text": "اليوم الحال باهي برشة، نمشيو نقعدو في القهوة",
"duration_seconds": 6.4,
"quality_score": 0.94,
"country_code": "TN",
"language_code": "ar-TN",
"submitted_at": "2026-08-19T09:41:22Z",
"region": "Sfax",
"gender": "f",
"age": "25-34",
"tier": 3
}Structure is real, values are illustrative. Full sample package with audio and manifest on request, free of charge.
Where the data comes from
Every recording is made in our own app by a registered, paid contributor — never scraped, never bought in. The correction step is what makes the material worth training on: a machine draft is not a label, a human correction is.
How this compares to freely available corporaEvery contributor consents per recording, with a timestamped record. The consent trail ships with the delivery and stands up to an audit.
No scraping, no crowdsourced leftovers. Contributors earn 1.6–2.6× the legal hourly minimum wage of their country for verified work.
Each record carries the region it was produced in, checked against the contributor's registered location — dialect zone is a fact, not an assumption.
Every export is watermarked and download-tracked, so a leaked corpus can be traced back to the licence it left under.
First deliveries from Q3 2026 — production is running in Tunisia.
Tell us language, domains, regions and volume. You get a sample package and a quote — collection is scheduled against your specification.