Dataset

Tunisian Arabic (Derja) Speech Corpus

Audio recordings by paid native speakers of Tunisian Derja, each with a transcription in Arabic script that a human produced or corrected. Built for training and evaluating speech recognition on a dialect that published models get wrong 30–50% of the time.

Datasheet
Audio format
M4A, AAC 192 kbit/s, 44.1 kHz, mono. WAV on request.
Transcription
UTF-8, Arabic script, human-produced or human-corrected
Languages
ar-TN live; ar-DZ, ar-LY, ar-MA and four Tamazight varieties on order
Regional coverage
24 mapped regions in Tunisia, 21 dialect zones across four countries
Delivery
JSON, JSONL or CSV manifest plus audio files, watermarked
Licence
Commercial, evaluation and white-label terms — agreed per project

What one delivered record looks like

These are the actual columns of our export view, not a proposed schema. You know before the first email whether the fields fit your training pipeline.

Fields per record
  • id
  • audio_url
  • text
  • duration_seconds
  • quality_score
  • country_code
  • language_code
  • submitted_at
  • region
  • gender
  • age
  • tier
Example record
{
  "id": "sub_8f3a1c07",
  "audio_url": "audio/tn/2026/08/rec_8f3a1c07.m4a",
  "text": "اليوم الحال باهي برشة، نمشيو نقعدو في القهوة",
  "duration_seconds": 6.4,
  "quality_score": 0.94,
  "country_code": "TN",
  "language_code": "ar-TN",
  "submitted_at": "2026-08-19T09:41:22Z",
  "region": "Sfax",
  "gender": "f",
  "age": "25-34",
  "tier": 3
}

Structure is real, values are illustrative. Full sample package with audio and manifest on request, free of charge.

Where the data comes from

Every recording is made in our own app by a registered, paid contributor — never scraped, never bought in. The correction step is what makes the material worth training on: a machine draft is not a label, a human correction is.

How this compares to freely available corpora
Documented consent

Every contributor consents per recording, with a timestamped record. The consent trail ships with the delivery and stands up to an audit.

Paid native speakers

No scraping, no crowdsourced leftovers. Contributors earn 1.6–2.6× the legal hourly minimum wage of their country for verified work.

Region on record

Each record carries the region it was produced in, checked against the contributor's registered location — dialect zone is a fact, not an assumption.

Watermarked delivery

Every export is watermarked and download-tracked, so a leaked corpus can be traced back to the licence it left under.

Availability

First deliveries from Q3 2026 — production is running in Tunisia.

Tell us language, domains, regions and volume. You get a sample package and a quote — collection is scheduled against your specification.