Speech and language data
for North African dialects.
Derja Data produces training datasets for Derja, Darija and Tamazight — recorded, transcribed and verified on the ground by paid native speakers. Production is live in Tunisia; Algeria, Morocco and Libya follow on request. Licensed for ASR, LLM and TTS development.

What you can license
Five product lines from one production pipeline — with a sample available before any commitment.
Speech corpora
Audio, human-verified transcription and speaker metadata, per country and dialect.
per verified hourDerived datasets
Arabizi, transliteration pairs, annotation labels, code-switching — by-products of the same recording.
per corpusSpeech API
Speech-to-text tuned for Maghreb dialects, retrained continuously on newly verified corrections.
per minuteAnnotation services
Our trained, trust-scored team labels your own data.
per taskCustom collection
You specify exactly the speech you need; we produce it on the ground.
per campaignEvery delivery includes audio (M4A; WAV on request), human-verified transcripts with JSON/CSV manifests, speaker and region metadata, consent documentation and watermarked files.
Full product catalogue →
How an hour of data is made
Our own smartphone application runs the entire pipeline — recording, machine draft, speaker correction and review.
Record
The app presents a prompt drawn from 6,347 templates; the contributor speaks freely in their own dialect.
Transcribe
Our own speech model drafts the transcript on our own servers. The data never leaves our infrastructure.
Correct
The same speaker corrects the draft — only they know what was said. The edit itself is the training signal.
Verify
Four independent quality layers determine whether the hour becomes a sellable product.

Every correction also trains our internal draft model, so each subsequent hour is produced faster and more consistently.
Produced in our own app
Every hour we deliver pairs audio with a transcript verified by the speaker themselves: the model drafts, the speaker corrects. The transcript records what was actually said.
Recording, machine draft, speaker correction, quality score and payout — the pipeline every delivered hour passes through.
Every hour is audited before delivery
Four independent control layers stand between a recording and a delivered hour.
Automatic anomaly detection
Automated integrity and plausibility checks run on every submission — before any person sees it.
Continuous quality measurement
Objective accuracy measurements the contributor can neither see nor predict feed a personal trust score.
Tiered human review
Newcomers are checked on every submission, experienced contributors by sample. Reviewers come from the top tier.
Independent consensus
Difficult cases are judged by several reviewers independently — only agreement lets them through.

- —Documented consent from every speaker
- —Provenance trail per dataset and model
- —Watermarking on every delivered file
- —Clean-room separation: commercial models are trained exclusively on our own consented data
Data you cannot scrape
Speech models have almost no Maghreb training data — because there is no public source to collect it from.
Own analysis based on the language table in the Whisper technical report (OpenAI, 2022).
There is no archive to buy and nothing to crawl. Derja and Darija are spoken languages that borrow from French, Italian and Tamazight, and Standard Arabic corpora do not cover them. The only way to obtain usable hours is to produce them — with native speakers, on the ground.
That is precisely what we do: paid contributors in four countries, with audit-ready consent and provenance documentation for every hour.
Where we collect
One platform across four countries — every recording tagged with GPS-verified region and dialect-zone metadata.




| Country | Arabic variety | Tamazight | Regions | Status |
|---|---|---|---|---|
| Tunisia | Tunisian Derja (ar-TN) | — | 24 governorates · 6 dialect zones | In production |
| Algeria | Algerian Darja (ar-DZ) | Kabyle | 58 wilayas · 6 dialect zones | On request |
| Morocco | Moroccan Darija (ar-MA) | Tashelhit · Tarifit · Central Atlas | 12 regions · 5 dialect zones | On request |
| Libya | Libyan Arabic (ar-LY) | — | 22 districts · 4 dialect zones | On request |
24 mapped regions in Tunisia, 21 dialect zones, 8 languages — one platform.
Coverage in detail →Produced by fairly paid people — and that is what keeps quality high.
What our data is used for
Wherever software has to understand how North Africa actually speaks.

Voice assistants
Natural Derja interaction for apps, smart devices and IVR — in place of menu systems that go unused.

Call-centre AI
Transcription, quality monitoring and live analytics in the language agents actually speak.

Media & subtitling
Dialect subtitling and searchable archives for streaming and broadcast across MENA.

Content moderation
Platforms cannot moderate what they cannot transcribe. Dialect data closes that gap.

Language learning
Courses for heritage speakers — a Maghrebi diaspora numbering in the millions across Europe.

Public services
Citizen services, emergency-call transcription and accessibility in the languages people actually use.
Tell us what you need to train.
Language, domain, volume — we reply with availability, a sample dataset and a quote. Custom collection typically begins within weeks.

