Speech & language data · North Africa

Speech and language data
for North African dialects.

Derja Data produces training datasets for Derja, Darija and Tamazight — recorded, transcribed and verified on the ground by paid native speakers. Production is live in Tunisia; Algeria, Morocco and Libya follow on request. Licensed for ASR, LLM and TTS development.

A contributor records a voice sample on a rooftop terrace in Tunis
4
countries on one platform
8
languages incl. four Tamazight varieties
24
mapped regions in Tunisia, GPS-verified
LIVE
production running in Tunisia
Data products

What you can license

Five product lines from one production pipeline — with a sample available before any commitment.

01

Speech corpora

Audio, human-verified transcription and speaker metadata, per country and dialect.

per verified hour
02

Derived datasets

Arabizi, transliteration pairs, annotation labels, code-switching — by-products of the same recording.

per corpus
03

Speech API

Speech-to-text tuned for Maghreb dialects, retrained continuously on newly verified corrections.

per minute
04

Annotation services

Our trained, trust-scored team labels your own data.

per task
05

Custom collection

You specify exactly the speech you need; we produce it on the ground.

per campaign

Every delivery includes audio (M4A; WAV on request), human-verified transcripts with JSON/CSV manifests, speaker and region metadata, consent documentation and watermarked files.

Full product catalogue
Two neighbours talking in an alley of the Tunis medina
EVERYDAY SPEECH — MEDINA, TUNIS
Production

How an hour of data is made

Our own smartphone application runs the entire pipeline — recording, machine draft, speaker correction and review.

01

Record

The app presents a prompt drawn from 6,347 templates; the contributor speaks freely in their own dialect.

02

Transcribe

Our own speech model drafts the transcript on our own servers. The data never leaves our infrastructure.

03

Correct

The same speaker corrects the draft — only they know what was said. The edit itself is the training signal.

04

Verify

Four independent quality layers determine whether the hour becomes a sellable product.

A contributor records a voice sample on her phone in a Tunis courtyard
FIELD RECORDING — TUNIS

Every correction also trains our internal draft model, so each subsequent hour is produced faster and more consistently.

Behind the data

Produced in our own app

whisper large-v3 → human
أريد أن أشرب قهوة
نحب نشرب قهوة

Every hour we deliver pairs audio with a transcript verified by the speaker themselves: the model drafts, the speaker corrects. The transcript records what was actually said.

Recording, machine draft, speaker correction, quality score and payout — the pipeline every delivered hour passes through.

Quality & compliance

Every hour is audited before delivery

Four independent control layers stand between a recording and a delivered hour.

01

Automatic anomaly detection

Automated integrity and plausibility checks run on every submission — before any person sees it.

02

Continuous quality measurement

Objective accuracy measurements the contributor can neither see nor predict feed a personal trust score.

03

Tiered human review

Newcomers are checked on every submission, experienced contributors by sample. Reviewers come from the top tier.

04

Independent consensus

Difficult cases are judged by several reviewers independently — only agreement lets them through.

A reviewer listens back to a recording and corrects the transcript
Included with every delivery
  • Documented consent from every speaker
  • Provenance trail per dataset and model
  • Watermarking on every delivered file
  • Clean-room separation: commercial models are trained exclusively on our own consented data
No public source

Data you cannot scrape

Speech models have almost no Maghreb training data — because there is no public source to collect it from.

Whisper training data (OpenAI, 2022), hours of speech
English438,000 h
98 other languages combined242,000 h
Maghreb dialects≈ 0 h

Own analysis based on the language table in the Whisper technical report (OpenAI, 2022).

There is no archive to buy and nothing to crawl. Derja and Darija are spoken languages that borrow from French, Italian and Tamazight, and Standard Arabic corpora do not cover them. The only way to obtain usable hours is to produce them — with native speakers, on the ground.

That is precisely what we do: paid contributors in four countries, with audit-ready consent and provenance documentation for every hour.

30–50%
typical word error rate reported for off-the-shelf models on Maghreb Arabic
~7%
our development target for a model trained on this corpus — methodology available on request
Coverage

Where we collect

One platform across four countries — every recording tagged with GPS-verified region and dialect-zone metadata.

TUNIS
TUNIS
ALGIERS
ALGIERS
CASABLANCA
CASABLANCA
TRIPOLI
TRIPOLI
CountryArabic varietyTamazightRegionsStatus
TunisiaTunisian Derja (ar-TN)24 governorates · 6 dialect zonesIn production
AlgeriaAlgerian Darja (ar-DZ)Kabyle58 wilayas · 6 dialect zonesOn request
MoroccoMoroccan Darija (ar-MA)Tashelhit · Tarifit · Central Atlas12 regions · 5 dialect zonesOn request
LibyaLibyan Arabic (ar-LY)22 districts · 4 dialect zonesOn request

24 mapped regions in Tunisia, 21 dialect zones, 8 languages — one platform.

Coverage in detail
The supply chain

Produced by fairly paid people — and that is what keeps quality high.

1.6–2.6×
the legal hourly minimum wage, per country
8+
payout channels incl. mobile money and cash
100%
of recordings covered by documented consent
Use cases

What our data is used for

Wherever software has to understand how North Africa actually speaks.

Voice assistants

Voice assistants

Natural Derja interaction for apps, smart devices and IVR — in place of menu systems that go unused.

Call-centre AI

Call-centre AI

Transcription, quality monitoring and live analytics in the language agents actually speak.

Media & subtitling

Media & subtitling

Dialect subtitling and searchable archives for streaming and broadcast across MENA.

Content moderation

Content moderation

Platforms cannot moderate what they cannot transcribe. Dialect data closes that gap.

Language learning

Language learning

Courses for heritage speakers — a Maghrebi diaspora numbering in the millions across Europe.

Public services

Public services

Citizen services, emergency-call transcription and accessibility in the languages people actually use.

Tell us what you need to train.

Language, domain, volume — we reply with availability, a sample dataset and a quote. Custom collection typically begins within weeks.