Data Products
Enterprise-grade datasets spanning speech, text, annotation and transliteration for North African language AI
First corpus deliveries begin in Q3 2026 — production is already running in Tunisia.
Early-access customers receive priority allocation and a voice in domain coverage. Collection is order-driven: name the domains, regions and volume, and we schedule production against them.

Tunisian Arabic Speech Corpora
Audio recordings with manually verified transcriptions — in production in Tunisia today; Algeria, Morocco and Libya are served on request through custom collection. Each recording pairs native-speaker audio with a human-corrected transcription, speaker metadata and regional dialect tags.

Annotation Datasets
Our annotators work from a pool of 2,028,674 Derja and Darija sentences assembled from openly licensed public corpora. What we license to you is the label layer produced on top of it — sentiment, dialect class, intent, topic and code-switching — established through multi-annotator consensus with a minimum agreement threshold of 66%. The source text stays under its original licence; the annotations are ours to license. Annotation of your own text is equally available.

Derja ASR API
A speech-to-text API built specifically for North African Arabic dialects. Models are self-hosted and retrained continuously on newly verified speaker corrections.
curl -X POST https://api.derjadata.com/v1/transcribe \ -H 'Authorization: Bearer djk_your_key' \ -F 'audio=@recording.m4a' \ -F 'language=ar-TN'

Transliteration Pairs
A dual-script dataset mapping Arabic script to Arabizi, the Latin-and-numeral writing system in which 3=ع, 5=خ, 7=ح and 9=ق. Example: "قهوة بالحليب" ↔ "9ahwa bel 7lib". The foundation for building transliteration models.
What a delivery looks like
One record in the actual export format — every field shown here ships with every hour.
{
"id": "a3f1c2e8-…",
"audio": "audio/…/2026-08/….m4a", // 192 kbit/s mono M4A
"duration_seconds": 47,
"language_code": "ar-TN",
"dialect_zone": "sahel",
"region": "Sousse", // GPS-verified
"transcription_corrected": "نحب نشري كرهبة مستعملة",
"transcription_arabizi": "n7eb nechri karhba mesta3mla",
"script_mode": "both",
"speaker": { "gender": "f", "age_band": "25-34", "consent_ref": "c-2026-…" },
"qa": { "status": "approved", "reviewer_verdict": "approved", "trust_score": 96 },
"provenance": { "correction_keystrokes": 41, "edit_distance": 6, "watermark": "wm-…" }
}A sample record with illustrative values. The full sample package (audio and manifest) is free on request — audio is never published openly, as our consent commitments require.
Additional Data Products
LLM Fine-Tuning Data
Instruction-tuned examples for Arabic dialect language models
TTS Training Data
Audio and text pairs prepared for text-to-speech synthesis
Dialect Identification
Multi-country classification with ground-truth labels
Code-Switching Data
Arabic/French/Italian mix patterns, pre-labelled
Named Entity Recognition
Names, places and organisations in Derja contexts
Emotion Recognition
Audio emotion labels building on sentiment annotations
RLHF Preference Data
Chosen/rejected pairs for alignment training
Machine Translation Pairs
Derja↔MSA, Derja↔French parallel corpora
Annotation Services
Our trained, trust-scored team labels your own data — priced per task
Delivery & Formats
Tell us what you need to train.
Language, domain, volume — we reply with availability, a sample dataset and a quote. Custom collection typically begins within weeks.
