Data Products

Enterprise-grade datasets spanning speech, text, annotation and transliteration for North African language AI

Availability

First corpus deliveries begin in Q3 2026 — production is already running in Tunisia.

Early-access customers receive priority allocation and a voice in domain coverage. Collection is order-driven: name the domains, regions and volume, and we schedule production against them.

6,347
recording prompts across 17 domains
194
subtopics with prepared prompts
24
GPS-verified regions in Tunisia
8
language varieties incl. four Tamazight
Tunisian Arabic Speech Corpora

Tunisian Arabic Speech Corpora

Audio recordings with manually verified transcriptions — in production in Tunisia today; Algeria, Morocco and Libya are served on request through custom collection. Each recording pairs native-speaker audio with a human-corrected transcription, speaker metadata and regional dialect tags.

M4A audio (192 kbit/s mono) + UTF-8 transcriptions
Delivered per country, dialect and domain — prepared for ASR training
Datasheet and schema
Annotation Datasets

Annotation Datasets

Our annotators work from a pool of 2,028,674 Derja and Darija sentences assembled from openly licensed public corpora. What we license to you is the label layer produced on top of it — sentiment, dialect class, intent, topic and code-switching — established through multi-annotator consensus with a minimum agreement threshold of 66%. The source text stays under its original licence; the annotations are ours to license. Annotation of your own text is equally available.

SentimentDialectIntentTopicCode-Switch
Datasheet and schema
Derja ASR API
Early access — launching with our first fine-tuned model

Derja ASR API

A speech-to-text API built specifically for North African Arabic dialects. Models are self-hosted and retrained continuously on newly verified speaker corrections.

curl -X POST https://api.derjadata.com/v1/transcribe \
  -H 'Authorization: Bearer djk_your_key' \
  -F 'audio=@recording.m4a' \
  -F 'language=ar-TN'
Transliteration Pairs

Transliteration Pairs

A dual-script dataset mapping Arabic script to Arabizi, the Latin-and-numeral writing system in which 3=ع, 5=خ, 7=ح and 9=ق. Example: "قهوة بالحليب" ↔ "9ahwa bel 7lib". The foundation for building transliteration models.

قهوة بالحليب9ahwa bel 7lib
Datasheet and schema

What a delivery looks like

One record in the actual export format — every field shown here ships with every hour.

{
  "id": "a3f1c2e8-…",
  "audio": "audio/…/2026-08/….m4a",          // 192 kbit/s mono M4A
  "duration_seconds": 47,
  "language_code": "ar-TN",
  "dialect_zone": "sahel",
  "region": "Sousse",                        // GPS-verified
  "transcription_corrected": "نحب نشري كرهبة مستعملة",
  "transcription_arabizi": "n7eb nechri karhba mesta3mla",
  "script_mode": "both",
  "speaker": { "gender": "f", "age_band": "25-34", "consent_ref": "c-2026-…" },
  "qa": { "status": "approved", "reviewer_verdict": "approved", "trust_score": 96 },
  "provenance": { "correction_keystrokes": 41, "edit_distance": 6, "watermark": "wm-…" }
}

A sample record with illustrative values. The full sample package (audio and manifest) is free on request — audio is never published openly, as our consent commitments require.

Additional Data Products

LLM Fine-Tuning Data

Instruction-tuned examples for Arabic dialect language models

TTS Training Data

Audio and text pairs prepared for text-to-speech synthesis

Dialect Identification

Multi-country classification with ground-truth labels

Code-Switching Data

Arabic/French/Italian mix patterns, pre-labelled

Named Entity Recognition

Names, places and organisations in Derja contexts

Emotion Recognition

Audio emotion labels building on sentiment annotations

RLHF Preference Data

Chosen/rejected pairs for alignment training

Machine Translation Pairs

Derja↔MSA, Derja↔French parallel corpora

Annotation Services

Our trained, trust-scored team labels your own data — priced per task

Delivery & Formats

CSV, JSON, or JSONL with comprehensive metadata
Watermarked exports with download tracking
Evaluation, commercial, and white-label licensing
Custom data collection campaigns available

Tell us what you need to train.

Language, domain, volume — we reply with availability, a sample dataset and a quote. Custom collection typically begins within weeks.