About Derja Data

Bridging the AI data gap for 100 million North African speakers

Mission

Our Mission

We are building the foundational data infrastructure that allows AI to understand, speak and serve North Africa. Local communities produce and verify their own language data and are paid fairly for the work — a durable ecosystem in which technological progress and fair compensation reinforce each other.

The Problem We Solve

Modern AI systems — voice assistants, translation tools, speech analytics — perform admirably in English and even Modern Standard Arabic. They fail on the dialects actually spoken by 100 million people across Tunisia, Algeria, Libya and Morocco, for a simple reason: almost no structured training data exists for these languages. We are changing that.

Our Approach

We built a mobile platform on which native speakers across four countries earn fair wages recording, transcribing, annotating and reviewing speech data in their local dialects. A four-layer quality-control system holds the data to enterprise standards, and our ASR flywheel compounds: the more data we collect, the better our models become, and the less correction each hour requires.

Speakers of different generations in conversation in a courtyard in Tunis
SPEAKER COMMUNITY — TUNIS
Who is behind it

A founder-led operation

Before Derja Data, the founder spent eight years as managing director building and leading a pharmaceutical group with seven subsidiaries. Its business — importing, processing and distributing strictly controlled medicines — is governed by some of the most stringent provisions of German law and required the full range of pharmaceutical authorisations: wholesale distribution and controlled-substances permits, a manufacturing licence, an end-to-end GDP quality system and recurring inspections by the authorities; on the medicines it produced, the group was formally named as the manufacturer. At its peak the group employed more than 35 people and generated annual revenues in the millions of euros — without outside capital. Professionally, that career rests on commercial training in wholesale and foreign trade and business administration studies with a focus on business informatics.

That background matters here, because training data at production scale is an operations and compliance business: consent records, provenance trails and audit-ready documentation for every single record. It is the same discipline a licensed pharmaceutical operation runs on, applied to language data.

German-Tunisian and a native speaker of Derja, with family and a working network across Tunisia, the founder is building the production operation in Tunis in person. Derja Data remains fully self-financed to this day.

8
years running a licensed pharma group
7
subsidiaries in the group
35+
employees at peak
100%
self-financed — then and now

Our Values

Fair Compensation

Our contributors earn 1.6–2.6× the legal hourly minimum wage of their country, paid per verified unit of work.

Data Sovereignty

All data processing takes place on our own infrastructure; no audio ever passes through third-party servers. Contributors remain informed about how their data is used.

Quality Over Quantity

A multi-layer verification system — automated screening, continuous quality measurement and independent human consensus — ensures every data point meets production standards.

Tell us what you need to train.

Language, domain, volume — we reply with availability, a sample dataset and a quote. Custom collection typically begins within weeks.