Open data

Open Derja and Darija corpora — what they cover, and what they leave open

Freely available corpora for North African Arabic exist, and they are genuinely useful. We work on them ourselves: our annotation pool of 2,028,674 sentences is assembled from openly licensed public corpora. This page is about which questions they answer and which they don't — so you can decide which part of your problem needs paid production and which doesn't.

What the open landscape holds

For Derja, Darija and Tamazight there are community translation sets, collections scraped from social media, read-speech corpora from volunteer projects and research releases from labs working on Maghreb languages. Our own annotation pool draws on darijaBridge, Tunisian Derja collections, MSA–Darija pair sets and Tatoeba-derived Kabyle material.

Each of them arrives with its own licence and its own consent story, and those differ far more between corpora than the file formats do. Both are worth reading before you train — not afterwards, when a customer asks.

Six questions to ask of any corpus

These apply to our material as much as to anyone's. If a supplier cannot answer them in writing, the answer is no.

01Can you show where each utterance came from?

Is there a record of who produced it, when, and under what agreement? A collection assembled from public posts usually cannot answer this, because the people who wrote those posts were never asked.

Every recording carries a timestamped consent record from a registered, paid contributor. The consent trail ships with the delivery and is built to survive an audit.

02Do you know the speaker and the region?

Without regional and speaker metadata you cannot balance a training set, and you cannot report accuracy per dialect. A model that works in Tunis and fails in Gafsa looks identical to one that works everywhere.

Each record carries the region it was produced in, checked against the contributor's registered location, plus the speaker band. 24 mapped regions in Tunisia, 21 dialect zones across four countries.

03Is the transcription human or machine?

Some collections contain machine transcription, occasionally without saying so. Training on the output of a speech recogniser teaches your model the errors of the recogniser that produced it.

Every transcription is produced or corrected by a human. Where a machine draft existed, the correction itself is recorded — that difference is the training signal you are actually paying for.

04Is audio paired with text from the same person, in the same moment?

Text-only corpora cannot train speech recognition or synthesis. Text matched to audio afterwards carries alignment errors you will not find until evaluation.

Audio and transcription come from one session and one contributor. For transliteration, both scripts come from the same speaker and the same utterance — not two texts paired later.

05Can you get more of exactly what you need?

A published corpus is fixed. If it lacks your domain — call-centre complaints, medical intake, banking vocabulary, a specific region — that gap stays a gap, however much data there is.

Collection is scheduled against your specification: domains, regions, speaker distribution, volume. 6,347 recording prompts across 194 subtopics exist today, and new ones are written to a brief.

06Does the licence carry your intended use?

"Open" spans everything from public domain to research-only. Check whether commercial training is permitted, whether attribution is required, and whether the terms still hold once the model you trained is sold.

Commercial, evaluation and white-label terms, agreed per project, in writing, with a named licensor you can point an auditor at.

When open corpora are the better choice

We would rather you spend nothing than spend it in the wrong place. There are cases where paid production is not the answer:

  • Feasibility studies and first baselines, where any dialect data beats none and the point is to learn whether the problem is solvable at all.
  • Academic work that publishes its results and can cite the source — reproducibility is worth more there than provenance.
  • Tasks where volume matters more than precision, such as language identification or coarse topic sorting.
  • Anything without a budget. Paid production earns its cost once the result has to hold up in front of a customer, not before.
The short version

Open corpora tell you whether the problem is solvable. They rarely tell you whether your solution is defensible.

If a customer, an auditor or a regulator will ask where your training data came from, that question has to be answerable before you train, not after. That is the part we sell — the recordings are how it is delivered.

Tell us what you need to train.

Language, domain, volume — we reply with availability, a sample dataset and a quote. Custom collection typically begins within weeks.