Maghreb Arabic Annotation Dataset
Label layers for North African Arabic: sentiment, dialect class, intent, topic and code-switching. Each label is set by several trained annotators independently and only counts as settled at 66% agreement or above.
- Label types
- Sentiment, dialect class, intent, topic, code-switching
- Consensus rule
- Several independent annotators, settled at ≥66% agreement
- Languages
- ar-TN, ar-DZ, ar-MA, Kabyle and Central Atlas Tamazight
- Domain structure
- 194 subtopics across 17 domains, from food to youth slang
- Delivery
- JSON, JSONL or CSV, one row per sentence and label type
- Your own text
- Annotation as a service on your corpus — priced per task
What one delivered record looks like
These are the actual columns of our export view, not a proposed schema. You know before the first email whether the fields fit your training pipeline.
- sentence_id
- text_ar
- annotation_type
- consensus_label
- consensus_score
- times_annotated
- domain_code
- country_code
- language_code
- region_origin
{
"sentence_id": "ann_41d9b2",
"text_ar": "ما نجمش نصبر أكثر من هكة",
"annotation_type": "sentiment",
"consensus_label": "negative",
"consensus_score": 1,
"times_annotated": 3,
"domain_code": "emotions_express.frustration",
"country_code": "TN",
"language_code": "ar-TN",
"region_origin": "Tunis"
}Structure is real, values are illustrative. Full sample package with audio and manifest on request, free of charge.
Where the data comes from
Our annotators work from a pool of 2,028,674 Derja and Darija sentences assembled from openly licensed public corpora. The source text stays under its original licence — what we license to you is the label layer our team produces on top of it. We equally annotate text you supply yourself.
How this compares to freely available corporaEvery contributor consents per recording, with a timestamped record. The consent trail ships with the delivery and stands up to an audit.
No scraping, no crowdsourced leftovers. Contributors earn 1.6–2.6× the legal hourly minimum wage of their country for verified work.
Each record carries the region it was produced in, checked against the contributor's registered location — dialect zone is a fact, not an assumption.
Every export is watermarked and download-tracked, so a leaked corpus can be traced back to the licence it left under.
First labelled batches from Q3 2026 — the pool is prepared and ready to label.
Tell us language, domains, regions and volume. You get a sample package and a quote — collection is scheduled against your specification.