Dataset

Maghreb Arabic Annotation Dataset

Label layers for North African Arabic: sentiment, dialect class, intent, topic and code-switching. Each label is set by several trained annotators independently and only counts as settled at 66% agreement or above.

Datasheet
Label types
Sentiment, dialect class, intent, topic, code-switching
Consensus rule
Several independent annotators, settled at ≥66% agreement
Languages
ar-TN, ar-DZ, ar-MA, Kabyle and Central Atlas Tamazight
Domain structure
194 subtopics across 17 domains, from food to youth slang
Delivery
JSON, JSONL or CSV, one row per sentence and label type
Your own text
Annotation as a service on your corpus — priced per task

What one delivered record looks like

These are the actual columns of our export view, not a proposed schema. You know before the first email whether the fields fit your training pipeline.

Fields per record
  • sentence_id
  • text_ar
  • annotation_type
  • consensus_label
  • consensus_score
  • times_annotated
  • domain_code
  • country_code
  • language_code
  • region_origin
Example record
{
  "sentence_id": "ann_41d9b2",
  "text_ar": "ما نجمش نصبر أكثر من هكة",
  "annotation_type": "sentiment",
  "consensus_label": "negative",
  "consensus_score": 1,
  "times_annotated": 3,
  "domain_code": "emotions_express.frustration",
  "country_code": "TN",
  "language_code": "ar-TN",
  "region_origin": "Tunis"
}

Structure is real, values are illustrative. Full sample package with audio and manifest on request, free of charge.

Where the data comes from

Our annotators work from a pool of 2,028,674 Derja and Darija sentences assembled from openly licensed public corpora. The source text stays under its original licence — what we license to you is the label layer our team produces on top of it. We equally annotate text you supply yourself.

How this compares to freely available corpora
Documented consent

Every contributor consents per recording, with a timestamped record. The consent trail ships with the delivery and stands up to an audit.

Paid native speakers

No scraping, no crowdsourced leftovers. Contributors earn 1.6–2.6× the legal hourly minimum wage of their country for verified work.

Region on record

Each record carries the region it was produced in, checked against the contributor's registered location — dialect zone is a fact, not an assumption.

Watermarked delivery

Every export is watermarked and download-tracked, so a leaked corpus can be traced back to the licence it left under.

Availability

First labelled batches from Q3 2026 — the pool is prepared and ready to label.

Tell us language, domains, regions and volume. You get a sample package and a quote — collection is scheduled against your specification.