Dataset

Arabic–Arabizi Transliteration Pairs

The same utterance in two scripts: Arabic, and Arabizi — the Latin-and-numeral writing that North Africans actually type, where 3=ع, 5=خ, 7=ح and 9=ق. Training material for transliteration models and for any system that has to read what people write in chat.

Datasheet
Pairing
Same speaker, same utterance, both scripts in one session
Scripts
Arabic (UTF-8) and Arabizi (Latin letters plus 2, 3, 5, 7, 9)
Languages
ar-TN, ar-DZ, ar-LY, ar-MA
Audio link
Each pair can ship with the underlying recording
Delivery
JSON, JSONL or CSV, one row per pair
Licence
Commercial, evaluation and white-label terms — agreed per project

What one delivered record looks like

These are the actual columns of our export view, not a proposed schema. You know before the first email whether the fields fit your training pipeline.

Fields per record
  • id
  • arabic
  • arabizi
  • country_code
  • language_code
  • submitted_at
  • region
  • gender
  • age
  • tier
Example record
{
  "id": "pair_2c77e4",
  "arabic": "قهوة بالحليب من فضلك",
  "arabizi": "9ahwa bel 7lib men fadhlek",
  "country_code": "TN",
  "language_code": "ar-TN",
  "submitted_at": "2026-08-19T09:44:03Z",
  "region": "Sfax",
  "gender": "f",
  "age": "25-34",
  "tier": 3
}

Structure is real, values are illustrative. Full sample package with audio and manifest on request, free of charge.

Where the data comes from

Both scripts come from the same speaker and the same utterance, produced in one session in our app. That is what makes the pair usable: it is not two texts matched afterwards, it is one person writing the same sentence twice.

How this compares to freely available corpora
Documented consent

Every contributor consents per recording, with a timestamped record. The consent trail ships with the delivery and stands up to an audit.

Paid native speakers

No scraping, no crowdsourced leftovers. Contributors earn 1.6–2.6× the legal hourly minimum wage of their country for verified work.

Region on record

Each record carries the region it was produced in, checked against the contributor's registered location — dialect zone is a fact, not an assumption.

Watermarked delivery

Every export is watermarked and download-tracked, so a leaked corpus can be traced back to the licence it left under.

Availability

Produced from Q3 2026 alongside the speech corpus.

Tell us language, domains, regions and volume. You get a sample package and a quote — collection is scheduled against your specification.