Arabic–Arabizi Transliteration Pairs
The same utterance in two scripts: Arabic, and Arabizi — the Latin-and-numeral writing that North Africans actually type, where 3=ع, 5=خ, 7=ح and 9=ق. Training material for transliteration models and for any system that has to read what people write in chat.
- Pairing
- Same speaker, same utterance, both scripts in one session
- Scripts
- Arabic (UTF-8) and Arabizi (Latin letters plus 2, 3, 5, 7, 9)
- Languages
- ar-TN, ar-DZ, ar-LY, ar-MA
- Audio link
- Each pair can ship with the underlying recording
- Delivery
- JSON, JSONL or CSV, one row per pair
- Licence
- Commercial, evaluation and white-label terms — agreed per project
What one delivered record looks like
These are the actual columns of our export view, not a proposed schema. You know before the first email whether the fields fit your training pipeline.
- id
- arabic
- arabizi
- country_code
- language_code
- submitted_at
- region
- gender
- age
- tier
{
"id": "pair_2c77e4",
"arabic": "قهوة بالحليب من فضلك",
"arabizi": "9ahwa bel 7lib men fadhlek",
"country_code": "TN",
"language_code": "ar-TN",
"submitted_at": "2026-08-19T09:44:03Z",
"region": "Sfax",
"gender": "f",
"age": "25-34",
"tier": 3
}Structure is real, values are illustrative. Full sample package with audio and manifest on request, free of charge.
Where the data comes from
Both scripts come from the same speaker and the same utterance, produced in one session in our app. That is what makes the pair usable: it is not two texts matched afterwards, it is one person writing the same sentence twice.
How this compares to freely available corporaEvery contributor consents per recording, with a timestamped record. The consent trail ships with the delivery and stands up to an audit.
No scraping, no crowdsourced leftovers. Contributors earn 1.6–2.6× the legal hourly minimum wage of their country for verified work.
Each record carries the region it was produced in, checked against the contributor's registered location — dialect zone is a fact, not an assumption.
Every export is watermarked and download-tracked, so a leaked corpus can be traced back to the licence it left under.
Produced from Q3 2026 alongside the speech corpus.
Tell us language, domains, regions and volume. You get a sample package and a quote — collection is scheduled against your specification.