Guide
Transliteration
import { transliterate } from '@devix-labs/arabic-slugify';
transliterate('مُحَمَّد'); // 'muhammad'
transliterate('چگونه', { lang: 'fa' }); // 'chguneh'
What it can and cannot do
Arabic script writes consonants and long vowels. Short vowels — the ones that
turn م-ر-ح-ب into "marhaba" — are marks, and almost nobody writes them.
So:
- Vocalised text romanises well.
مَرْحَبًا→marhaban, because the vowels are in the text. - Unvocalised text gives the consonants.
مرحبا→mrhba. That is not a shortcoming of this library; it is what the writing system provides.
Every other library returns the same kind of string and none of them explains why. If you need a readable Arabic slug, keep the script — see Getting started.
Where the romanisation earns its keep is as a stable, unique, ASCII-safe key: a filename, a database column, a legacy system that will not take Unicode.
The rules it does apply
The definite article. A leading ال becomes al-, which is most of what
makes an Arabic title recognisable in Latin letters.
transliterate('المقالة'); // 'al-mqala'
transliterate('المقالة', { article: false }); // 'almqala'
transliterate('الله'); // 'allah'
الله is the one word kept as an exception, because al-llah reads as a
mistake. The table is deliberately that small: a romanisation full of exceptions
is a dictionary pretending to be an algorithm.
Ta marbuta. ة closing a word is a; inside one it is t.
transliterate('مدرسة'); // 'mdrsa'
Three other libraries give mdrs, mdrsh and mdrst for that word.
Shadda doubles the consonant. After Unicode normalisation the vowel is
ordered before the shadda, which is how implementations end up doubling the
vowel instead — muhamaad rather than muhammad.
Tanween is followed by a silent alef, so مَرْحَبًا is marhaban, not
marhabana.
Where the languages differ
| Arabic | Persian | Urdu | |
|---|---|---|---|
و |
w |
v at the start of a word, u inside |
o |
ی / ي |
y |
y at the start, i inside |
i |
ث |
th |
s |
s |
ذ ض ظ |
dh d z |
z |
z |
ق |
q |
gh |
q |
final ه |
h |
eh |
eh |
ع |
a |
dropped | dropped |
پ چ ژ گ |
— | p ch zh g |
p ch zh g |
ٹ ڈ ڑ ں |
— | — | t d r n |
A table built only for Arabic drops پ, چ, ژ and گ entirely, which is why
پاکستان comes back as akstan from some libraries.
Positional readings
Persian and Urdu treat و and ی as consonants when they open a word and as
long vowels inside one:
transliterate('وزير', { lang: 'fa' }); // 'vzir' — opens the word
transliterate('چگونه', { lang: 'fa' }); // 'chguneh' — inside it
Arabic keeps w and y throughout, which is its own convention.