Skip to content
Devix Open Source

Guide

Transliteration

import { transliterate } from '@devix-labs/arabic-slugify';

transliterate('مُحَمَّد');                    // 'muhammad'
transliterate('چگونه', { lang: 'fa' });       // 'chguneh'

What it can and cannot do

Arabic script writes consonants and long vowels. Short vowels — the ones that turn م-ر-ح-ب into "marhaba" — are marks, and almost nobody writes them.

So:

  • Vocalised text romanises well. مَرْحَبًا → marhaban, because the vowels are in the text.
  • Unvocalised text gives the consonants. مرحبا → mrhba. That is not a shortcoming of this library; it is what the writing system provides.

Every other library returns the same kind of string and none of them explains why. If you need a readable Arabic slug, keep the script — see Getting started.

Where the romanisation earns its keep is as a stable, unique, ASCII-safe key: a filename, a database column, a legacy system that will not take Unicode.

The rules it does apply

The definite article. A leading ال becomes al-, which is most of what makes an Arabic title recognisable in Latin letters.

transliterate('المقالة');                      // 'al-mqala'
transliterate('المقالة', { article: false });  // 'almqala'
transliterate('الله');                          // 'allah'

الله is the one word kept as an exception, because al-llah reads as a mistake. The table is deliberately that small: a romanisation full of exceptions is a dictionary pretending to be an algorithm.

Ta marbuta. ة closing a word is a; inside one it is t.

transliterate('مدرسة');   // 'mdrsa'

Three other libraries give mdrs, mdrsh and mdrst for that word.

Shadda doubles the consonant. After Unicode normalisation the vowel is ordered before the shadda, which is how implementations end up doubling the vowel instead — muhamaad rather than muhammad.

Tanween is followed by a silent alef, so مَرْحَبًا is marhaban, not marhabana.

Where the languages differ

Arabic Persian Urdu
و w v at the start of a word, u inside o
ی / ي y y at the start, i inside i
ث th s s
ذ ض ظ dh d z z z
ق q gh q
final ه h eh eh
ع a dropped dropped
پ چ ژ گ — p ch zh g p ch zh g
ٹ ڈ ڑ ں — — t d r n

A table built only for Arabic drops پ, چ, ژ and گ entirely, which is why پاکستان comes back as akstan from some libraries.

Positional readings

Persian and Urdu treat و and ی as consonants when they open a word and as long vowels inside one:

transliterate('وزير', { lang: 'fa' });    // 'vzir'   — opens the word
transliterate('چگونه', { lang: 'fa' });   // 'chguneh' — inside it

Arabic keeps w and y throughout, which is its own convention.

Updated 15 Sep 2026