← Article directory

One God, One Alphabet, Four Languages: A Short Guide to the Middle East for the Confused European

19. 4. 2026
One God, One Alphabet, Four Languages: A Short Guide to the Middle East for the Confused European
Image from the original article on Médium.cz

The article challenges the widespread European notion of a uniform "Arab-Islamic world," showing that script, religion, and language are three independent layers: the Middle East interweaves three mutually unrelated language families—Semitic (Arabic), Indo-European (Persian and other Iranian languages, distantly related to Czech), and Turkic—unified only by a borrowed Arabic alphabet. The author explains Arabic diglossia (standard MSA with no native speakers alongside mutually unintelligible dialects, likened to medieval Latin), the grammatical kinship of Persian to European languages, and the reach of the "Persosphere" from Tajikistan to Pakistan. The conclusion wryly notes that the true pan-regional lingua franca of today's Middle East is English.

The European view of the Middle East has one mercifully soothing simplification: the "Arab-Islamic world." In our heads, Morocco merges with Saudi Arabia, Iran with Egypt, and Pakistan with Turkey into a single colorful blob that we label "Muslim countries" and stop poking at. Yet once you find yourself in Dubai — the present-day crossroads of the entire region — you discover that no such bloc exists. An Arab and an Iranian understand each other no better than a Czech and a Hungarian do. An Egyptian and a Moroccan also have to work at it. And Persian, which sounds "Middle Eastern," is in fact an Indo-European language, a distant cousin of Czech — even though it is written in the same alphabet as Arabic.

To make sense of it, you only need to grasp one thing: script, religion, and language are three different layers that need not overlap. Europe knows this anyway — after all, German, Hungarian, and Turkish are written in the same Latin alphabet, yet they belong to three completely different language families. The same holds in the Islamic world, only with the Arabic alphabet instead of the Latin one.

Three maps, not one

In the Middle East and its surroundings, three entirely unrelated language families are interwoven.

The Semitic family includes Arabic, Hebrew, Amharic (Ethiopia), Aramaic, and Maltese. It is a subfamily of the even larger Afroasiatic family, which reaches all the way into North Africa. Semitic languages have a characteristic grammar built on roots of three consonants: from k-t-b, Arabic forms kataba (he wrote), kitāb (book), kātib (writer), maktab (office), and a dozen other words. A principle unusual for a European — as if Czech were to generate from the consonants "p-s" the words psal, píseň, pes, pasáž, pošta, and then expected you to recognize a family in them.

The Indo-European family stretches from Iceland to Bangladesh, and in the Middle East it is represented by Persian, Kurdish, Pashto, and Balochi. All four are Iranian languages (note: "Iranian" is a linguistic term, not a political one) — close cousins of Hindi and Urdu, and, within the wider family, distant relatives of Czech. The Iranian and Indo-Aryan languages form two sister branches of the Indo-Iranian subfamily. When a Persian says pedar (father), mādar (mother), barādar (brother), dokhtar (daughter), nām (name), or now (new), you hear in it the echo of words you know from Czech, Latin, and German. Persia and Bohemia share a deep kinship of which only the skeleton remains — but that skeleton is there.

The Turkic family includes Turkish, Azerbaijani, Turkmen, Uzbek, Kazakh, Kyrgyz, and other languages from the Balkans to Siberia. They are agglutinative — they build words from long chains of suffixes, much like Finnish or Hungarian. They are related to neither Arabic nor Persian; they came into the Arabic alphabet only through Islam, and in some cases stepped back out of it again.

On top of all this, as a fourth layer, hangs the Arabic alphabet — a calligraphic creation that Islam brought with it and that the most varied languages adapted. Just as the Latin script took into its ranks Slavic (Czech, Polish), Finno-Ugric (Hungarian, Finnish), Turkic (Turkish), and Austronesian (Malay) languages, the Arabic alphabet clothed Persian, Urdu, Pashto, Uzbek (until the late 1920s), and Ottoman Turkish. The alphabet is the suit, not the body.

The Arab world: a language with no native speakers

Now to Arabic itself. The twenty-two states of the Arab League stretch from the Atlantic (Morocco, Mauritania) to the Gulf of Oman, and are home to more than 300 million native speakers of all the Arabic varieties combined (Ethnologue and Babbel 2026 estimates go as high as 360 million). That is more than French, more than German — Arabic is among the five most widely spoken languages in the world.

The paradox is that the Arabic that is written and spoken on the news has no native speaker anywhere at home. It is a medieval Latin that never went out of fashion.

This state of affairs has a technical name: diglossia — two versions of the same language exist side by side, a "high" and a "low" one, and each serves a different purpose. The "high" version is Modern Standard Arabic (MSA), in Arabic al-fusha ("the eloquent"), a modernized version of classical Quranic Arabic. It serves for writing, news, official speeches, and school instruction. The "low" version is the local dialect — ammiya in Egypt, darija in Morocco, shami in the Levant — which is spoken at home but traditionally not much written. You do not learn MSA at home but at school, much as you learned standard Czech alongside your colloquial Czech or Brno slang (hantec).

A Central European analogy: imagine that medieval Latin had never receded, and that in 2026 the television news in Madrid, Lisbon, Bucharest, and Paris were broadcast in Latin. At home, the Spanish, Portuguese, Romanians, and French would speak their vernaculars, which arose from Vulgar Latin and drifted apart from one another, but in writing, at school, in the mosque — pardon, the church — and in statesmen's speeches everyone would use the same codified Latin. That is exactly the situation of the Arab world.

The dialects, meanwhile, have drifted apart as dramatically as the Romance languages: a Saudi or an Iraqi will not spontaneously understand Moroccan darija, whereas a Moroccan understands Egyptian fairly well — but asymmetrically, because he spends his entire childhood watching Egyptian films and Syrian-Lebanese series, while a Cairene has no reason to learn darija. Thanks to the Cairo film industry of the 20th century, Egyptian Arabic became the informal lingua franca of the Arab world — when educated Arabs from different corners meet, they usually "smooth" their dialect toward Egyptian. MSA as such is rarely used in speech — it comes across as formal, bookish, as if you spoke Czech in a pub in the register of the Lord's Prayer.

Persian: the Middle Eastern cousin of Czech

Now we throw the switch and cross over to Iran. Here, alongside Arab Egypt and Turkish Istanbul, the Middle East becomes something you know from the European linguistic landscape as "Indo-European." Persian (Iranians themselves call it Farsi) is nothing but an older cousin of Hindi and, around a few corners, of Czech as well. It is joined to Czech by a proto-language four to six thousand years old — the same one that stands behind Latin and Sanskrit.

Look at the basic kinship terms: pedar (father), mādar (mother), barādar (brother), dokhtar (daughter) — the same root as Latin pater, mater, frater, English father, mother, brother, daughter, Czech mateř, bratr, dcera. The numerals yek, do, se, chahār (1, 2, 3, 4) correspond directly to Hindi ek, do, tīn, chār and, in the distance, to our jeden, dva, tři, čtyři. The word now means "new" — the same as English new, Latin novus, Czech nový. A Persian does not demand of you the inheritance of Semitic triconsonantal roots: the grammar is Indo-European, with prefixes and suffixes, without grammatical gender (there is no distinction between "he" and "she"), with fairly regular conjugation.

And that is precisely why Persian is grammatically far more accessible to a Czech than Arabic, even though it looks just as exotic, because you see it in the Arabic script. It is a trick similar to Yiddish: a Germanic language, structurally close to German, but written in the Hebrew script — at first glance it looks like Hebrew, but beneath the surface it is Germanic.

Script as a false friend

Let us return to that Dubai café. If an Arab and an Iranian saw a foreign text, both would recognize the letters — but they would read it as two different languages, each of which understands nothing of the other. Why?

Persian adapted the Arabic alphabet because, after the Islamization of Iran in the 7th–10th centuries, the prestige of the Arabic script was so great that the older Persian scripts (Pahlavi) were gradually abandoned. In doing so it added four of its own characters for sounds that classical Arabic does not have: a character for "p" (pe), for "ch" (che), for "zh" (zhe), and for "g" (gaf). Arabic lacks these four sounds — which is why, when Arabs borrow a Persian or European word, they mangle it for a long time (tilifūn instead of telefon, Bākistān instead of Pakistan). Persians, on the other hand, do not complain in the Arabic alphabet about missing letters, but about the opposite problem: Arabic has three different characters for the s sound (in Arabic sin, sad, tha) and even four for z (zay, dad, za, dhal) — and Persian pronounces them all the same.

What is more, Persian took on an enormous layer of Arabic loanwords — estimates range from roughly 30% of everyday speech to more than half the vocabulary of classical literature. So a Persian recognizes hundreds of words in an Arabic text as "foreign familiars" — much as a Czech recognizes telefon, doktor, systém, profesor in an English text. But when an Arab writes with Arabic grammar and a Persian with Persian grammar, the result is two incompatible sentences with plenty of shared words. Without study, the one does not understand the other.

The script, then, is not the language. A Frenchman reads the German text "Ich bin hungrig" in disbelief: he knows the letters, but the meaning passes him by entirely. The same holds for the relationship between Arabic and Persian.

The Persosphere: a surprisingly wide neighborhood

Because of the way a Czech looks at the map of the Middle East, he often has the feeling that Iran is a lonely island. In reality, Persian has an extensive "neighborhood" called the Persosphere.

Dari in Afghanistan and Tajik in Tajikistan are essentially three standards of the same language. The difference between Iranian Persian, Afghan Dari, and Tajik Tajik is about like the one between British and American English — a local accent, slightly different vocabulary, but fully mutually intelligible. The only catch: Tajik is written in Cyrillic, because Tajikistan was part of the Soviet Union. Moscow rewrote its alphabet — first into the Latin script in the late 1920s, then definitively into Cyrillic in 1939. Imagine if Slovak were written in the Cyrillic alphabet — the spoken form the same, the written form foreign. Together, there are around 110 million Persian speakers — more than the population of Germany.

Urdu in Pakistan is an Indo-Aryan language (a sister of Hindi), but it has an enormous layer of Persian and Arabic loanwords and is written in a Persian variant of the Arabic alphabet. Why? Because the Mughal Empire (1526–1857) adopted Persian as its administrative and literary language, and for several centuries a Persian-literate intellectual in Delhi was as much a matter of course as a Latin-literate intellectual in Prague. Urdu is essentially Hindi in Persian clothing. The parallel with Yiddish is striking: Yiddish is a Germanic language in the Hebrew alphabet with Hebrew-Aramaic loanwords; Urdu is an Indo-Aryan language in the Persian alphabet with Perso-Arabic loanwords. Both arose at cultural crossroads where a prestigious religious language cloaked the domestic vernacular.

Turkish is a Turkic language, grammatically completely different from both Arabic and Persian. Until Atatürk's reforms (1928), however, it was written in a Persian variant of the Arabic alphabet, and throughout the Ottoman era its vocabulary was largely Perso-Arabic — in literary and official texts the share of Persian and Arabic loanwords reached as much as 88 percent. Atatürk's alphabet reform and the subsequent language purification did lower the proportion, but modern Turkish still has thousands of Persian words: düşman (enemy, from Persian doshman), hafta (week), kitap (book), bahçe (garden), şehir (city), dost (friend, from Persian dust). When a Persian and a Turk meet, they cannot converse fluently — but they will find hundreds of shared words that both understand.

And a Central European analogy to close

If a Central European wanted to condense this whole map into a single image, it would look like this: the Arab world is like a great Latin-speaking empire in which every province eventually developed its own Romance language, yet everyone still writes and watches television in Latin. Persian is their Indo-European neighbor, which borrowed their alphabet and half their vocabulary but, beneath the surface, is closer to Czech than to Arabic. Turkish is a Turkic visitor that for centuries dressed in Persian clothes, then cast them off and today uses the Latin script. Urdu is the heir of this cultural blend at the opposite end of the Islamic world — an Indian language in Persian garb.

And the Pakistani, the Arab, the Iranian, and the Turk in that Dubai café? They are chatting in English. Because the only truly pan-regional language of the present-day Middle East — as Latin was in the Middle Ages, or French in the 19th century — today comes from elsewhere. From London and Washington. A historical irony worthy of a good sit-down over coffee.

The concept, structure, and editorial line of this article are the work of the author, who prepared the content outline, established the key theses, and directed the entire creative process. Generative AI (Claude, Anthropic) was used as a tool for research, fact-checking, and fleshing out the author's draft.

The author edited the outputs throughout, verified the key findings, and approved the final wording. No part of the text was published without human review. All factual data were verified against the publicly available sources cited in the text.

The procedure complies with the transparency requirements for AI-generated content under Art. 50 of EU Regulation 2024/1689 (AI Act). #poweredByAI

Read the Czech original on Médium.cz.

AI · Claude — machine translation, may contain inaccuracies.