Building an Arabic chatbot that Gulf customers actually use starts with a corpus of real customer messages, not a model's default Arabic output. Here is how to build an Arabic chatbot that handles dialect, Arabizi and code switching without guessing: create an evaluation set, label what you are actually dealing with, decide the register per intent, fix right-to-left rendering in the widget, and review misses on a fixed weekly cadence. None of these steps is exotic engineering, but skipping one is usually why an Arabic bot reads like it was translated from an English script rather than written for the customer in front of it.
How to build an Arabic chatbot, step by step
The method has six parts, in order:
- Build an evaluation set from real customer messages.
- Label dialect, script and code switching in every example.
- Decide the answer register per intent.
- Get right-to-left and mixed-direction rendering correct in the widget.
- Route WhatsApp's specific limits correctly.
- Log misses every week and store transcripts under PDPL.
Each step is detailed below.
Start with an evaluation set built from real messages
Before you write a single prompt or intent, pull a broad, representative batch of real customer messages from your support inbox, WhatsApp thread or call transcripts turned into text. Do not write synthetic Arabic examples first. Customers write the way they type on a phone keyboard, which is different from how a translator or a language model writes when asked to "respond in Arabic."
Group the messages by intent (order status, return, complaint, booking, general question) and keep the raw text, not a cleaned-up version. This becomes your evaluation set: the fixed list you run every candidate model or prompt change against before it ships. Without it you are guessing whether a change made things better or worse, and this is standard practice in Arabic NLP customer support work, not a Gulf-specific trick.
Label dialect, script and mixing before you label intent
Tag every message in the eval set with two things: dialect and script. Skip this and your intent accuracy numbers hide a language problem underneath.
For dialect, use a simple five-way scheme that mirrors what Arabic NLP researchers use for classification: Modern Standard Arabic (MSA), Egyptian, Gulf, Levantine and North African ACL Anthology. Treat these as linguistic buckets, not passport categories: Gulf Arabic is generally described as covering the Arabian Peninsula, and Iraqi Arabic sits between Levantine and Gulf rather than fitting neatly into either arXiv. A Gulf Arabic chatbot built for UAE, Saudi, Kuwaiti or Omani customers will mostly see Gulf and MSA messages, with Egyptian and Levantine appearing from expatriate customers, delivery partners and staff.
For script, mark whether the message is written in Arabic script, Latin-letter Arabizi, or a mix of both within the same sentence.
Arabizi and code switching are the default, not the exception
Arabizi (also called Arabish or Franco-Arabic) is the transliteration system that grew out of early phones and chat apps that could not render Arabic script. It uses Latin letters and numerals for sounds that have no Latin equivalent, so "3" stands in for ain and "7" for a hard h sound TranslatorHub. If your eval set has no Arabizi in it, you built an Arabizi chatbot test suite from a channel your customers do not actually use for quick messages.
Code switching, mixing Arabic with English or French inside the same message, is a routine feature of Arabic digital communication and a distinct problem from dialect identification, not a corner case you can defer arXiv. A message that mixes an Arabic sentence with an English product name or the word "invoice" is normal, not broken. Your intent classifier and your fallback logic both need to be tested against it, and your eval set should carry a code-switch tag alongside the dialect tag.
Decide the register per intent, not for the whole bot
Do not pick one "Arabic voice" for the bot and apply it everywhere. Decide register at the intent level. A returns policy explanation reads better in a neutral, MSA-leaning register because it is a formal statement the customer may reread. A delivery status quick reply reads more naturally in Gulf colloquial because that is how the answer would be said out loud.
Build a small register table next to your intent list: intent, target dialect, formal or colloquial, and whether Arabizi input should get an Arabic-script or Latin-script reply. Route on the input signal you actually detected, not on a guess made once at design time.
Get the Arabic RTL chat widget's text direction right
A chat widget carries Arabic replies, English product names, order numbers and Arabizi input in the same thread, often in the same message. Setting the container direction and leaving individual strings unset is the most common bug in an Arabic RTL chat widget.
Set the container's dir="rtl" when the conversation is Arabic, and use dir="auto" on individual message bubbles so the browser's Unicode Bidirectional Algorithm can order left-to-right and right-to-left runs correctly without you manually reordering characters MDN. This matters most for punctuation, numbers and short English fragments sitting inside an Arabic sentence, which is exactly what a code-switched customer message looks like.
For dynamically inserted text whose direction you cannot predict at render time, an order number, a customer's name, a product title pulled from a database, wrap it in a bdi element so its direction does not bleed into the surrounding text and scramble punctuation at the edges MDN. This is a small addition and it fixes a visible, recurring bug: a phone number or order ID that displays with digits reversed inside an otherwise correct Arabic sentence.
Route WhatsApp with its own rules
If the bot lives on WhatsApp, two limits shape the design. Outbound template messages use the generic language code ar for Arabic, not a country-specific variant Meta for Developers, so template libraries do not need a separate Gulf, Egyptian or Levantine file. Free-form session replies are capped at 4,096 characters Meta for Developers, which matters for Arabic more than English because RTL punctuation and longer word forms can push the same idea past the limit sooner. Build a truncation and "read more" pattern into any answer that pulls from a long policy document. For the cost side of running this on WhatsApp, see WhatsApp Business API pricing UAE, explained.
Log and review misses every week
Every unanswered or misrouted message goes into a review queue, tagged with dialect, script and intent. Review it on a fixed weekly cadence, not "when there's time." Look for patterns: a specific dialect the bot keeps failing on, an Arabizi spelling variant it does not recognize, a code-switched phrase it routes to the wrong intent.
Fix the pattern, not the single message, and re-run your eval set before the next release. This is the step most teams skip after launch, and it is the one that actually improves the bot over time. Nobody can promise a resolution rate up front, and you should not accept one as a specification.
Store transcripts under the UAE PDPL
If your bot handles customer names, phone numbers, order details or complaints, those transcripts are personal data under the UAE's federal data protection law, Federal Decree-Law No. 45 of 2021, which has applied since 2 January 2022, with the UAE Data Office acting as the federal regulator UAE government. As of a January 2025 legal review, the law's Executive Regulations had still not been published, which affects how some provisions are interpreted in practice DLA Piper.
This is not legal advice. Confirm your retention period, storage location and consent language with counsel before you ship, and check current guidance before relying on anything written here. Practically, that means deciding retention windows up front, storing transcripts in a country-appropriate location, and building a deletion path rather than treating chat logs as permanent training data by default.
Country notes: UAE, Saudi Arabia, Kuwait and Oman
Dialect mix, regulatory detail and channel preference shift by country even within the Gulf. Rather than repeat those specifics here, they are covered on the conversational AI service page, alongside the engagement models and starting scope for building one of these.
Novamind Technologies builds Arabic and bilingual chat systems as part of its conversational AI work for GCC companies, treating right-to-left rendering and dialect handling as first-class requirements rather than something bolted on after launch. The same bilingual and RTL discipline shows up in the GSG Academy platform, built for German Standard Group across iOS, Android, web and an admin panel. For more on how Arabic support bots hold up under real customer service load, see Arabic AI customer service: what works in the Gulf in 2026.
Frequently asked questions
What is the difference between Gulf Arabic and Modern Standard Arabic for a chatbot?
MSA is the formal, standardized register customers expect in written policy or legal text. Gulf Arabic is the everyday spoken variant, generally described as covering the Arabian Peninsula, and it reads naturally in quick replies like order status or a booking confirmation. A working bot needs both, applied per intent rather than as one fixed voice.
Can a chatbot understand Arabizi (Arabic written in Latin letters)?
Arabizi uses Latin letters and numerals such as 3 and 7 to represent Arabic sounds with no direct Latin equivalent, and it is common in quick chat and WhatsApp messages. A model can be tested against Arabizi input, but no dialect or script variant should be described as fully supported; build an evaluation set that includes real Arabizi messages and track misses weekly.
Do I need separate WhatsApp templates for each Arabic dialect?
No. WhatsApp's template system uses the generic two-letter code ar for Arabic rather than a country or dialect-specific code, so one Arabic template library covers Gulf, Egyptian and Levantine customers. Register differences between dialects are a content decision inside that template, not a separate template file.
How long can a WhatsApp chatbot's Arabic reply be?
Free-form (session) text replies on the WhatsApp Cloud API are capped at 4,096 characters. Because Arabic script and RTL punctuation can run longer than an equivalent English sentence, build a truncation or read-more pattern for any answer pulled from a longer policy document.
Does UAE law cover chatbot transcript storage?
UAE-based personal data, including names, phone numbers and complaint details captured in a chat transcript, generally falls under the UAE's federal data protection law, which has applied since 2 January 2022 with the UAE Data Office as regulator. This is background information, not legal advice, so confirm retention, storage location and consent wording with counsel before you launch.
Sources
- Languages - WhatsApp Cloud API - Meta for Developers
- Text - WhatsApp Cloud API - Meta for Developers
- Data protection laws | The Official Platform of the UAE Government
- Data protection laws in UAE - General - Data Protection Laws of the World
- dir HTML global attribute - MDN Web Docs - Mozilla
- bdi HTML bidirectional isolate element - MDN Web Docs
- QCRI @ DSL 2016: Spoken Arabic Dialect Identification Using Textual Features - ACL Anthology
- A Novel Dialect-Aware Framework for the Classification of Arabic Dialects and Emotions
- Arabizi Translator Free - TranslatorHub
- Arap-Tweet: A Large Multi-Dialect Twitter Corpus for Gender, Age and Language Variety Identification
Novamind Technologies builds web, mobile, eCommerce, AI and automation, desktop and custom software for GCC businesses: fixed-scope pricing with published starting figures, bilingual delivery, and code you own outright. See the services page for starting prices, or contact us to brief a project.