Artificial intelligence
Artificial intelligence (AI) is the field of computer science that builds systems to perform tasks that normally need human intelligence, such as understanding language, recognising images or deciding from data.
In customer service, AI answers customer questions or sorts incoming messages and routes them to the right person.
Generative AI
Generative AI is AI that produces new content, such as text, images or audio, rather than only classifying data or picking a pre-written answer.
A system that writes a fitting reply about opening hours, instead of selecting a canned sentence, is being generative.
Large language model (LLM)
A large language model (LLM) is an AI model trained on vast amounts of text, so it can understand questions and write answers in natural language. The models behind ChatGPT and Gemini are examples.
Inside a voice agent, the LLM is the brain: it decides what to say and when to fetch information from a tool first.
AI agent
An AI agent uses a language model not just to answer but to take steps toward a goal, such as searching a menu or recording a booking, within permissions its owner sets.
Unlike a plain assistant, an agent acts: it calls tools, reads the results and decides the next step.
AI voice agent
An AI voice agent talks with people by voice: it hears the caller, turns speech into text, decides on a reply and speaks it in a natural voice, all in the same live conversation.
A restaurant busy with phone orders at peak hours might use one to answer menu questions, take the order and read it back before sending it to the kitchen.
Chatbot
A chatbot is a program that converses in text, usually on a website or messaging app. Some follow fixed scripts; others use a language model that handles open questions.
The key difference from a voice agent is the channel: a reader tolerates a paragraph, a listener needs short sentences.
IVR (interactive voice response)
Interactive voice response (IVR) is a phone system that plays recorded prompts and asks callers to press keys or say specific words to route their call, as in “press 1 for orders”.
IVR handles simple routing but cannot understand a free-form request like “two shawarma meals, no garlic, and two drinks”, which a voice agent can.
Speech recognition (ASR/STT)
Speech recognition, also called ASR (automatic speech recognition) or STT (speech to text), converts spoken audio into text a system can process.
Accuracy depends on noise, line quality and dialect, so test it on real calls in the local dialect, not just clean recordings.
Text to speech (TTS)
Text to speech (TTS) reads written text aloud in a synthetic voice; modern TTS sounds close to human speech in tone and rhythm.
The practical challenge is pronouncing local dish names, numbers and addresses correctly, which often means hand-tuning specific words.
Latency
Latency is the time between the caller finishing a sentence and the agent starting to reply. The shorter it is, the more natural the call feels; the longer, the more the caller suspects they were not heard.
It adds up across stages: speech recognition, model reasoning, tool calls and voice generation. Measure each stage separately to find where time goes before cutting it.
Barge-in
Barge-in is a voice agent’s ability to stop speaking as soon as the caller starts talking and listen instead, as people do in normal conversation.
If the customer cuts in while the agent lists dishes, it should stop and take the choice:
AgentToday we have mansaf, maqluba, chicken kabsa, and…
CustomerSorry, hold on. Maqluba, please.
AgentSure, one maqluba. Anything to go with it?
CustomerNo, that’s it, thanks.
Turn-taking and voice activity detection (VAD)
Turn-taking is how a conversation decides who speaks when. It relies on voice activity detection (VAD), which tells speech apart from silence or noise.
Tuning is delicate: wait too long and the agent feels slow; reply too fast and it talks over a customer who was pausing to think.
Grounding
Grounding means tying a language model’s answers to trusted data, such as the actual menu or opening hours, instead of what it absorbed in training.
For example, in ABRAJRUM’s voice agent Sara, prices come only from menu tools, so the agent never speaks a price no tool returned. That reduces errors, but it also means answers are only as accurate as the menu the restaurant keeps up to date.
Tool calling
Tool calling is a language model’s ability to request an external function, such as looking up an item or checking table availability, and use the result in its reply.
So the agent does not guess whether a table is free at 8 pm; it asks the booking system. Permissions stay with the owner: a tool does only what it is allowed to.
Hallucination
A hallucination is when a language model confidently states something false or invented, such as a dish not on the menu or a wrong price.
They are reduced by grounding answers in real data, forbidding answers without a source, and reading the order back for the customer’s yes before confirming.
Dialect and code-switching
A dialect is the spoken Arabic of a region, such as Levantine, Gulf or Egyptian, differing from Modern Standard Arabic in vocabulary, pronunciation and structure.
Code-switching is mixing two languages in one sentence, such as ordering in Arabic with the word “delivery” or an English dish name. A good voice agent follows without asking the caller to repeat.
SIP and the phone line
SIP is a signalling protocol for voice calls over the internet and the usual way to connect software, such as a voice agent, to a real phone line from a carrier.
The customer dials the restaurant’s usual number and the call reaches the agent through a SIP-capable carrier. This needs an arrangement with a carrier in each country.
Escalation (human handoff)
Escalation, or human handoff, moves a conversation from the agent to a person when the caller asks or the case is outside the agent’s permissions, such as a complaint.
It can be a live call transfer or an alert so staff call back. At ABRAJRUM, for example, the agent alerts the restaurant manager on WhatsApp and logs the request for follow-up.
POS and KDS
A point-of-sale (POS) system records orders and payments; a kitchen display system (KDS) shows orders to cooks on screens, split by station, instead of printed tickets.
For a voice agent to be useful, its orders should reach the POS and kitchen display directly, not be re-keyed by staff, which invites mistakes.
Quick reference table
Each term in one line, for quick lookup.
| Term | In one line |
|---|---|
| Artificial intelligence | Systems that perform tasks that normally need human intelligence. |
| Generative AI | AI that creates new text, images or audio. |
| Large language model (LLM) | A model that understands language and writes natural answers. |
| AI agent | A system that takes steps toward a goal within set permissions. |
| AI voice agent | An agent that listens, understands and replies by voice in real time. |
| Chatbot | A program that converses in text. |
| IVR | Recorded phone menus driven by keypresses or set words. |
| Speech recognition (ASR/STT) | Turning spoken audio into text. |
| Text to speech (TTS) | Reading text aloud in a synthetic voice. |
| Latency | The gap between the caller finishing and the agent replying. |
| Barge-in | The agent stops talking when the caller speaks. |
| Turn-taking and VAD | Deciding who speaks when by telling speech from silence. |
| Grounding | Tying answers to trusted data such as the real menu. |
| Tool calling | The model requests an external function and uses the result. |
| Hallucination | False or invented information stated confidently. |
| Dialect and code-switching | Regional spoken Arabic, and mixing two languages in one sentence. |
| SIP | The protocol that connects software to a real phone line. |
| Escalation (human handoff) | Passing the conversation or request from the agent to a person. |
| POS and KDS | The order and payment system, and the kitchen order screen. |
Common questions
What is the difference between an AI voice agent and a chatbot?
Both may run on a language model, but a chatbot works in text while a voice agent speaks. A voice agent therefore also needs speech recognition, voice generation, low latency and the ability to handle interruptions.
Can a voice agent be stopped from hallucinating completely?
No one can guarantee that, but it can be reduced a great deal: take prices and items from tools linked to real data, forbid answers without a source, and read the order back and wait for an explicit yes before sending it.
Why does latency matter more on calls than in text chat?
Silence on a call is noticed immediately. A caller has no typing indicator, so a pause sounds like a dropped line or an unheard sentence, and they repeat themselves or talk over the agent.
Does a voice agent understand every Arabic dialect?
That depends on the speech recognition model and on testing it with real calls in each dialect. It is best to set the agent to the dialect of the market it serves and test it on local dish names and addresses before going live.
What is the difference between ASR and TTS?
They run in opposite directions. ASR (speech recognition) turns the caller’s voice into text the system can read; TTS (text to speech) turns the agent’s written reply back into spoken audio. A voice agent needs both, with a language model in between.
