From customer conversations to operational intelligence.
Sections one to nine of this report use only Khedmah's production chatbot Statistics dashboard for the periods stated. The add-on deep dive into “Other” and free text uses a separate conversation-level intent analysis of production messages from 27 July 2026 onward, and is labelled as such throughout; its figures are approximate classifications and are never mixed with dashboard counts. Where the dashboard cannot establish causality, findings are described as patterns, signals or hypotheses requiring validation. Food Issues and Payment Issue are parent categories and are never summed with their children. A 31 August–2 September screenshot exists but covers three days and crosses into September; its absolute volumes are excluded from the five-week comparison.
Adoption is established. Weekly volumes held around the 500-chat level for five consecutive weeks, and the question has moved from whether customers will use the channel to where the channel should be optimised. Five measures define that agenda.
The largest free-text intent already has a structured workflow — showing that better AI routing may deliver substantial value without requiring a completely new automation flow.
Refund status is one of the clearest candidates for a new self-service journey.
Two of these are demand-shaping opportunities — order visibility and unstructured intent. Two are workload measures that have not yet bent downward. One is a policy question that has grown sevenfold in five weeks.
103 unique users opened 71 chats on day one, 58 of them single-issue. Language split almost evenly: 40 English, 31 Arabic. Twelve refunds were issued and 31 tickets raised.
Children sit inside the parent total. They are not additional complaints.
Day-one tickets were led by Item Missing (13) and Other (6), with Quality No Photo, Misconduct and Subjective Feedback at three each, and single tickets for Payment, Food Not Received and a Refund Guardrail Block.
Raw counts moved within a narrow band, so the shape of demand is best read as a share of chats. Launch day is shown for reference but is a single day and not comparable to a seven-day period.
| Measure | Launch day | Week 1 | Week 2 | Week 3 | Week 4 | Week 5 |
|---|---|---|---|---|---|---|
| Total chats | 71 | 533 | 578 | 489 | 492 | 535 |
| English / Arabic | 40 / 31 | 331 / 202 | 379 / 199 | 320 / 169 | 283 / 209 | 349 / 186 |
| Refunds issued | 12 | 49 | 51 | 54 | 42 | 56 |
| Tickets raised | 31 | 199 | 222 | 227 | 201 | 247 |
| Track Order share | 42.3% | 41.7% | 52.9% | 48.5% | 50.6% | 53.8% |
| Other Free Text share | 22.5% | 25.7% | 26.3% | 26.8% | 28.3% | 27.9% |
| Food Issues share (parent) | 26.8% | 15.2% | 17.0% | 18.4% | 16.3% | 17.9% |
| Payment Issue share (parent) | — | 8.4% | 9.2% | 7.4% | 6.7% | 6.0% |
| Multi-issue chat rate | 18.3% | 21.6% | 23.5% | 25.6% | 25.6% | 27.1% |
| Ticket intensity (per 100 chats) | — | 37.3 | 38.4 | 46.4 | 40.9 | 46.2 |
Table 1. Weekly demand and workload measures. Parent-category shares exclude their child categories from the count.
Track Order moved from 41.7% of chats in Week 1 to 53.8% in Week 5 — and these queries are handled end to end by the bot, so the growth is growth in fully self-served contact.
This is not a chatbot failure — it is the opposite. Track Order is fully self-serve: the bot resolves these journeys without human intervention, so the largest single block of customer demand is being absorbed by automation rather than by the support team. The next opportunity is to move from reactive status answering to proactive communication, and prevent part of the demand altogether.
Track Order queries are handled entirely by the AI bot. The rise from 41.7% to 53.8% represents an expanding share of contact deflected from human agents.
Can Khedmah answer “Where is my order?” before the customer needs to ask?
Other Free Text held between 25.7% and 28.3% of chats every full week — a remarkably persistent share.
This report does not contain the content of those conversations, so it cannot say what the missing intents are. It can say that this is the largest single coverage gap.
In addition to the Statistics dashboard, a separate conversation-level analysis was performed on actual production chatbot messages from 27 July 2026 onward. It looked at customers who selected “Other / My issue is not listed here” and then explained their problem, and customers who typed a meaningful problem directly in free text without first selecting a predefined issue category.
chats involved an explicit Other / issue-not-listed journey
chats contained meaningful free-text intent before a predefined issue selection
order-level conversations useful for understanding natural-language demand outside button/menu behaviour
This is a conversation-level qualitative and intent analysis, separate from the production dashboard metrics. The 1,365 figure should not be reconciled directly with the dashboard's Other Free Text count: the two populations answer different questions. Category counts below are analytically classified dominant intents.
Which predefined reason or category was recorded?
What were customers actually trying to tell us in natural language?
Free-text conversations broadly fall into two types, and each has a different solution. The distinction is strategically important because one requires no new backend workflow at all.
The customer simply typed naturally instead of navigating through the predefined buttons.
The customer is asking for something that does not currently have a clear structured journey.
The free-text data reveals that a relatively small set of repeat customer needs accounts for a large proportion of natural-language support demand.
Figure 4. Analytically classified dominant intents. Wording is deliberately approximate because some conversations are classified with judgement. Smaller categories include pickup-related queries, technical issues, pricing and charge questions, complaint follow-ups and miscellaneous customer-service requests.
Khedmah already provides a predefined Cancel My Order journey, so the problem is not a missing category. The category exists. The customer should not need to find the button: when cancellation intent is typed, the chatbot should recognise it and take the customer directly into the existing workflow.
The customer should not need to find the button.
Four clusters stand out as candidates for new structured journeys, each with enough conversation volume to justify design work.
12.9% of analysed demand; 90+ specifically about refunds pending or not received
order changes plus address, phone and delivery-instruction changes
availability issues plus orders rejected, cancelled or not accepted
rider needs well beyond the current misconduct journey
Customers ask about refunds not received, refunds pending, when a refund will arrive, cancelled orders where money has not returned, missing or partial refunds, and expected settlement time. Approximately 90+ conversations related specifically to a refund being pending or not received.
69 conversations involved adding or removing an item, changing a selected item, flavour or preference, or adding restaurant instructions. A further 58 involved changing a delivery address or phone number, adding a contact number, or apartment, gate and delivery-location instructions.
Restaurant, store and availability issues at approximately 66 conversations, and orders rejected, cancelled or not accepted at approximately 57: restaurants not accepting orders, unavailable restaurants or items, orders awaiting confirmation, and unexpected rejections. A distinct customer state from general order tracking.
Riders going in the wrong direction, unreachable riders, orders not picked up, no rider assigned, rider changes, location confusion, requests for rider contact details, payment interaction with the rider, and actual misconduct. The existing Report Delivery Partner Misconduct journey covers only part of this.
Loyalty points, missing or deducted points, points after failed orders, coupons, discount codes, Buy 1 Get 1, promotional offers, free delivery, and membership questions — a different need from delivery, food complaints or payment failure.
Clear requests for an agent, customer service, a support representative or phone support. Not automatically a signal of dissatisfaction — some issues genuinely require human intervention.
Within the large cancellation cluster, accompanying customer language indicates multiple possible reasons. These counts are a qualitative signal only — not every cancellation conversation states a reason, so they should not be turned into a complete cancellation-reason breakdown.
The system is not always converting that natural language into a structured journey.
Map natural-language requests automatically into workflows that already exist.
Build workflows for high-frequency customer needs that are not yet covered.
Four vendors and one branch stand well outside the norm on a specific, named problem.
~3.3× higher concentration at KD-Mart than across Food-vendor support chats. At the MBD branch, 22 of 51 support chats involved Item Missing — and in Week 5 alone, 10 of 14.
The chatbot is moving beyond customer support and becoming an operational early-warning system — identifying which vendors and branches Khedmah should investigate first.
Vendor percentages represent issue concentration among chatbot and support conversations, not percentage of total vendor orders.
Multi-issue chats rose from about 22% to about 27% across the five full weeks.
Figure 1. Multi-issue chat rate, Weeks 1–5. Launch day was 18.3% and is excluded as a single-day observation.
Customers genuinely have multiple problems.
An unresolved first journey causes customers to select another issue.
Analyse multi-issue journey sequences to determine which explanation dominates. The answer changes what to fix in the chatbot's UX.
Human and operational workload remains a major optimisation opportunity.
Tickets generated per 100 chats rose from 37.3 in Week 1 to 46.2 in Week 5, with a dip in Week 4. This is deliberately called ticket intensity rather than an escalation rate: the dashboard does not establish a guaranteed one-to-one mapping between a chat and a ticket.
Figure 2. Tickets generated per 100 chats.
Which ticket categories represent unavoidable human intervention, and which can be automated further?
Six categories carry most of the ticket load. Other, Item Missing and Refund Guardrail Block are the movers; Payment and Food Not Received are the two clear improvements.
| Ticket category | W1 | W2 | W3 | W4 | W5 | Direction |
|---|---|---|---|---|---|---|
| Other | 31 | 52 | 52 | 43 | 67 | Rising — largest category |
| Item Missing | 38 | 35 | 44 | 33 | 38 | Flat at a high level |
| Food Not Received | 33 | 30 | 21 | 25 | 23 | Lower than Week 1 |
| Refund Guardrail Block | 5 | 12 | 24 | 24 | 36 | Rising 7.2× since Week 1 |
| Payment | 28 | 23 | 26 | 12 | 10 | Down ~64% |
| Refund Escalation | 13 | 12 | 15 | 20 | 17 | Broadly stable |
| Full Item Refund Review | 15 | 17 | 15 | 17 | 16 | Stable |
| Quality No Photo | 12 | 13 | 11 | 6 | 17 | Volatile |
| Subjective Feedback | 10 | 12 | 10 | 6 | 8 | Low, stable |
| Misconduct | 10 | 5 | 5 | 8 | 2 | Down from 10 to 2 |
| Cold Food | 3 | 7 | 1 | 3 | 5 | Low volume |
| Refund API Error | 1 | 4 | 3 | 4 | 8 | Week-5 high |
Table 2. Ticket categories by week. Shading is relative to the highest cell in the table (Other, Week 5 — 67).
Figure 3. Refund Guardrail Block tickets per week.
Safeguards firing is not a failure. Guardrails may be protecting Khedmah from inappropriate automated refunds. But activation at this rate warrants a policy review, because the same mechanism can also stop legitimate customers.
Are predominantly risky claims being blocked, or are legitimate customers increasingly reaching the safeguard?
Absolute volume remains small, but Week 5 was the highest level recorded. This report does not speculate about root cause and recommends engineering investigation now, while the numbers are still small.
Failures at the final refund action stage can undermine an otherwise successful automated resolution journey.
| Item Missing | W1 | W2 | W3 | W4 | W5 |
|---|---|---|---|---|---|
| Reason volume | 49 | 55 | 67 | 61 | 70 |
| Share of chats | 9.2% | 9.5% | 13.7% | 12.4% | 13.1% |
| Tickets | 38 | 35 | 44 | 33 | 38 |
Reason incidence increased in the later weeks while ticket volume did not rise proportionately between Week 1 and Week 5. That gap should not be read as the chatbot resolving more cases automatically — the dashboard cannot support that conclusion.
The relationship between increased Item Missing demand and relatively stable ticket volume deserves deeper containment analysis.
Three categories ended the period materially below Week 1. These are positive operational signals; none is yet attributable to a specific system improvement without further evidence.
Roughly 64% lower by Week 5. Payment reason volume also fell from 45 to 32, and its share of chats from 8.4% to 6.0%.
Approximately 30% below Week 1 while reason volume stayed broadly stable at 16 → 15.
Lowest level of the period, from a low base throughout.
Food Not Received demand remained broadly stable while associated ticket volume was lower by Week 5. This may indicate improving handling, but it must be validated with conversation-level linkage before any automation claim is made.
Food Issues took 26.8% of chats on launch day, then 15.2%, 17.0%, 18.4%, 16.3% and 17.9% across Weeks 1 to 5. That is a settled band, not a continuous improvement.
The largest continuing issue inside the category, rising to 70 reason instances in Week 5.
Volatile but meaningful: 22, 38, 28, 19 and 36 across the five weeks.
Relatively low volume throughout: 8, 17, 7, 10 and 11.
Weekly chat volumes stayed roughly around the 500-chat level.
The largest use case is resolved end to end by the bot, rising to 53.8% of Week-5 chats.
Issue and ticket taxonomies are generating usable operational signal.
From 28 tickets in Week 1 to 10 in Week 5.
33 tickets to 23, with stable underlying demand.
Khedmah can now see week-on-week changes in what customers need.
The chatbot is not only a customer service interface. It is a real-time customer experience sensor.
The first five weeks answered “Will customers use the chatbot?” Six more valuable questions now define the roadmap.
How much order-status demand can be prevented proactively?
How much of Other can be converted into structured automation?
Why are multi-issue journeys increasing?
Which ticket categories can be resolved automatically?
Are refund guardrails optimally balancing CX and financial protection?
Can final-action technical failures be eliminated?
Month 1 established the channel. The data now shows where to optimise.
The next phase is to reduce avoidable contact, increase automation, and turn support data into proactive customer experience.
A 31 August–2 September dashboard view exists but covers only three days and crosses into September. Its absolute volumes are deliberately excluded from every comparison in this report and should be read only as an early pulse on the next period.