An AI avatar kiosk demos brilliantly. A human-looking face on a screen, speaking naturally, answering questions in Hindi and English. Everyone who sees it wants one. Then the practical questions arrive, and most of those projects quietly become a touchscreen. We build one of these — RIYA — so treat this as an interested party being honest about where the format actually pays for itself.

What an avatar adds that a touchscreen doesn’t

It handles the unstructured question. A touchscreen answers what you anticipated. A conversational system handles “I’ve come for my father’s cataract follow-up, which floor?” — a question no menu tree contains.

It works for people who don’t want to navigate a UI. A meaningful share of visitors at hospitals, government offices and banks won’t touch an unfamiliar interface, and will wait twenty minutes for a human instead. Voice removes that barrier in a way touch never has.

It works in Hindi and regional languages naturally, including for users who can read comfortably in neither English nor a Devanagari UI.

It absorbs repetitive questions. In most public-facing environments, a small number of questions account for the large majority of front-desk interactions. Offloading those is the entire business case.

Where it works

Banks. High volume of repetitive procedural queries — documentation for a loan, account opening requirements, cheque status, branch services. Long dwell time in a seated waitingarea. Strong brand incentive to look modern.

Airports. Wayfinding, terminal and gate queries, transfer and facility questions, and multilanguage demand by design. Extremely high query repetition. Add an accurate wayfinding engine with QR handover to the traveller’s phone and the utility is immediate.

Hospitals. Departments, doctor availability, OPD process, floor routing, insurance and admission documentation — asked hundreds of times a day, often by anxious first-time visitors. This is the single highest-value environment we see, and also the one that demands the most care around what the system is allowed to say. It must route and inform. It must never give clinical advice.

Government and citizen services. Document checklists, process explanation, queue and token guidance, form assistance. Massive query repetition, high language diversity, high visitor anxiety.


The pattern is consistent: high query repetition + high language diversity + captive waiting time + procedural (not clinical or financial-advisory) answers.

Where it doesn’t

QSR ordering. Ordering food is a structured transaction that a well-designed touchscreen completes faster than a conversation. Speed matters more than warmth at a lunch rush. Use a self-ordering kiosk.

Fashion and apparel retail. The genuinely useful applications here are visual — try-on, size, styling — which is an AR problem more than a conversational one, and there are specialists focused on it.

Anywhere the answer must be authoritative and liability-bearing. Medical advice, legal interpretation, financial product recommendation, insurance eligibility determination. A conversational system should hand these to a human, clearly and quickly.

Low-footfall locations. The economics need volume of interactions. A kiosk answering forty questions a day doesn’t repay itself. Forty an hour does.

What actually makes these hard

Anyone can wire a language model to a face. The difficulty is in the physical deployment:

Latency. Conversation collapses if the response takes too long. This needs local inference tuning, aggressive caching of common answers, and a graceful degraded mode when the network is poor — which, in an Indian airport basement or a hospital’s ground floor, it will be.

Interruption. Real people interrupt. A system that has to finish its sentence before listening feels broken within two exchanges.

Ambient noise. A concourse at peak hour is a hostile acoustic environment. Microphone array placement matters more than model choice.

Grounding. The avatar must answer from your verified data — your departments, your doctors, your process — not from general knowledge. Every deployment needs a controlled knowledge base and clear refusal behaviour outside it.

Idle and reset behaviour. What happens when someone walks away mid-conversation? If the next person inherits the last person’s context, you have a privacy problem.

The hardware around it. Thermals for a GPU running continuously, power, enclosure, service access. The AI is software. The thing in the lobby is a machine, and it will need someone to fix it.


How to evaluate a vendor

  1. Ask to see it in a live public environment, not a demo room. Noise and network conditions are the test.

  2. Ask what happens when the internet drops mid-conversation.

  3. Ask how the knowledge base is updated when a doctor’s schedule changes — by whom, in how long.

  4. Ask what the system does when asked something outside its scope. The right answer is a fast, clear handoff.

  5. Ask who services the physical unit, where, and how fast.

The honest summary

An avatar kiosk is not a better touchscreen. It’s a different tool, worth its premium in environments with heavy repetitive questioning, language diversity, and visitors who won’t use a UI. In a transaction-first environment, buy a self-ordering kiosk and spend the difference on uptime.

RIYA by Digitos is an AI avatar kiosk built on hardware we manufacture and service in India, with multilingual conversation, integrated wayfinding and offline degraded operation.

Leave a Reply

Your email address will not be published. Required fields are marked *