Find out where AI would save your business the most time. Take the free AI audit →

←Case studies

Live AI on stage

An AI took a seat on the panel at ForsaTEK. The hard part was teaching it when to stay quiet.

100-200ms

Response time, live on stage

Emirates Group case study

ForsaTEK is Emirates Group's flagship innovation event. One of its panels, Building the New AI Industrial Revolution, ran with four names on the stage screen. Three were executives. The fourth was Maryam 3.0, a virtual AI panelist that answered the moderator out loud, in the room, for the full hour. We built it with KalaMena AI, who led the engagement with Emirates Group.

The ForsaTEK stage screen listing Maryam 3.0 as a virtual AI panelist alongside the human panelists and the moderator
The panel lineup on the stage screen. Maryam 3.0 third from left, developed by KalaMena AI and Hephon.

The brief was a conversation, not a demo

Playing a recorded clip on a screen is the safe way to put AI in front of an audience. It is also obvious, and a room full of senior leaders can tell the difference. The brief here was a real seat on a real panel. The moderator asks, the AI answers, and the audience watches it happen at conversational speed with everything that can go wrong still able to go wrong.

Answering was the easy part

Current models answer well. That was the part of the problem we worried about least. Almost all of the engineering went into everything that surrounds the answer: hearing the room correctly, working out who is speaking, deciding whether anyone is actually waiting on a reply, and getting sound out of the PA fast enough that the exchange still reads as a conversation.

A panel is also not a tidy sequence of questions. People talk over each other. A sentence trails off and somebody else finishes it three seconds later. A rhetorical question gets thrown at the room with no answer expected. A good share of what is said is not addressed to anyone in particular.

An open microphone, and no button to press

Most voice assistants get an easy start in life. A wake word or a button tells the system that speech is about to arrive and that it is meant for the assistant. On a panel there is no button. The microphones are open for the whole hour, every voice in the room is a potential input, and the system has to decide for itself what is a question, what is conversation between two humans, and what is noise.

So the pipeline runs continuously rather than in request and response turns. Audio streams in, transcription runs live against it, and the decision about whether to answer is made on a transcript that is still being written. Nothing waits for a complete utterance before it starts thinking about it, because waiting is where the seconds go.

Endpointing: deciding when a question is actually finished

The classic way to decide that a speaker has finished is to wait for a fixed stretch of silence. On stage that fails in both directions. Panellists pause mid-sentence to think, which a silence timer reads as the end of a turn, and they run two questions together with no gap at all, which the same timer reads as one. Answering into a thinking pause is the single most damaging thing a system like this can do, because interrupting an executive in front of an audience is the kind of thing people remember.

Endpointing therefore has to look at what was said, not only at the gaps. Whether the sentence is grammatically complete, whether it ended on a question, whether the speaker is mid-clause, and how long this particular person usually pauses. Tuning that against real rehearsal audio rather than against a threshold in a config file is most of the work.

Working out when a question is meant for it

Being addressed is a separate problem from a turn ending. The moderator can hand a question to Maryam by name, put the same question to a human panelist, or open it to everyone. The system has to read a direct address, notice when a question has already been handed to somebody else, and stay quiet when the room is talking among itself. A system that answers everything it hears becomes unbearable inside two minutes.

Voice identity, so it knows who it is answering

Maryam 3.0 recognises the people on the panel by their voice. Every stretch of speech is attributed to a specific person as it is transcribed, which does three jobs at once. It lets the system answer the right person. It keeps the running transcript labelled, so a reference to something said earlier resolves to whoever said it. And it lets the model behave differently towards the moderator running the session than towards a panelist making a point.

Not talking over anyone, and not listening to itself

Two failure modes sit on either side of the microphone. If a human starts speaking while Maryam is answering, she has to yield rather than talk through them. And because the answers come out of the same PA that the microphones are listening to, the system will hear itself speak. Without handling that, its own answer arrives back as a new question, and the panel spends the rest of the hour listening to a machine interviewing itself.

The latency budget

Response time landed between 100 and 200 milliseconds, measured from the end of the question to the first sound of the answer. That is roughly the gap a person leaves before replying in ordinary conversation. Above a second, the room notices the delay and starts treating the thing as a machine being operated. Under 200 milliseconds it reads as a panelist taking a beat.

A budget that tight only works if the stages overlap. Transcription runs while the speaker is still talking, so the question is largely understood by the time it ends. Generation starts on the partial transcript rather than after it. Speech synthesis streams, so the first words are already in the air while the rest of the answer is still being produced. Nothing in the chain is allowed to wait for the stage before it to finish, and every hop, including the one out to the sound desk, is part of the budget.

Holding an hour of context without drifting

The panel ran for an hour, and a live conversation refers back to itself constantly. Someone picks up a point made forty minutes earlier, or asks a follow-up that only makes sense against what was already said. The running labelled transcript is what the model reasons over, which keeps answers consistent with the rest of the session rather than treating each question as a fresh start.

The same window is what keeps the persona stable. An hour is long enough for a system to slide out of character or to start agreeing with whatever was last said to it, and on a stage that drift is visible to everyone in the room.

Guardrails, on a stage with no undo

Guardrails were set before the event and held for the duration. In front of senior leadership, industry partners and government stakeholders, a hallucination is not a bug report. It is the clip that circulates afterwards. The system stayed inside its brief for the full hour, said nothing it could not support, and made nothing up.

A live transcript of the whole room

Everything said during the panel was transcribed as it happened, attributed by speaker, including Maryam's own answers. That transcript is what turns a stage moment into something an organisation can review afterwards: what was asked, who asked it, what the system said back, and at what point in the hour.

Delivered with KalaMena AI

KalaMena AI led the engagement with Emirates Group and brought Hephon in as the technology partner. Maryam 3.0 was built and run jointly, and the stage credit reflected that. Hephon engineering was on site for the event, on the console for the length of the panel.

What we took from it

Putting an AI into a live conversation with people is a different discipline from putting one into a product. The model is a component, and by now a fairly reliable one. The engineering sits in the timing, in knowing who is speaking, in deciding when the system should say nothing at all, and in having a plan for the moment something breaks in front of a full room.

Ready to kickstart your project?

Speak with us

Find out where AI would fit in your business.

A calm landscape in warm light

Already know what you need?Let’s talk.

Josef, co-founderAyman, co-founderBook a call30 min with a founder

Hephon Agent

By chatting you agree to our Privacy Policy.

We use cookies for analytics and advertising, to understand how the site is used and improve it. You can accept or keep them off. The site works either way. See our privacy policy.