





At events, an AI avatar is most often a booth for guests: guests have their photo taken, record a short voice sample and about five minutes later get a vertical video in which their digital twin speaks in their own voice. That is how we worked at the WMT AI conference: 100+ personal avatars during the event, about five minutes per guest. The two other formats are a digital host on screen who opens sessions and announces speakers, and an avatar of a speaker who could not attend. Below, we cover what we need from the organizer, what it costs and where an avatar does not work at a venue.
















It is an attraction and a way to collect content at the same time: guests have their photo taken, record a short voice sample and a few minutes later get a video in which their digital twin says a short text about the event that the guest has approved. The organizer gets a library of the participants’ personal videos with signed consent forms; the guests get a reason to post the video on their own channels.
The avatar opens the event, announces sessions, and reads out the house rules and partner messages. These are recorded videos based on an agreed script: we produce them before the event, and they are played from the AV desk. The upside: the text can be rewritten the day before the event without a reshoot. The downside: the avatar does not react to what happens in the room, so you still need a live moderator.
The speaker records a smartphone video and a voice sample following our instructions; we build the avatar and voice their talk in their own voice. We use this format on one condition: the audience knows it is watching a generated video. We do not pass an avatar off as a live broadcast.
The organizer wanted the booth to draw a crowd at the conference and offer more than handouts. We set up an express pipeline on site: a photo of the guest → voice cloning with ElevenLabs in one to two minutes → AI avatar generation → a finished vertical testimonial video for Reels and Shorts.
During the conference we made 100+ personal avatars, about five minutes per guest. The side benefit turned out to be worth more than the main one: WMT AI now has a library of the participants’ personal videos with signed consent forms. It uses them in its marketing, labeled as created with AI.
What worked: a short testimonial script of 15 to 20 seconds, a quiet corner for voice recording and a consent form printed in advance. What got in the way: venue noise on the recordings and guests who wanted to improvise a long text (the model turns that into a flat, lifeless delivery).
We price the format for your event date
A host on screen, a speaker who cannot attend in person, or a booth for guests: we quote the price and timeline for your program.
A host on screen needs the program and the house rules two weeks ahead: we write the script, and the host announces real people and real sessions, so a speaker moved at the last minute means a new video. The host’s face can be the organizer, a partner, a person the guests know, or a mascot if a human face is not needed.
The booth needs a spot with even light for photos, a quiet corner or a small recording booth for the voice, a stable internet connection, a screen to show the finished videos and a table for the consent forms. Each guest gives written consent on a separate form: we provide the template, and the guest signs it before recording.
For an absent speaker, we need their smartphone video and voice recording, made following our instructions, and the text of the talk approved by the speaker. This takes 14 days, as with a regular avatar.
Questions from the audience, a discussion, reacting to a delay or a mix-up: that is the job of a live moderator. The avatar in a package is a set of recorded videos and does not talk with anyone in real time; we build an interactive format that responds to the audience as a project from €15,000.
A voice recorded two meters from the stage produces a model with artifacts; without a quiet zone the booth does not work, and that has to be planned into the venue layout.
Reportage, interviews, audience reactions and shots of people need a camera crew. The avatar covers the talking head and the guests’ personal videos.
Small events of twenty people, where a live host costs less and fits better, and events with no plans for content afterward, since the booth pays off through the videos guests share on their social media.
We price an event by format. Host videos and an absent speaker’s avatar fall under the regular AI avatar packages: START includes the avatar, the voice clone and 20 videos for €2,900, PRO 30 videos for €3,900, EXPERT 50 videos for €5,500; the timeline is 14 days (EXPERT 21 days). For one event START is usually enough: the opening, the house rules, session announcements and partner messages fit into twenty short videos.
We price the guest booth by the number of guests and the hours on site: they determine how many workstations and staff we set up. We name the exact figure based on the brief, once we know the format and the expected number of attendees.
Book the host at least three weeks ahead: two weeks go into the avatar and the videos, one week into program changes, which always happen. Book the booth a month ahead, so there is time to test the pipeline and agree on the consent form with the organizer’s lawyer. Contact us on Telegram at @clyovo_ai, by phone at +7 919 997 99 62 or by email at german@clyovo.ru.
Questions clients ask before ordering, with our answers.
The avatar in our packages does not answer questions: it is a set of recorded videos based on a pre-written script. It can open the event, announce a session and read a partner message, but it will not answer a question from the audience. We build an interactive avatar that answers the audience as a project from €15,000.
At the WMT AI conference each guest took about five minutes, and we made 100+ avatars during the conference. Throughput depends on the number of workstations at the booth and on whether there is a quiet zone for recording: without one, the pace drops because recordings have to be redone.
Yes, in writing, as a separate document. We provide the form template, you check it with your lawyer, and each guest signs it before recording. We do not start without the guest’s signature.
The guest gets a vertical video for Reels and Shorts in which their digital twin says a short text in their own voice: a testimonial about the event or a partner, or an answer to the question of the day. The organizer gets a library of these videos, with the participants’ consent, to promote the next event; every post has to say that the video was created with AI.
Yes, the voice-over can be produced in several languages; for FORCELAB we made a multilingual avatar for different markets. What matters: a native speaker should approve the text in each language, and we write out abbreviations and names the way they should sound and listen to every replacement.
The organizer’s, or a face the guests recognize: that works better than the other options. If a human face is not needed or the person does not want to be on camera, we use a mascot: for SIBUR’s internal AI course, a character turned out to fit better than a host. People trust a drawn character less, so for partners’ sales pitches a human face is better.
The host: three weeks ahead, which covers 14 days for the avatar and the videos plus a buffer for program changes. The booth: a month ahead, since the pipeline has to be tested at the venue or a similar space and the consent form agreed. The absent speaker’s avatar: 14 days from receiving their video and voice recording.