





Voice cloning means training a model on a recording of your speech; the model then reads any written text aloud in your voice. We do it with ElevenLabs as part of our AI avatar and AI video work: the voice clone is included in every AI avatar package, including START at €2,900. Promo videos use a narrator voice-over. We clone your own voice or the voice of a person who has given written consent; we do not work with recordings of anyone else. Below, we cover what to record, how it sounds on long texts and where the technology reaches its limits.












CLYOVO avatar
Standart Ecology · waste management, demolition and land reclamation
website





You send a voice recording, we build a voice model from it and use it to voice your scripts. The output is either an audio track or a finished video with an AI avatar whose lips are synced to that track.
In phrases and paragraphs up to half a minute long, the difference from a live recording is rarely audible. In long monologues the model gives itself away by its evenness: it does not stumble, does not take random breaths and does not speed up where a person would out of excitement. That is why we place pauses and emphasis by hand and split long texts into blocks that we then join.
What the model does well: calm, informative delivery, explaining a service, answering a frequent customer question, a training segment. What it does poorly: acting, shouting, whispering, comic intonation. That is work for a human performer, and replacing it with synthesis makes no sense.
We need a voice recording made following our instructions. We provide the text; any decent microphone will do, even a phone in a quiet room. Silence matters more than equipment: echo, an air conditioner and street noise outside the window spoil the model more than a basic microphone does.
Usually a few minutes of voice audio is enough, and we do not name an exact length before we have heard you. You start with a short test recording; we listen to it and build a draft voice. If the timbre comes through clearly, there is nothing to add. If your voice sounds tired in the recording or there is noise on the track, we ask you to re-record specific parts and tell you which ones. This way you do not spend an evening on a recording that will not work later.
On timing, the voice fits into the schedule of the main project: an AI avatar package takes 14 days from the start (EXPERT 21 days), promo videos and AI videos 10 to 21 days. There is no separate “voice day” in the plan: the model is built in the first few days, since otherwise we cannot check the lip sync.
Send a short voice message, and we will tell you whether your voice works for a model
No extra recording, no upfront payment: we listen to the sample and tell you which package makes sense.
We work with your own voice or the voice of a person who has signed consent. Consent is given in writing, as a separate document, by the person whose voice we clone.
In practice, this means the client company’s signature alone is not enough. If it is an employee’s voice, the employee signs. If you bring a recording of a narrator, a performer or a public figure, we decline the job.
For Igor Nikitin, founder of WMT, the voice was a requirement: his audience knows this voice, and a different voice would break that recognition. We cloned the voice with ElevenLabs, turned his face into an avatar and then released 11 vertical episodes for Reels.
The main gain from cloning was speed from the second episode on. Once the voice is built and the script approved, producing an episode takes under an hour: the text goes to voice-over, the track goes to lip sync, the video is exported. At this stage, Igor’s only involvement is approving the text: no recording, no calls.
Another example: Standart Ecology (waste management, demolition and land reclamation). Eight avatar videos are embedded on the service pages of its website, and visitor-to-inquiry conversion on the website rose by 36% (results from a single project, with no A/B test). The avatar is an AI-generated spokesperson for the company and does not depict a real employee.
It does not hold conversations. It voices a pre-written text: the avatar will not answer a question on a call, run a live stream or react to a viewer’s comment. An interactive avatar that responds in real time is not part of the package; we build it as a project from €15,000.
It cannot sing, it is not suitable for film dubbing, and a human performer beats it at any emotional peak.
The voice model is stored in the account of the service it was built on. We delete it at your written request and do not use it in other projects. If your company policy restricts where such data may be stored, raise it before the start: the service may not meet strict storage requirements, and it is better to find that out on the intro call.
Who does not need it: if you need one video and are ready to record it yourself, recording it costs less. Cloning pays off when there are many videos and they come out regularly. It also does not fix a weak offer: the voice makes a text speakable, and how convincing it is depends on the text.
Voice cloning is part of the AI avatar packages.
AI avatar: START includes the avatar, the voice clone and 20 videos for €2,900, PRO 30 videos for €3,900, EXPERT 50 videos for €5,500. The timeline is 14 days, 21 days for EXPERT. Promo videos and AI videos without an avatar: 20, 30 or 50 videos for €2,900, €3,900 or €5,500, with a timeline of 10 to 21 days. If you need videos on a steady basis, there is the Creative Unlimited content subscription: €4,500 per month for up to 40 videos in any format, with one manager and one monthly payment.
In the number of videos, the formats and how deeply the scripts are developed. Voice quality does not depend on the package: the model is the same.
Send a short voice message on Telegram to @clyovo_ai or +7 919 997 99 62, or email a recording to german@clyovo.ru. From it we will tell you whether your voice works for a model without extra recording and which package makes more sense for it.
Questions clients ask before ordering, with our answers.
Technically, yes, but we do not take on such work. We need written consent from the person whose voice is cloned, as a separate document. We do not work from recordings found online, old ads, podcasts or phone calls made without consent.
Usually a few minutes, recorded following our instructions. We do not promise an exact figure before we hear your voice: you record a short test, we listen to it before you record everything and tell you whether it is enough or which specific parts to add.
In short phrases, usually not. In a monologue of several minutes, an attentive listener notices the evenness of breathing and pace. We compensate by placing pauses by hand and splitting long texts into blocks, though the effect does not disappear completely.
Nobody can fully rule out a breach of an external service, and we do not promise that. What we do: we give no third parties access to the model, do not use it in other projects, delete it at your written request, and fix the term and the territory of use in the contract. If your data storage requirements are stricter, discuss them with us before the start.
They need manual work: abbreviations, brand names, foreign words and numbers are written into the script the way they should sound, and we listen to every replacement.
We build the voice as part of an avatar or video package. If you need only the audio track, write to us: we will tell you directly whether we take on such a task and on what terms.
We delete the model at your written request. Videos already released are a different matter: they have to be taken down from the platforms where they are published, and that is work on your side. That is why it is better to agree on the term and the territory of use in the contract from the start.