Appearance
AI Voices
AI Voices lets you clone your own voice, or design a synthetic one, for use in personalised video voiceovers in Studio.
Recording the voice
Follow these guidelines to maximise the quality of the AI voice:
Record at least 1 minute of audio
Avoid recording more than 3 minutes, this will yield little improvement and can, in some cases, even be detrimental to the clone.
How the audio was recorded is more important than the total length (total runtime) of the samples. The number of samples you use doesn't matter; it is the total combined length (total runtime) that is the important part.
Approximately 1-2 minutes of clear audio without any reverb, artifacts, or background noise of any kind is recommended. When we speak of "audio or recording quality," we do not mean the codec, such as MP3 or WAV; we mean how the audio was captured. However, regarding audio codecs, using MP3 at 128 kbps and above is advised. Higher bitrates don't have a significant impact on the quality of the clone.
Keep the audio consistent
The AI will attempt to mimic everything it hears in the audio. This includes the speed of the person talking, the inflections, the accent, tonality, breathing pattern and strength, as well as noise and mouth clicks. Even noise and artefacts which can confuse it are factored in.
Ensure that the voice maintains a consistent tone throughout, with a consistent performance. Also, make sure that the audio quality of the voice remains consistent across all the samples. Even if you only use a single sample, ensure that it remains consistent throughout the full sample. Feeding the AI audio that is very dynamic, meaning wide fluctuations in pitch and volume, will yield less predictable results.
Replicate your performance
Another important thing to keep in mind is that the AI will try to replicate the performance of the voice you provide. If you talk in a slow, monotone voice without much emotion, that is what the AI will mimic. On the other hand, if you talk quickly with much emotion, that is what the AI will try to replicate.
It is crucial that the voice remains consistent throughout all the samples, not only in tone but also in performance. If there is too much variance, it might confuse the AI, leading to more varied output between generations.
Find a good balance for the volume
Find a good balance for the volume so the audio is neither too quiet nor too loud. The ideal would be between -23 dB and -18 dB RMS with a true peak of -3 dB.
How to prepare your voice training samples
When creating your custom voice you will need one or several training samples (Recordings). These should be recorded using the same equipment (microphone etc) and on the same set that the rest of the video is.
Must haves:
- Needs to have the same audio grading/mixing that the final video will have.
- Needs to have emotional range that reflects the rest of your video, can't be overly monotone.
Nice to haves:
- If the intention is to use AI-voice for name reading, the training data should aim to include a small set of actual name readings, e.g. 5-10 greetings. This helps increase consistency.
Training sample requirements
- Accepted file-types:
wav,mp3 - Minimum length (per sample):
10 seconds - Maximum length (per sample):
3 minutes
Assuring high quality voice generation
The quality of the generated audio is mainly dependent on two factors:
- The quality of the original training data
- The creative itself (e.g. how the generated audio is implemented in the script)
The more weaknesses there are in these two areas, the less life-like the end result will be. For ideal implementation especially the creative/script should be considered closely.
Language support
Audio generation is supported in multiple languages, although quality may vary across. Currently, we have the most confidence in English scripts.
List of supported languages
- Arabic (
ar) - Bulgarian (
bg) - Chinese (
zh) - Croatian (
hr) - Czech (
cs) - Danish (
da) - Dutch (
nl) - English (
en) - Filipino (
fil) - Finnish (
fi) - French (
fr) - German (
de) - Greek (
el) - Hungarian (
hu) - Hindi (
hi) - Indonesian (
id) - Italian (
it) - Japanese (
ja) - Korean (
ko) - Malay (
ms) - Norwegian (
no) - Polish (
pl) - Portuguese (
pt) - Romanian (
ro) - Russian (
ru) - Slovak (
sk) - Spanish (
es) - Swedish (
sv) - Tamil (
ta) - Turkish (
tr) - Ukrainian (
uk) - Vietnamese (
vi)
A note on languages and QA (Only applicable for managed projects)
Using scripts in languages that are foreign to Seen's CS-team will reduce their capability to perform quality assurance. In these cases you as a customer would be informed and invited to participate in the QA to ensure the best possible outcome.
Cloning the voice
Make sure to have an approval from the voice actor before cloning any voice.
Creating a custom voice inside of the platform
Navigate to AI Voices through the left sidebar.
Click Clone Voice in the top right hand corner.
Upload the recorded training samples you want to use to instruct the cloning. Multiple samples can be used for a single cloning. Read more on how to prepare your voice training samples above.

Add a name, description and tags to your new voice to make it easier to find in the future.

You can upload two different audio files with different tonality, for example "conversational" and "energised", and mix between them inside the same Studio project.
Designing the voice
If you are not able to record a voice, but still want a unique voice for your project, you can use our Design Voice tool.
Prompt the voice
Give it a name, select a language, and write a description following these guidelines:

- Be specific: Include details like age, tone, gender, and accent. For example, instead of "old woman," use "sarcastic old woman with a thick New York accent".
- Combine attributes: Use multiple traits together for better control (e.g., "A sarcastic old woman with a thick New York accent, speaking slowly").
- Use emotional tags: Incorporate emotional tags like
[excited]or[calm]directly into your text to influence delivery. - Avoid vague descriptors: Steer clear of imprecise terms like "foreign" or "exotic".
- Specify audio quality: You can ask for a specific quality, like "studio quality microphone" or "old phone".
Select one of the previews and press Create Voice.

Your voice is now ready to be used in your next Seen Studio project.
Using the voice
Use personalised audio as early in the video as possible for a great hook!
Your new voice can be found under Select Voice in your Text to Speech panel on the left side of your Seen Studio panel.

Select the correct language for your voice.

We have three different models supporting different languages:
- Swift: the fastest one
- Flexible: natural tone
- Max: highest quality

You can generate a new version of the same text until you are happy with it. To finish, press Add to Project.

Writing the text
Start with a test name, to understand the tone of voice and the length of your AI voice. Use both shorter and longer names, to understand how long your block needs to be. The full sentence will be generated for each variable, so you are only listening to one of the output examples.
Sentence structure
- Replace periods: For a more natural and conversational tone, replace periods at the end of sentences with exclamation marks.
- Add emphasis: For extra emphasis, you can add multiple exclamation marks (e.g.,
!!!). - Combine with capitalization: Use capitalization for words you want to emphasize, like
THIS. - Keep each sentence as short as possible, for example "Hi, {first_name}", "Hello, {first_name}" or "Welcome, {first_name}".
- Keep each personalised text as a full sentence, sentences that end with a period or exclamation mark work well.

Use it in the timeline
Keep enough room in the timeline for longer variables. "Hi, Victoria-Helena" takes longer to read than "Hi, Bob". Keeping enough room before/after a personalised audio will also make it sound more natural between the sentences.

Use properties
To finalize your project, go to the Personalisation panel and Add existing or create New properties for your project, for example first_name.

Change your test data with the correct property, generate and replace it in the timeline.

Your audio layer will now have a purple outline to indicate that you have personalised the layer.
