Frequently Asked Questions

Overview

What is TTS?

TTS stands for Text-to-Speech – a production method that transforms written text into spoken dialogue using a previously created voice.

What is STS?

STS stands for Speech-to-Speech – a production method that applies a different voice to an existing performance while preserving timing, rhythm, and expression. See our Speech-to-Speech guide for production use cases.

What is STT?

STT stands for Speech-to-Text. This feature converts spoken audio into written text, automatically recognizing the language used in the recording. STT is available exclusively in the ARA2 version of the plugin.

What are Voice Operations?

A voice operation refers to either creating or editing a voice or a voice design/remix. Each subscription includes a specific number of voice operations, which counts toward your monthly quota. You can find the exact number included in your subscription details.

What is Voice Synthesis?

Voice synthesis is a virtual reproduction of a real voice, created using AI technology. It mimics the tone, style, and character of the original speaker so naturally that it sounds like the real person is talking.

Each subscription includes a curated library of high-quality voices. In addition, every plan comes with a quota for custom synthetic voices that users can create themselves.

What is Voice Design?

Voice Design is a feature that allows users to create custom voices from scratch by describing them through text prompts. You can specify attributes such as age, gender, accent, tone and emotion. This enables you to generate entirely new, realistic voices tailored precisely to your needs.

The Voice Design feature is not available in the Voice Talent plan.

What is Voice Remixing?

Voice Remixing allows you to modify a synthesized or custom-created voice by entering a text prompt. You can specify attributes such as age, gender, accent, tone, and emotion, making it easy to adjust a voice to fit your exact needs.

Voice Remixing is available for voices in the VoiceWunder library. It is not supported for shared voices.

The Voice Remixing feature is not available in the Voice Talent plan.

All about Voice Sharing

What is Voice Sharing?
Voice Sharing allows voice talents to create licensed digital versions of their own voice and grant selected studios access. Studios can use these licensed voices for production, while licensing and ownership remain with the voice talent.

Who owns the voice?
The voice talent always remains the owner of their voice and their synthetic voices. Voice Sharing provides the technology. Ownership does not change.

Who controls access?
Only the voice talent can grant or revoke access. Studios cannot access a voice unless the talent has explicitly approved it.

Can I see when my voice is used?
Yes. Every use of your voice is logged and visible to you. You can see who used your voice, when it was used, and what was created with it.

Can my voice be used without my knowledge?
No. Studios can only use your voice after you grant access. All usage is visible in your activity log.

Can I revoke access?
Yes. You can revoke access at any time.

Who defines the license terms?
License terms, scope, and fees are agreed directly between the voice talent and the studio. Voice Sharing does not set license fees and does not take commissions.

Do I get paid when my voice is used?
Voice Sharing enables licensing between you and the studio. Compensation is agreed directly between both parties.

Is there a recommended pricing model?
Voice Sharing provides an example price matrix as a reference. You are free to define your own pricing.

Can my voice still be used for regular recordings?
Yes. Voice Sharing complements traditional voice production. It does not replace recording sessions.

Can someone use my voice without permission?
No. Voice Sharing is limited to voices created and shared by the voice talent. Studios cannot create or access a voice without permission.

Can Voice Sharing voices be resold?
No. Voice Sharing voices are licensed directly by the voice talent. Voice Sharing does not sell or sublicense voices.

How do I share my voice? 
Upload up to 20 different versions of your voice and grant access to selected studios – conveniently within a DAW or via a simple browser upload. You remain in control at every step.

How do studios get access?
Studios provide their VoiceWunder ID to the voice talent. The talent grants access.

Still have questions?
Contact our support team.

New to Voice Sharing?
See our
AI Voice Casting guide for an overview of how studios and voice talent work together.

The Voice Library

VoiceWunder offers a range of default voices in its Voice Library, tailored to each subscription plan.

Basic Plan: Includes a curated selection of professional-quality default voices.

Studio Plan: Includes all Basic voices plus an additional extended voice library featuring exceptionally high-quality voices. 

For both the Basic and Studio plans, commercial use of the default voices is permitted and royalty-free – even after the subscription has ended.

What factors should I keep in mind when developing a synthetic Voice?

For best results, use the highest possible audio quality – free from background noise. We recommend a loudness level of approximately -23 dB to -18 dB RMS, with a True Peak of -6 dB to prevent distortion. The recording should be spoken in a consistent tone and volume throughout. The more consistent the input, the more natural and coherent the voice output will sound. A recording duration of 1–3 minutes is usually sufficient.

Coins/Billing

Each subscription includes a specific number of coins, which are used for different forms of speech synthesis:

  • TTS (Text-to-Speech): 1 character = 1 coin
  • STS (Speech-to-Speech): 1 minute of audio = 1,000 coins, regardless of how much
    text it contains
  • STT (Speech-to-Text): 1 minute of audio = 1,000 coins, regardless of how much
    text it contains
  • RESCUE DIALOGUE: 1 minute of audio = 1,000 coins, regardless of how much text it contains

Unused coins expire at the end of the month, and the coin balance is refilled at the beginning of each month according to the subscription plan.

What happens after I use my 5,000 welcome coins? Your Community plan continues with the included monthly coin allocation.

Consent

Due to legal requirements, it is essential to obtain the explicit consent of the respective speaker before creating a synthetic voice. Use of the plug-in or platform is permitted only with material for which the user holds the appropriate rights or authorization.

Where required by applicable law or platform policies, AI-generated speech content should be appropriately disclosed or labeled by the deploying party. VoiceWunder supports available industry-standard provenance and transparency technologies where technically feasible. Downstream editing, transcoding and distribution workflows may alter or remove metadata outside the control of VoiceWunder.

Compliant with EU AI Act Article 50 (effective 2 August 2026). For audio content using generic voices, no disclaimer is required.

For a broader overview of what to check when evaluating AI voice vendors for compliance, see our AI Voice Compliance buyer’s guide.

Optimizing TTS results

If the voice output doesn’t meet your expectations, try the following adjustments:

  • Re-render the text: The voice often adjusts its emphasis slightly with each rendering.
  • Correct pronunciation: Spell words phonetically to match how they should sound.
  • Use SSML Phoneme Tags: For English, phonetic notation allows more precise control.
  • Adjust rhythm and emphasis: Add punctuation like “, . ; : – ! ?” to influence pauses and tone.
  • Use CAPITAL LETTERS: This can emphasize certain words or syllables.
  • If you experience issues with pronunciation or language usage when using real voice talents, we recommend enabling Safe Mode under Speech Generation in the settings for improved stability and consistency.
  • Use advanced mode and emotional tags. Emotional tags are inline cues placed in square brackets (e.g., [sigh], [excited]) within the text you want to synthesize. They guide how the voice speaks by adding emotional, non-verbal, or stylistic elements. To enable advanced mode, open the settings page and select Advanced Mode under Speech Generation. Once enabled, you can also choose emotional tags from a pop-up list by clicking the icon in the lower-right corner of the text box. You can also try specifying a language here if the voice has a particular accent or coloration.

Optimizing STS results

To achieve the best results with Speech-to-Speech (STS), consider the following guidelines:

  • Match vocal characteristics: The original speaker and the synthetic voice should sound similar in tone and style.
  • Mimic the synthetic voice: Have the speaker imitate the tone of the target voice for a more natural output.
  • Maintain proper audio levels: Both recordings should be between -23 dB and -18 dB RMS, with a True Peak no higher than -6 dB.
  • Clean the audio: Remove background noise, mouth clicks, and other artifacts before use.
  • Use consistent training data: Upload 1-3 minutes of speech spoken in a steady pitch and tone.
  • Avoid mixed sources: Do not use recordings with multiple speakers.

Optimizing Voice Design

If the voice you designed doesn’t sound quite right, try these tips to fine-tune the result:

  • Adjust Loudness: If the voice sounds distorted or overly intense, lower the Loudness value. This affects both the speaker’s tone and overall volume. Default: 50
  • Increase Prompt Strength: If the voice doesn’t match your prompt closely enough, raise the Prompt Strength value. This makes the system rely more heavily on your text description. Default: 0.5 – very high values may produce unintended results.
  • Regenerate Previews: If you’re not satisfied with the suggested voices, click “Create Previews” again to generate new variations.
  • Refine Your Prompt: Experiment with different character traits and emphasize key qualities using modifiers like “very,” “slightly,” “deep,” “warm,” or “gentle.”
  • Use Advanced Mode and Emotional Tags: You can add emotional or performance cues directly into the preview text using square-bracket tags such as [sigh], [excited], or [whispering]. These guide the voice’s emotion, delivery, and style. To enable this feature, go to Settings → Speech Generation → Advanced Mode. Once enabled, you can insert emotional tags using the icon in the lower-right corner of the preview text field.
  • Specify the Language: If the voice will speak a particular language, include it in your prompt so the voice is optimized for pronunciation and cadence.

Optimizing Voice Remixing

If the voice you remixed doesn’t sound quite right, try these tips to fine-tune the result:

  • Lower Loudness if the voice sounds distorted or too intense. This affects both tone and volume. Default: 50.
  • Increase Prompt Strength if the voice doesn’t match your description closely enough (Default: 0.5 – very high values may produce unintended results).
  • Click “Create Previews” again to generate new voice variations if you’re not satisfied.
  • Refine your prompt by adding clear character traits and emphasizing words like “very,” “warm,” “soft,” “deep,”or “energetic.”
  • Use Advanced Mode and emotional tags in the Preview Text field. Emotional tags are placed in square brackets (e.g., [sigh], [excited]) to guide emotion, delivery, and style. Enable this by going to Settings → Speech Generation → Advanced Mode, then use the emotion tag icon in the preview field.
  • Specify the language if you want the voice optimized for a particular language.

Can I install the plug-in on multiple workstations?

Yes. With a paid plan, you can install the plug-in on multiple workstations (Basic: up to 3, Studio: up to 5). However, the plug-in can only be active on one workstation at a time. To switch devices, simply close the plug-in window on your current workstation before opening it on another. See our Pro Tools guide for the full workflow.

Commercial usage

All voice output generated using voice design/remix or the provided voice library can be used for your professional and commercial end-use projects (e.g., videos, films, games, podcasts, advertisements) if created under a Basic or Studio subscription. This usage remains valid even after the subscription ends.

However, the license is strictly limited to end-use only.

It is strictly prohibited to:

  • Use any VoiceWunder outputs (audio files, synthetic voices, voice models, embeddings, or any derived data) for training, fine-tuning, improving, or developing any artificial intelligence, machine learning, or speech synthesis models (including but not limited to TTS, STS, voice cloning, or any generative audio systems).
  • Extract, reverse-engineer, or derive voice features, embeddings, or training data from the outputs.
  • Distribute, share, or make available any outputs in a manner that enables third-party AI training.

For custom voices or shared voices, additional restrictions from your agreement with the speaker apply.

Please note: Commercial use of any material generated under a Free subscription is strictly prohibited.

Can I use Community voices commercially? No. Commercial use requires a paid VoiceWunder plan.

Prohibited use

Any use that violates ethical standards or applicable laws is strictly prohibited. This includes, but is not limited to:

  • Using any generated outputs (audio, voices, models) for training, fine-tuning, or improving any AI/ML models or systems
  • Inciting violence, discrimination, or criminal behavior
  • Promoting drug use
  • Supporting fraudulent, exploitative, or abusive practices

Troubleshooting OMFs

Pro Tools often has issues accessing audio files within an embedded OMF via ARA. When importing the OMF, please select “Copy from source media” under Audio Media Options. This will copy the audio files from the OMF into the session’s Audio Files folder.

System requirements

The plug-in is compatible with any internet-enabled Mac or Windows PC running Adobe Premiere Pro (version 26.2), Avid Pro Tools (version 2025.6), Steinberg Nuendo/Cubase (version 14), Presonus Studio One Pro (version 7) or Cockos Reaper (7.43). A high-speed internet connection is strongly recommended, as audio transmission may involve large data volumes.

Ending a Subscription

You can easily cancel your subscription at https://account.voicewunder.ai/ in the user area under “Subscription|Manage Subscription….”

What happens to my voices after I cancel my subscription? Voices created under paid plans remain available after your subscription ends. If they remain inactive for an extended period, VoiceWunder will notify you before any permanent deletion.

What happens if I stop using my Community account? Community voices may be removed after extended periods of inactivity to make voice resources available for active users. Your Community account remains available, and you can create a new voice at any time.

Security

Your personal data is used exclusively for generating the requested speech synthesis and is never utilized for model training purposes. All data is transmitted securely using SSL encryption.
We offer European data residency and adhere fully to GDPR and CCPA regulations. For your security, our team will never request your password.

Which languages are supported?

Currently up to 74 languages are supported: Afrikaans, Arabic, Armenian, Assamese, Azerbaijani, Belarusian, Bengali, Bosnian, Bulgarian, Catalan, Cebuano, Chichewa, Croatian, Czech, Danish, Dutch, English, Estonian, Filipino, Finnish, French, Galician, Georgian, German, Greek, Gujarati, Hausa, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Irish, Italian, Japanese, Javanese, Kannada, Kazakh, Kirghiz, Korean, Latvian, Lingala, Lithuanian, Luxembourgish, Macedonian, Malay, Malayalam, Mandarin Chinese, Marathi, Nepali, Norwegian, Pashto, Persian, Polish, Portuguese, Punjabi, Romanian, Russian, Serbian, Sindhi, Slovak, Slovenian, Somali, Spanish, Swahili, Swedish, Tamil, Telugu, Thai, Turkish, Ukrainian, Urdu, Vietnamese, Welsh.

Who issues my invoice?

If you subscribed before 10.02.2026, your invoices are issued by Link, LLC f/k/a Lemon Squeezy LLC.
If you subscribed after 10.02.2026, invoices are issued by VoiceWunder GmbH.

© 2026 VoiceWunder® GmbH · All rights reserved.

All prices are net prices and may be subject to applicable taxes. Professional use only.