Troubleshooting Audio Generation Issues and Credit Consumption
Last updated: September 2, 2026
If your audio generation has produced problematic output — such as inaudible audio, skipped words, cutoff sentences, incorrect voice assignments, or mispronunciations — this guide will help you identify the root cause and resolve it, while minimizing unnecessary credit usage.
Understanding Credit Usage and Free Re-generations
Each generated clip includes 2 free re-generations at no additional credit cost. This allows you to re-roll a bad take twice before credits are consumed. Always use your free re-generations first when you encounter a problematic output before making script edits or switching modes.
Note that switching between Convo Mode and Standard Mode after generating audio requires a full regeneration in the new mode, which will consume additional credits.
Voice Compatibility: Convo Mode vs. Standard Mode
Many audio quality issues stem from using incompatible voice types with a given generation mode.
Convo Mode is built on Google Gemini's multi-speaker audio engine and works natively with Google Gemini TTS voices.
ElevenLabs voices (including cloned voices) are best used in Standard Mode. When used in Convo Mode, they may sound distorted, lose their accent and tone, cause generation failures, or switch unexpectedly during generation.
To find compatible voices quickly, use the voice search bar:
Search
Google Gemini TTSto find voices optimized for Convo Mode.Search
ElevenLabsto find voices optimized for Standard Mode.
Regional Accent Issues in Convo Mode
Convo Mode defaults to a standard American English accent regardless of the voice selected. This means regional accents — such as British, Quebec French, or other non-American accents — will not be preserved automatically.
Solutions
Use Standard Mode for guaranteed accent preservation. Standard Mode respects the native accent of the selected voice.
Add Delivery Instructions in the Delivery Instructions box at the top center of the screen to guide accent behavior in Convo Mode.
Use in-line prompts at the beginning of segments, such as
[speak with a Quebec accent]or[speak with a British accent].Use a detailed prompt template in the Delivery Instructions box. Here is a tested example for an academic podcast with British accents — adapt it to your content type and requirements:
Scene: Academic Podcast Episode
This is a professional two-person academic podcast where two speakers engage in a rich, in-depth conversation. Both speakers are very knowledgeable and experienced in their field. They explore complex ideas, challenge each other with thoughtful questions, and explain concepts clearly but without oversimplification.Director's Notes:
Speaker 1:
Style: speaks with a confident, warm, and articulate academic tone. Slightly reflective and curious.
Pace: balanced and steady, neither rushed nor overly slow. Natural pauses at sentence ends.
Accent: speaks in a very strong British accent.Speaker 2:
Style: speaks with a thoughtful, analytical, and slightly more measured tone. A touch of enthusiasm when explaining key points.
Pace: speaks at a moderate, deliberate pace.
Accent: speaks with a British English accent, from the London region.
Audio Cutoff and Compression Issues
Convo Mode can occasionally cut off audio at the end of sentences, particularly words ending in 's' or proper names at sentence boundaries. This is a known limitation of the underlying Gemini audio engine that is actively being worked on.
Solutions
Use your 2 free re-generations first before making any script changes.
Make a small edit to the affected paragraph — such as adding a space, changing punctuation, or capitalizing a letter — to mark it for regeneration. Then click the orange Generate changes button to regenerate only that section.
Modify your script to avoid sentences ending with 's' sounds where possible.
Mispronunciation and Pronunciation Inconsistency
Voices can be inconsistent with pronunciation across different generations, even with identical input. Specific words, dates, or phrases may be mispronounced repeatedly.
Solutions
Edit script to the way you want it to pronounce( e.g.
July 23→July twenty-third)Use the Pronunciation Settings feature for persistent issues. Go to Settings → Pronunciation, add the problematic word under 'Word' and the correct pronunciation under 'Pronunciation' (e.g.
IKEA→ee-KAY-uh). Save and regenerate. This rule applies across your entire project.Make small edits to the problematic segment — such as adding or removing commas or changing punctuation — then regenerate.
Use capital letters for difficult words (e.g.
PORTUGUESE) to improve pronunciation accuracy.
Incomplete Sentence Generation
Sometimes audio generation may start a few words into a sentence rather than at the beginning, resulting in incomplete or nonsensical output.
Solutions
Make small edits to the affected segment (add a comma, period, or adjust punctuation).
Speaker Assignment and Voice Switching Issues
In Convo Mode, voices may switch between speakers or the wrong voice may bleed in at speaker transitions. This is a known limitation of the Google Gemini voice engine (currently in beta), which may occasionally disregard speaker assignment instructions. The AI assistant cannot detect or fix these issues since it does not hear the generated audio.
Solutions
Use your free re-generations to re-roll the affected clip — speaker assignments may resolve correctly on a subsequent attempt.
Make script changes manually rather than asking the AI assistant to fix them, as manual edits are more reliable.
If speaker assignment issues persist, consider switching to Standard Mode, which offers more predictable voice assignment behavior.
Script Formatting Issues
Poorly formatted scripts are a common cause of garbled audio, random sounds, skipped lines, or voices reading unintended text. Review the following before generating audio:
Stray Characters After Punctuation
Quotation marks or other special characters placed after punctuation marks (e.g., after a question mark or period) can cause unintelligible sounds or garbled audio at the end of a speaker's output. This can also occur after editing or deleting content from your script.
Fix: Carefully inspect your script for any quotation marks or special characters that appear after punctuation marks. Remove them and regenerate the affected clip.
Speaker Labels Embedded in Dialogue Text
When using the AI assistant to generate a script, speaker labels (e.g., Ricky: or Diane:) may sometimes be left inside the dialogue text itself, causing the voice to read the name aloud.
Fix: Review your script before generating. Edit any clip text directly in the timeline to remove embedded speaker labels before clicking Generate all or Generate convo.
Mixed-Language Content
Embedding a single foreign-language word mid-sentence within otherwise same-language text can cause the audio pipeline to fail matching the generated speech to the word, causing the entire generation turn to fail.
Fix: Separate mixed-language content into individual lines or segments rather than embedding foreign words inline within a sentence.
General Script Review Tips
Always review your full script before clicking Generate all to avoid spending credits on an incorrect generation.
Edit clip text directly in the timeline to fix errors before generating.
Make script corrections manually rather than relying on the AI assistant for formatting fixes.
Regeneration Skipping Lines
When using Generate changes, random lines that were not marked for regeneration may occasionally be skipped or affected. If this occurs, identify the specific skipped lines, make a small edit to mark them, and regenerate those sections individually.
Summary: Quick Reference
Issue | First Step | Recommended Fix |
Distorted/failed audio with ElevenLabs voice | Check generation mode | Switch to Standard Mode; use Google Gemini TTS voices for Convo Mode |
Wrong accent in Convo Mode | Check mode setting | Use Standard Mode, or add Delivery Instructions / in-line accent prompts |
Audio cut off at end of sentences | Use 2 free re-generations | Make small script edit to affected paragraph; use Generate changes |
Mispronounced words | Try small punctuation edits | Use Settings → Pronunciation for persistent issues |
Incomplete sentence audio | Make small edit to segment | Regenerate only the affected clip |
Wrong speaker voice assignment | Use free re-generations | Make manual script edits; consider switching to Standard Mode |
Garbled audio / random sounds | Inspect script formatting | Remove stray characters after punctuation; regenerate clip |
Voice reads speaker label aloud | Review script in timeline | Edit clip text to remove embedded speaker labels before generating |