Clean

Paste one narration script - an audiobook chapter, an explainer voiceover, a podcast cold open, a radio advertisement, an e-learning module - and take it to the point of recording with a synthetic voice. Four lanes over the same script: prepare the read into synthesis blocks with every acronym, number, date and two-pronunciation word decided; cast the voices with an engine family and four settings each; write the sound-effect cues as generation prompts with durations and levels; and brief one music bed with sections, ducking and a real alternate. A free in-browser script reader parses the script first - spoken characters, duration, synthesis blocks and twenty-seven deterministic flags - and hands its findings to the model as facts it must reconcile one by one. Nothing is synthesised here; the output is a plan you execute in your own tool. Derived from four agent skills published by ElevenLabs in @elevenlabs/skills: @elevenlabs/text-to-speech, @elevenlabs/speech-engine, @elevenlabs/sound-effects and @elevenlabs/music. A derived work, not affiliated with or endorsed by ElevenLabs, and calling no ElevenLabs API.

Share

Details

PricingUsage-based + 10% creator margin
Billed model rate$2.75 in / $16.50 out per 1M tokens
Creator margin+10%
Effective rate$3.00 in / $18.00 out per 1M tokens
Security scanClean — skill and frontend scanned
Model gpt-terra
Created2026-08-19
Updated2026-08-21

Every public app is built from a security-scanned skill and must pass a clean scan — skill and frontend — before it can be listed. Have a skill of your own? Turn it into an app — or read the step-by-step walkthrough.