Neutral pacing, no tone shift, default voice. The line lands legibly and evenly, but without commitment.
Prosody first. Text second.
The rendering engine treats a script the way a director treats a table read - not as a stream of tokens to pronounce, but as a shape of pacing, breath and stress. That shift is what makes the output stop sounding read.
Your table is ready. Right this way.
Warm, calm, neutral urban
Deep, formal, broadcast
Bright, expressive, retail
Custom, consented voice
One script. Four passes.
The same 26 words below run through four rendering passes. Each pass changes one variable, so you can hear what each variable actually does.
"The city was quiet, the sort of quiet that follows a long rain. She turned the key, opened the door, and stepped inside."
The rendering engine inserts a soft breath before "She turned the key" and a beat between the two clauses. Same words, different rhythm.
"Long rain" and "stepped inside" carry weight. The voice does not stress every noun - it stresses the two that carry the moment.
The tone parameter moves from neutral to hushed. Rate drops by a fraction, pitch lowers, breath deepens. The line becomes a scene.
What the engine actually models. In plain language.
Prosody prediction
A separate model predicts pauses, stress and pitch contour from the script before any waveform is generated - so pacing and rhythm are decided at the sentence level, not per word.
Emphasis modeling
Rather than pronouncing every noun with equal weight, the engine promotes one or two words per clause. Which ones is a function of context, not brute frequency.
Timing metadata
Every rendered file ships with phoneme, word and sentence boundaries in JSON. Downstream systems like IVR menus or captions can slot in without re running the audio.
