Studio/Text to speech generation

Prosody first. Text second.

The rendering engine treats a script the way a director treats a table read - not as a stream of tokens to pronounce, but as a shape of pacing, breath and stress. That shift is what makes the output stop sounding read.

Script

Your table is ready. Right this way.

Aanya

Warm, calm, neutral urban

EN-IN4.2s
Rohan

Deep, formal, broadcast

HI-IN3.8s
Meera

Bright, expressive, retail

EN-IN4.0s
Your clone

Custom, consented voice

EN-IN4.5s
A worked example

One script. Four passes.

The same 26 words below run through four rendering passes. Each pass changes one variable, so you can hear what each variable actually does.

Script

"The city was quiet, the sort of quiet that follows a long rain. She turned the key, opened the door, and stepped inside."

Pass 01 - Baseline

Neutral pacing, no tone shift, default voice. The line lands legibly and evenly, but without commitment.

Pass 02 - Breath and pause

The rendering engine inserts a soft breath before "She turned the key" and a beat between the two clauses. Same words, different rhythm.

Pass 03 - Emphasis modeling

"Long rain" and "stepped inside" carry weight. The voice does not stress every noun - it stresses the two that carry the moment.

Pass 04 - Tone shift

The tone parameter moves from neutral to hushed. Rate drops by a fraction, pitch lowers, breath deepens. The line becomes a scene.

Under the hood

What the engine actually models. In plain language.

Prosody prediction

A separate model predicts pauses, stress and pitch contour from the script before any waveform is generated - so pacing and rhythm are decided at the sentence level, not per word.

Emphasis modeling

Rather than pronouncing every noun with equal weight, the engine promotes one or two words per clause. Which ones is a function of context, not brute frequency.

Timing metadata

Every rendered file ships with phoneme, word and sentence boundaries in JSON. Downstream systems like IVR menus or captions can slot in without re running the audio.