Voice picks — round 7

Three brands, three complaints, one shared cause — and it was never the voice. In edge-tts every full stop buys a flat ~1.05 seconds of dead air. The Docket was spending 30% of its runtime in silence. Grit & Glory's copy had been deliberately chopped into six short fragments, which bought six identical pauses — a metronome. That is what you heard as "robotic" and "weird pauses". Fixed by rewriting the copy and re-measuring, not by swapping voices again. And there is a bigger answer at the top of this page.

The bigger answer — a free hosted engine, working today

Chatterbox vs. what we have YOUR CALL
Resemble AI · MIT licence · free hosted API · no account, no key, no install, no spend
You asked whether we could connect to something online. We can, and I proved it — it is already running. Same Desmond Doss copy, same processing, only the engine changed:
EngineSilence gaps ≥0.25sAudio
edge-tts (Guy) — today'sfour, ~1.05s each48 kbps lossy MP3
Chatterboxzero16-bit PCM, 384 kbps
Not fewer gaps. None. It does not stitch sentences together with a fixed pause, so the mechanical dead air simply does not exist. Whisper confirms every word is spoken. This is the engine blind-tested 65.3% preferred over ElevenLabs — the one you were leaning toward.
A · edge-tts, our current best (20.4s)
B · Chatterbox, same words (10.8s)
The honest catch: we got one render through and then the free tier cut us off. A plain one-line test fails too, so it is a per-address GPU quota, not our text. It is fine for auditions; it is not reliable enough to hard-wire into the nightly loop as-is. Two ways forward, both free, both yours to decide: (1) a free Hugging Face account raises that quota — no card, but creating accounts is your call, not mine; (2) install it locally (MIT, runs on the CPU) so we own it and never depend on someone else's server. I would do both: local for production, hosted for quick auditions.

Needs your ear — round 7

Grit & Glory ROUND 7
en-US-GuyNeural · rate −6% · pitch −2Hz · copy rewritten
The pauses were the script, not the voice. Round 5 chopped the copy into six short fragments believing that created trailer air. It bought six identical ~1.05s gaps instead — 1.06 / 1.05 / 1.11 / 1.11 / 1.03 / 0.97. Rate was proven innocent: −10% to −4% moved the gap by 0.06s. Rewritten as four long declaratives with no commas, the gaps now fall at 4.75 / 12.07 / 14.50 — irregular, with the last one landing directly in front of the Medal of Honor line.
The Docket ROUND 7
en-US-AndrewMultilingualNeural · rate −2% · pitch −2Hz
"Robotic" was dead air, not flatness — Roger actually had the widest pitch range of every voice tested. He was also pausing 0.77s after every single sentence: 30% of the runtime was silence. Andrew cuts that to 0.42s and finishes the same script six seconds faster. Guy was measured as a literal metronome (1.0s gaps, 3.9% variance) and is out. One trade, stated plainly: Andrew is 2.9 dB darker than Roger in the presence band, and that is built into the voice — I tested rate changes and they bought back 0.3 dB. You asked for warmth, so I took warmth knowingly. If this reads as "whispered" again, the fix has to be in the shared audio chain, not the voice. All six legal terms verified; "precedent" lands clean twice; no lexicon changes.
Sleep Audio (Dusk) ROUND 7
en-US-AvaMultilingualNeural · rate −16% · pitch −2Hz
You were right, and the paperwork was wrong. My theory was that the de-esser in our audio chain was eating the /s/ sounds — measured, it moves sibilance by 0.0 dB. Wrong theory. The real finding: the sample you heard was not rendered at the settings round 6 claimed. Duration proves it — 14.21s, but −16% renders at 13.10s and −22% at 14.11s. It went out at roughly −22%. Two words were genuinely garbled ("brown noise" → "ground noise") and both render clean at every rate when redone. Ava measured the brightest /s/ of the field (−0.5 vs Brian −2.8); Andrew was the dullest, so that swap would have made it worse. Brian control below — if you prefer him, say so.
Control · Brian, same line
Penloom History YOUR QUESTION, ANSWERED
current pick: en-US-EricNeural · rate −6% · pitch −2Hz
No — we had not studied your sixteen actors. Only Jeremy Irons overlapped; the research had gone to documentary narrators instead. Now done, and your list splits three ways: classically-trained RP (McKellen, Stewart, Hopkins, Irons), low warm gravel (Neeson, Elba, Caine, Connery), and light/young (Holland, Garfield, Hiddleston, Craig, Pegg) which is wrong for dark history. Your brand wants the first group. So I measured every British voice we have against it: Eric sits at 107 Hz with a full low end; the best British option, Ryan, is 121 Hz and noticeably thinner; Thomas is thinner still. Microsoft's British males are read-aloud assistants, not stage-trained instruments — and I will not tell you Ryan is Ian McKellen. Hear him yourself below. That gravitas is exactly what the engine question above is about.
The best British we can actually get · en-GB-RyanNeural

Approved by you — locked

Faith & Founding APPROVED
en-US-BrianMultilingualNeural · rate −6% · pitch −2Hz
Meme Vault APPROVED
en-US-BrianMultilingualNeural · comedic timing
Penloom Works APPROVED
en-US-BrianMultilingualNeural · rate −2% · pitch −3Hz
Cogloom APPROVED
en-US-AndrewMultilingualNeural · rate −4% · pitch −3Hz
Liberty Line APPROVED
en-US-AndrewMultilingualNeural · rate −4% · pitch −6Hz
Pet Legacy APPROVED
en-US-BrianMultilingualNeural · rate −6% · pitch −2Hz
SWFL Happy Hour APPROVED
en-US-AndrewMultilingualNeural · rate +4% · pitch −4Hz

One more thing, and it affects every finished video

Chasing the pauses turned up something bigger than the samples. Our renderer (cinematic.mjs) inserts its own 0.8-second silence block per beat — on top of the ~1.05s the full stop already bought. That is roughly 1.9 seconds of dead air per reveal in the actual published videos, not just these auditions. Every brand using that renderer has been paying it. I have left it alone: it changes how finished videos time out, which is a bigger decision than a voice pick. Say the word and I will cut it to ~0.35s across the board.