Skip to main content

Decision

The voices DeckTalk can call are a closed set of adapters inside the engine. Each section of narration ends as audio plus a start and an end time for every word. The words come from the voice itself, and once alignment is built they may come from an aligner run over the audio instead. Everything a vendor is particular about, its settings, its key, its hosts, its pause markup, its bill and its file format, is declared by its adapter and read from there, so nothing above speech/ names a vendor.

The closed set

The engine ships two speech adapters and names each one in PROVIDERS in src/decktalk/speech/__init__.py:
  • elevenlabs, the cloud voice, over https.
  • dtsp, one HTTP client that talks to decktalk-voice, a separate local server that is not part of this repository and runs the local speech and alignment models. It needs no key and speaks to the address the machine names, a loopback address by default.
There are no entry points and no plugin loading. A project file names an adapter by its key in PROVIDERS and can never name code. A host that embeds the library may hand Machine.of its own table, which is how a test runs the real narrate against a voice that spends nothing, and that table is code the host wrote and imported itself. The Python protocol types are public, in the decktalk.speech module: the provider protocol, the request it receives and the speech it returns, and the factory. A host and a contributor therefore have a type to build against, and a new adapter still arrives as a change to this tree.

What an adapter declares

Each adapter declares, in one place:
  • Its settings table. [elevenlabs] and one table per adapter, an XConfig whose fields are declared with tune() and layered by tomlmap like every other table, with the adapter’s own bounds. Every vendor field lives there, stability, similarity_boost, style and speaker_boost among them.
  • Its model, set in its own table alone, so changing [voice] provider never sends one vendor’s model id to another.
  • Its key variable, such as ELEVENLABS_API_KEY, or none. doctor, the plan and every hint name the variable the adapter declares.
  • Its base URL, base_url in its own table, which is machine-scoped, so a project file that names a host is refused and only the machine decides where the key and the script go.
  • How it renders pauses, NATIVE or STITCHED, per model where a vendor’s models differ.
  • Its billing: per character, per second of audio, or free, with the rate in its own table.
  • Its take identity, the mapping of its own fields that enters the take digest.
  • Its output format and the file suffix a take is written under, so a take is named by what it holds.
[voice] keeps the three keys every voice has: provider, id and speed. Each adapter bounds speed to its own range. The protocol carries no cache_key. The digest is taken in one place above the boundary, from the take identity the adapter declares. An adapter is built from its own settings table, a source of secrets, a timeout and a retry count, and never from a project, a stage, a path or a build directory. It receives a request and answers with audio and, when it has them, timed words.

Pauses are data

A request carries pieces: each is a run of text and the pause after it in seconds. No request carries markup. A NATIVE adapter writes the pause in its own syntax, which for ElevenLabs is a <break time="0.7s" /> tag. A STITCHED adapter is sent one piece at a time, and DeckTalk joins the audio with measured silence and moves each piece’s word times by where that piece starts. The take digest is taken over the adapter’s canonical rendering of the pieces. ElevenLabs renders exactly the text it is sent now, so every ElevenLabs take digest stays byte for byte what it is, and tests/contract/test_take_hash.py holds the digests of really voiced films against tests/data/take_hash.json. check refuses a script with pauses on a model whose adapter declares no way to render them, so a pause is never dropped without a finding. ElevenLabs v3 and v4 do not read <break> tags, so they are refused on a script with pauses until their adapter renders those pauses STITCHED.

Credentials

Every credential is a Secret, and only the adapter’s header builder reveals it to send it. Every request names the secrets it carries, which is the secrets argument of post, and scrubbing works on their values, not on a list of header names: every such value, in whatever header, query parameter or body field it travelled, and the token after an Authorization scheme, is replaced in any reply body or error text before it is quoted. A redirect that leaves the request’s origin, its scheme, host and port together, drops every header that holds one of those values, whatever it is called, along with Authorization, Proxy-Authorization and Cookie, which are the second line for a credential no secret names.

Take identity and money

A take’s digest is the sha256 of six fields joined by newlines: the adapter’s name, the voice id, the model, the output format, the take identity as compact JSON with sorted keys, and the rendered text. Every field but the text refuses a newline, and the text is last, so no two sets of inputs share a payload and no existing digest moves. A paid request whose reply broke after it was sent is never sent again, and it is reported as possibly charged. A request is retried only when nothing connected, or on a 408, a 429 or a 5xx. A run either may spend or may not, which is spend: bool, and spend gates money and nothing else. The voice is built only when a take must be made, so a run that plays what is on disk reads no key. A free adapter never asks for approval and needs no --spend: it makes every missing take whether or not the run may spend, so --no-spend, the Action and --watch all voice with dtsp. A run that may not spend never calls an adapter that bills. It plays every take on disk and gives each missing section a placeholder and a TAKE_MISSING finding whose hint is --spend. A free adapter that nothing answers is a server that is not running, so it degrades the same way rather than failing the run: each section plays a placeholder, its TAKE_MISSING hint says to start decktalk-voice, and the run exits 0.

The voice id

The voice id is [voice] id in decktalk.toml, and DECKTALK_VOICE_ID in the environment overrides it for anyone who keeps it out of the file. The id is a published name, not a credential. It names which voice reads the script, the way a model name names which model answers, and it grants nothing on its own: a voice cloned on an ElevenLabs account is usable only with that account’s key. The take identity always carries the adapter’s name beside the id, because an id means something only to the vendor that issued it.

Sound is its own seam

The sound stage (score) buys music, ambience and effects through a SoundProvider, a second protocol with its own table of adapters beside PROVIDERS. It never borrows the speech registry and never checks which class a voice is. Sound is billed by the second of audio, because that is how sound is billed, at a rate in US dollars per minute in the sound adapter’s own table. A sound’s ledger digest is taken over the endpoint as the service publishes it and the request body, without the configured base_url, so moving to another host of the same service or to a local mock buys nothing again.

Alignment

Alignment is planned and not built yet. Every voice DeckTalk ships today returns its own word times, so nothing below runs in a build. What is decided is its shape. Alignment is a seam inside narrate, not a seventh stage. Words a voice returns with its audio are used as they come, and are kept beside the take. A voice that returns no word times is paired with an aligner, which reads the take’s audio and the script’s words and returns a time for every word. An aligner’s words are kept in a words cache of their own, keyed by the digest of the audio plus the aligner’s id and revision, so a better aligner re-aligns every take and buys none of them again. A voice without timestamps will therefore be a usable voice. DeckTalk will own the client side of /v1/align: the request and response types it sends to decktalk-voice and reads back, in decktalk.speech, with a fake server that proves the whole path in CI. No adapter that reads an author’s own recording is planned with it. This record states only the shape decided so far: the seam sits in narrate, the words have their own key, and the server is out of process.

Why

The cut is made on words, so word times are the whole requirement. Who supplies them is not. Forced alignment against a script DeckTalk already holds is accurate well inside the cue tolerance, so a voice with no timestamps stops being a different feature and becomes one more source of audio. The set is closed because what a build can call should be exactly what a reader of this tree can see. A host that runs projects it did not write must not offer a plugin surface, and an entry point is one. The loopback server is what keeps the set closed without keeping models out: torch, MLX, espeak and a model’s weights live in another process, behind one reviewed HTTP client. That process boundary is also a safety boundary. The espeak front end that local voices use can call exit() on a long data path, which inside DeckTalk’s process would end the build without a Python exception, and outside it ends only the server, which the adapter reports as a provider error. The protocol types are public because a closed set still needs a type to build against. A host embedding the library, a test and a contributor all write against the same protocol, and none of them can load code through a project. Every vendor-shaped fact sits on its adapter because each one, left elsewhere, is a bug waiting for a second vendor: a credential header not on a list leaks in an error, a vendor’s settings in every take digest re-voice a film for a field another vendor ignores, a pause tag is read aloud, a format string is sent to a service that does not know it, and a sound is priced at the speech rate. Declaring them once means check can refuse a mismatch before anything is bought, rather than a request failing after. Scrubbing by value is the design fix for credentials because a header list is only as good as the vendor its author knew. A value is the thing that must not leak, whatever header carried it. Take identity is where money is. A digest that moves buys every take again, so the ElevenLabs rendering is held byte for byte by a golden test, and the newline rule keeps the payload unambiguous without changing any digest that exists. The take digest leaves out the neighbouring text sent for prosody, so a take is kept when the section before it changes, which costs a little continuity and saves buying a take. The words have their own key because the audio does not depend on the aligner. Putting the aligner in the take digest would buy the audio again to re-align it. Alignment is not a stage because the six-stage rule makes a step a stage when a user runs it alone on files another step did not just write. Re-aligning is narrate finding the audio cached and the words missing, which it already runs alone, and cue still reads one artifact. Building an adapter from values rather than a Project keeps speech/ in the leaves layer, which the layering test holds, and lets a test build one from a few numbers and a source of secrets.

What it rules out

  • No entry points, no plugin discovery and no adapter named in a project file that is not in the table its machine answers with.
  • No vendor name above speech/. A layer test holds it.
  • No new cloud speech provider before the first public launch. The local server and the aligner supply the new voices.
  • No markup in a request, and no pause dropped without a finding.
  • No cache_key on the protocol, and no digest spelled outside the one place takes are named.
  • No retry of a paid request whose reply broke after sending.
  • Cloning is a separate command that records the speaker’s consent. A build never clones, so a build can never spend on a clone.
  • No licence gate. Each model entry in the local server’s registry carries a plain licence field, any model may be used locally, and nothing in DeckTalk filters on that field.
  • No adapter reaches the build directory or caches on its own. Caching is one rule above the boundary.

What would change it

A model that has to run in DeckTalk’s own process, with no way to serve it on loopback, would reopen how local voices load. A vendor whose only timing path is a websocket would add a transport to http.py. A source of phoneme or character times that cues should name directly would widen what a word is, which the timing note covers.