> ## Documentation Index
> Fetch the complete documentation index at: https://docs.decktalk.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Providers, local voices and forced alignment

> The closed set of voices, what each one declares, and where word times come from.

## Decision

The voices DeckTalk can call are a closed set of adapters inside the engine. Each section of
narration ends as audio plus a start and an end time for every word. The words come from the voice
itself, and once alignment is built they may come from an aligner run over the audio instead.
Everything a vendor is particular about, its settings, its key, its hosts, its pause markup, its
bill and its file format, is declared by its adapter and read from there, so nothing above `speech/`
names a vendor.

### The closed set

The engine ships two speech adapters and names each one in `PROVIDERS` in
`src/decktalk/speech/__init__.py`:

* `elevenlabs`, the cloud voice, over https.
* `dtsp`, one HTTP client that talks to `decktalk-voice`, a separate local server that is not part
  of this repository and runs the local speech and alignment models. It needs no key and speaks
  to the address the machine names, a loopback address by default.

There are no entry points and no plugin loading. A project file names an adapter by its key in
`PROVIDERS` and can never name code. A host that embeds the library may hand `Machine.of` its own
table, which is how a test runs the real `narrate` against a voice that spends nothing, and that
table is code the host wrote and imported itself.

The Python protocol types are public, in the `decktalk.speech` module: the provider protocol, the
request it receives and the speech it returns, and the factory. A host and a contributor therefore
have a type to build against, and a new adapter still arrives as a change to this tree.

### What an adapter declares

Each adapter declares, in one place:

* **Its settings table.** `[elevenlabs]` and one table per adapter, an `XConfig` whose fields are
  declared with `tune()` and layered by `tomlmap` like every other table, with the adapter's own
  bounds. Every vendor field lives there, `stability`, `similarity_boost`, `style` and
  `speaker_boost` among them.
* **Its model**, set in its own table alone, so changing `[voice] provider` never sends one
  vendor's model id to another.
* **Its key variable**, such as `ELEVENLABS_API_KEY`, or none. `doctor`, the plan and every hint
  name the variable the adapter declares.
* **Its base URL**, `base_url` in its own table, which is machine-scoped, so a project file that
  names a host is refused and only the machine decides where the key and the script go.
* **How it renders pauses**, NATIVE or STITCHED, per model where a vendor's models differ.
* **Its billing**: per character, per second of audio, or free, with the rate in its own table.
* **Its take identity**, the mapping of its own fields that enters the take digest.
* **Its output format and the file suffix** a take is written under, so a take is named by what it
  holds.

`[voice]` keeps the three keys every voice has: `provider`, `id` and `speed`. Each adapter
bounds `speed` to its own range. The protocol carries no `cache_key`. The digest is taken in one
place above the boundary, from the take identity the adapter declares.

An adapter is built from its own settings table, a source of secrets, a timeout and a retry count,
and never from a project, a stage, a path or a build directory. It receives a request and answers
with audio and, when it has them, timed words.

### Pauses are data

A request carries pieces: each is a run of text and the pause after it in seconds. No request
carries markup. A NATIVE adapter writes the pause in its own syntax, which for ElevenLabs is a
`<break time="0.7s" />` tag. A STITCHED adapter is sent one piece at a time, and DeckTalk joins the
audio with measured silence and moves each piece's word times by where that piece starts.

The take digest is taken over the adapter's canonical rendering of the pieces. ElevenLabs renders
exactly the text it is sent now, so every ElevenLabs take digest stays byte for byte what it is, and
`tests/contract/test_take_hash.py` holds the digests of really voiced films against
`tests/data/take_hash.json`.

`check` refuses a script with pauses on a model whose adapter declares no way to render them, so a
pause is never dropped without a finding. ElevenLabs v3 and v4 do not read `<break>` tags, so they are
refused on a script with pauses until their adapter renders those pauses STITCHED.

### Credentials

Every credential is a `Secret`, and only the adapter's header builder reveals it to send it. Every
request names the secrets it carries, which is the `secrets` argument of `post`, and scrubbing works
on their values, not on a list of header names: every such value, in whatever header, query parameter
or body field it travelled, and the token after an `Authorization` scheme, is replaced in any reply
body or error text before it is quoted. A redirect that leaves the request's origin, its scheme, host
and port together, drops every header that holds one of those values, whatever it is called, along
with `Authorization`, `Proxy-Authorization` and `Cookie`, which are the second line for a credential
no secret names.

### Take identity and money

A take's digest is the sha256 of six fields joined by newlines: the adapter's name, the voice id,
the model, the output format, the take identity as compact JSON with sorted keys, and the rendered
text. Every field but the text refuses a newline, and the text is last, so no two sets of inputs
share a payload and no existing digest moves.

A paid request whose reply broke after it was sent is never sent again, and it is reported as
possibly charged. A request is retried only when nothing connected, or on a 408, a 429 or a 5xx.

A run either may spend or may not, which is `spend: bool`, and spend gates money and nothing else.
The voice is built only when a take must be made, so a run that plays what is on disk reads no key.
A free adapter never asks for approval and needs no `--spend`: it makes every missing take whether or
not the run may spend, so `--no-spend`, the Action and `--watch` all voice with `dtsp`. A run that
may not spend never calls an adapter that bills. It plays every take on disk and gives each missing
section a placeholder and a `TAKE_MISSING` finding whose hint is `--spend`. A free adapter that
nothing answers is a server that is not running, so it degrades the same way rather than failing the
run: each section plays a placeholder, its `TAKE_MISSING` hint says to start `decktalk-voice`, and
the run exits 0.

### The voice id

The voice id is `[voice] id` in `decktalk.toml`, and `DECKTALK_VOICE_ID` in the environment
overrides it for anyone who keeps it out of the file. The id is a published name, not a credential.
It names which voice reads the script, the way a model name names which model answers, and it grants
nothing on its own: a voice cloned on an ElevenLabs account is usable only with that account's key.
The take identity always carries the adapter's name beside the id, because an id means something
only to the vendor that issued it.

### Sound is its own seam

The sound stage (`score`) buys music, ambience and effects through a `SoundProvider`, a second
protocol with its own table of adapters beside `PROVIDERS`. It never borrows the speech registry and
never checks which class a voice is. Sound is billed by the second of audio, because that is how
sound is billed, at a rate in US dollars per minute in the sound adapter's own table. A sound's
ledger digest is taken over the endpoint as the service publishes it and the request body, without the configured `base_url`, so
moving to another host of the same service or to a local mock buys nothing again.

### Alignment

Alignment is planned and not built yet. Every voice DeckTalk ships today returns its own word times,
so nothing below runs in a build. What is decided is its shape.

Alignment is a seam inside `narrate`, not a seventh stage. Words a voice returns with its audio are
used as they come, and are kept beside the take. A voice that returns no word times is paired with
an aligner, which reads the take's audio and the script's words and returns a time for every word.
An aligner's words are kept in a words cache of their own, keyed by the digest of the audio plus the
aligner's id and revision, so a better aligner re-aligns every take and buys none of them again.

A voice without timestamps will therefore be a usable voice.

DeckTalk will own the client side of `/v1/align`: the request and response types it sends to
`decktalk-voice` and reads back, in `decktalk.speech`, with a fake server that proves the whole path
in CI. No adapter that reads an author's own recording is planned with it. This record states only
the shape decided so far: the seam sits in `narrate`, the words have their own key, and the server is
out of process.

## Why

The cut is made on words, so word times are the whole requirement. Who supplies them is not.
Forced alignment against a script DeckTalk already holds is accurate well inside the cue tolerance,
so a voice with no timestamps stops being a different feature and becomes one more source of audio.

The set is closed because what a build can call should be exactly what a reader of this tree can
see. A host that runs projects it did not write must not offer a plugin surface, and an entry point
is one. The loopback server is what keeps the set closed without keeping models out: torch, MLX,
espeak and a model's weights live in another process, behind one reviewed HTTP client. That process
boundary is also a safety boundary. The espeak front end that local voices use can call `exit()` on a
long data path, which inside DeckTalk's process would end the build without a Python exception, and
outside it ends only the server, which the adapter reports as a provider error.

The protocol types are public because a closed set still needs a type to build against. A host
embedding the library, a test and a contributor all write against the same protocol, and none of them
can load code through a project.

Every vendor-shaped fact sits on its adapter because each one, left elsewhere, is a bug waiting for a
second vendor: a credential header not on a list leaks in an error, a vendor's settings in every take
digest re-voice a film for a field another vendor ignores, a pause tag is read aloud, a format string is
sent to a service that does not know it, and a sound is priced at the speech rate. Declaring them
once means `check` can refuse a mismatch before anything is bought, rather than a request failing
after.

Scrubbing by value is the design fix for credentials because a header list is only as good as the
vendor its author knew. A value is the thing that must not leak, whatever header carried it.

Take identity is where money is. A digest that moves buys every take again, so the ElevenLabs
rendering is held byte for byte by a golden test, and the newline rule keeps the payload unambiguous
without changing any digest that exists. The take digest leaves out the neighbouring text sent for
prosody, so a take is kept when the section before it changes, which costs a little continuity and
saves buying a take.

The words have their own key because the audio does not depend on the aligner. Putting the aligner in
the take digest would buy the audio again to re-align it.

Alignment is not a stage because the six-stage rule makes a step a stage when a user runs it alone on
files another step did not just write. Re-aligning is `narrate` finding the audio cached and the words
missing, which it already runs alone, and `cue` still reads one artifact.

Building an adapter from values rather than a `Project` keeps `speech/` in the leaves layer, which the
layering test holds, and lets a test build one from a few numbers and a source of secrets.

## What it rules out

* No entry points, no plugin discovery and no adapter named in a project file that is not in the
  table its machine answers with.
* No vendor name above `speech/`. A layer test holds it.
* No new cloud speech provider before the first public launch. The local server and the aligner
  supply the new voices.
* No markup in a request, and no pause dropped without a finding.
* No `cache_key` on the protocol, and no digest spelled outside the one place takes are named.
* No retry of a paid request whose reply broke after sending.
* Cloning is a separate command that records the speaker's consent. A build never clones, so a build
  can never spend on a clone.
* No licence gate. Each model entry in the local server's registry carries a plain licence field, any
  model may be used locally, and nothing in DeckTalk filters on that field.
* No adapter reaches the build directory or caches on its own. Caching is one rule above the
  boundary.

## What would change it

A model that has to run in DeckTalk's own process, with no way to serve it on loopback, would reopen
how local voices load. A vendor whose only timing path is a websocket would add a transport to
`http.py`. A source of phoneme or character times that cues should name directly would widen what a
word is, which [the timing note](/decisions/timing-from-the-spoken-words) covers.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.