> ## Documentation Index
> Fetch the complete documentation index at: https://docs.decktalk.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Timing comes from the spoken words

> Every time in the system is derived from word timestamps, so a cue names a phrase and never a second.

## Decision

Every time DeckTalk computes is derived from the start and end times of spoken words. A cue names a
phrase from the narration, never a second. Every take ends with a time for every word, supplied by
the voice itself or by an aligner run over the take's audio against the script, and that list is the
only clock in the system.

## Why

The thing that makes a narrated video expensive to change is that a picture is pinned to a second.
Rewrite one sentence and every picture after it is wrong, so the fix is to cut the whole video again
by hand. If a picture is pinned to a phrase instead, the same rewrite moves the pictures with it and
nobody touches a timeline.

This also makes the video checkable. `verify` knows which word a reveal was supposed to follow, so a
reveal that arrives late is a measured number carrying the code `CUE_OFF` rather than something a
person has to notice.

## What it rules out

* A voice with no word timestamps never supplies a clock of its own. An aligner times its audio
  against the script, so the clock is still the words, and
  [the provider note](/decisions/the-provider-interface) says where each source of word times sits.
* A cue cannot be placed between two words that are never spoken, so a picture that belongs to no
  phrase has to earn a phrase in the script.
* A section with no narration resolves no cues. The `cue` stage reports each one as `CUE_UNRESOLVED`
  rather than writing a cue list the recorder would then play blind.
* No attribute on a deck page says when a cue fires. The page owns what a thing looks like and what it
  means, and `cues.json` owns when it happens.

## What would change it

A voice or an aligner that returns phoneme or character timings rather than word timings would
widen what a cue may name. Nothing else would.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.