Skip to main content

Decision

Every time DeckTalk computes is derived from the start and end times of spoken words. A cue names a phrase from the narration, never a second. Every take ends with a time for every word, supplied by the voice itself or by an aligner run over the take’s audio against the script, and that list is the only clock in the system.

Why

The thing that makes a narrated video expensive to change is that a picture is pinned to a second. Rewrite one sentence and every picture after it is wrong, so the fix is to cut the whole video again by hand. If a picture is pinned to a phrase instead, the same rewrite moves the pictures with it and nobody touches a timeline. This also makes the video checkable. verify knows which word a reveal was supposed to follow, so a reveal that arrives late is a measured number carrying the code CUE_OFF rather than something a person has to notice.

What it rules out

  • A voice with no word timestamps never supplies a clock of its own. An aligner times its audio against the script, so the clock is still the words, and the provider note says where each source of word times sits.
  • A cue cannot be placed between two words that are never spoken, so a picture that belongs to no phrase has to earn a phrase in the script.
  • A section with no narration resolves no cues. The cue stage reports each one as CUE_UNRESOLVED rather than writing a cue list the recorder would then play blind.
  • No attribute on a deck page says when a cue fires. The page owns what a thing looks like and what it means, and cues.json owns when it happens.

What would change it

A voice or an aligner that returns phoneme or character timings rather than word timings would widen what a cue may name. Nothing else would.