TikTok captions, timed to the word
TikTok is watched with the sound off more often than not, and the first two seconds decide the rest. Captions are how the hook survives both. Ymotion times them to the word and sets the part that carries the hook larger than the part that leads up to it.
Word-level timing, from the speech itself
Captions change on the word rather than on a timer. The transcript carries the moment each word starts and ends, so a caption does not turn over mid-syllable, which is the thing a viewer notices without being able to say what is wrong.
The hook gets the size
Each caption is split into what leads up to the claim and the claim itself. The claim is set in the heavier face at a larger size and in the accent colour; the lead-in stays quiet. A caption that is all emphasis has nothing to stand against, so Ymotion refuses to make one.
Pacing that does not flinch
Per-stretch gestures are deliberately sparse. A video where every caption arrives differently is noise; most of a transcript is somebody talking, and it is left alone. The gesture changes where the feeling changes: a punchline landing, a question being asked, a let-down.
Questions
- Does Ymotion add captions word by word?
- Captions appear as short phrases of about five words, timed from the word-level transcript. One word at a time is readable only at slow speech and becomes a strobe at normal pace.
- Can I use it for videos that are not in English?
- Transcription is language-agnostic, so the words come back in whatever language was spoken. The caption faces cover Latin scripts; a script they cannot draw shows as empty boxes, and the editor says so.
- Is the export ready to upload?
- Yes. A finished MP4 at the clip’s own resolution and frame rate, encoded visually losslessly, with the captions rendered in.