How to sync lyrics to music
Line-level sync in one click, beat-accurate polish on the timeline, and word-level timing when you want karaoke highlights.
Lyric timing has three levels of fidelity: lines that appear roughly with the vocal, lines that land exactly on the phrase, and individual words that light up as they are sung. Motion Text supports all three, and you can stop at whichever level your video needs.
Start rough, refine where it matters. Choruses and hooks deserve word-level care; verses usually read fine at line level.
Rough pass: Auto-sync
Paste lyrics one line per row, set the sung range (skip the intro), and Sync spaces every line evenly across it. This alone gets you 80% of the way on most songs.
Line pass: drag against the waveform
Play back and watch the waveform. Vocal phrases show up as energy bursts. Drag each line so it starts on its phrase; beat markers snap your eye to the grid.
Word pass: LRC or transcription
For karaoke highlights, import an .lrc file with word timestamps, or use Pro transcription to time every word from the audio itself. Words stay individually adjustable.
Check the edges
Trim each line to disappear just before the next arrives. A 0.1–0.2 second gap reads cleaner than butted cues. Frame-step with arrow keys to check entrances.
Auto-sync or by hand: which one to start with
Start with auto-sync every time, even when you intend to hand-time the whole thing. Auto-sync isolates the vocal from the mix and places every line you pasted against what it hears, so the words are already in roughly the right place and in the right order. Hand-timing from zero means placing every line twice: once to get it near, once to get it right.
The exception is a track auto-sync struggles with, and the pattern is predictable. Heavily processed or layered vocals give it less to isolate; long instrumental intros throw off the range if you have not marked where the singing starts; spoken-word and rap sections with dense syllables tend to land early. Roughly one song in five needs real manual work. When you hit one, use Tap to sync: play the track and tap once per line as you hear it, which produces better timings in one pass than dragging ever will.
- Auto-sync first, always. It costs one click and it never makes hand-timing harder.
- Set the sung range before syncing. An unmarked 20-second intro is the single biggest source of bad output.
- Tap to sync beats dragging for a song that needs redoing wholesale.
- Drag only for the last 10%, where you are moving individual lines a few tenths of a second.
Diagnosing drift: constant offset vs. accumulating error
When timing is wrong, it is wrong in one of two ways, and they have completely different fixes. Telling them apart takes ten seconds: check a line near the start and a line near the end.
If both are out by about the same amount, you have a constant offset. Every line is early or late by the same margin, usually because the audio has silence at the head that the analysis counted, or because you set the sung range a beat off. The fix is one action: select every lyric line and drag them together. Relative timing is untouched.
If the first line is close and the last line is badly out, the error is accumulating, and dragging everything will fix one end while breaking the other. This means the spacing itself is wrong — commonly a trimmed or extended intro, a tempo change mid-song, or a version of the track that differs from the one the lyrics were timed against. Do not chase it line by line from the start. Re-run auto-sync with the sung range set correctly, or split the song at the section that changes and sync each half against its own range.
One more effect that reads as bad timing but is not: text needs a moment to be read before it can feel simultaneous. Lines that are technically perfect feel a fraction late. Pull them 100–200 ms ahead of the vocal and they will land right, which is why finished lyric videos almost always sit slightly early on paper.
The LRC file format, explained
LRC is the plain-text format karaoke tools and music players use for timed lyrics, and it is worth understanding because it is the interchange format between everything. It is readable enough to fix in any text editor. Each line begins with a timestamp in square brackets, in minutes, seconds, and hundredths, followed by the lyric.
That is standard, line-level LRC: it tells a player when each line begins, and the line stays up until the next timestamp. It is all you need for classic lyric videos where whole lines appear together.
Line-level LRC — one timestamp per line:
[00:12.50]I remember the night
[00:16.20]The city was still awake
[00:20.05]And nothing had a nameEnhanced LRC: word-level timing for karaoke highlights
For karaoke highlighting, where each word lights up as it is sung, you need enhanced LRC (sometimes called the A2 extension). It keeps the line timestamp and adds an inline timestamp in angle brackets before each word. Motion Text reads both forms, and a tag covering several words simply times them as a group.
Two practical notes. Motion Text ignores LRC metadata tags — [ti:], [ar:], [al:] and, importantly, [offset:] — and reads only the timestamps themselves, so a file that relies on an offset tag needs it applied to the timestamps first or corrected with one drag after import. And because LRC marks only where lines start, the final line has no following timestamp to end it; it is held for three seconds on import, which is usually the one cue worth trimming by hand.
Enhanced LRC — one timestamp per word:
[00:12.50]<00:12.50>I <00:12.78>remember <00:13.30>the <00:13.55>night
[00:16.20]<00:16.20>The <00:16.44>city <00:16.90>was <00:17.15>still <00:17.50>awakeTaking your timings elsewhere
Timings you make in Motion Text are not trapped in Motion Text, and timings you made elsewhere are not trapped either. Anything that exports line-level or enhanced LRC — desktop karaoke tools, the timing editors built into music players, the lyric sync features in streaming apps that allow export — imports directly. So do SRT and VTT subtitle files, which is the usual route if your lyrics started life as captions on a video.
Going the other way, the practical export is the video itself: a rendered MP4 with the words burned in, which is what every platform wants anyway. If you need the timings as data for a karaoke player or another editor, hand-editing an LRC file against the timings you can read off the timeline is a five-minute job for a three-minute song, and the format above is all you need to know to write one.
- Imports: .lrc (line-level and enhanced), .srt, .vtt.
- Word-level highlighting needs enhanced LRC or Pro transcription; plain LRC and SRT carry line timing only.
- Metadata tags, including [offset:], are ignored — bake any offset into the timestamps.
- The last cue in an LRC file has no end time and is held for three seconds; trim it if that runs long.
Questions
What is an LRC file?
A lyric file format with timestamps, exported by many karaoke tools. Motion Text imports line-level and word-level (enhanced) LRC.
Can AI do the whole sync for me?
Pro transcription listens to the track and returns word-timed captions automatically, and does best on clear vocals. You keep full manual control over every result.
Why do my lyrics feel late even when timed correctly?
Perception: text needs a beat to be read. Start lines 100–200 ms before the vocal and they will feel simultaneous.
Can I sync lyrics without uploading my song anywhere?
Yes. Auto-sync and the waveform analysis run in your browser on your machine, and browser MP4 export does too. Your audio file does not leave your computer unless you deliberately queue a server render.
My whole track is out by half a second. Do I have to redo it?
No — that is a constant offset, the easy failure. Select all the lyric lines and drag them together; the relative timing you already fixed is preserved. Only drift that grows through the song needs re-anchoring, and that is usually a tempo or trimmed-intro problem, covered above.
Does Motion Text read the [offset:] tag in an LRC file?
No. Motion Text reads the timestamps themselves and ignores LRC metadata tags, [offset:] included. If your file relies on one, apply it to the timestamps before importing, or import as-is and shift every line together on the timeline, which takes one drag.