aishifu Research H3 Case Library About Contact 中文

45% of clips tagged 'lip-sync' were never checked for speech

By Alpha Lay ·

This is the least flattering piece about our own data.

We ran a round of speech detection (ASR) across the case library to answer something basic: does this video have anyone speaking in it? With that field you can separate talking-head content from pure-visual content, and you can evaluate things like lip-sync capability.

After the run we reconciled the results, and found that two dimensions that should corroborate each other disagree more than half the time.

Where the data comes from

The two are produced independently, which makes cross-validation meaningful.

Finding 1: For more than half the corpus, the question was never answered

Speech detection resultCasesShare
not_run_this_round70455.3%
no_segments (audio, no speech)31424.6%
dialogue_detected25119.7%
suspect_hallucination30.2%
unmatched20.2%

55.3% of cases have no speech conclusion at all. And the gap is unevenly distributed — it is worst exactly where it matters most:

Priority tierSampleNever processed
P0 (current short drama)263194 (73.8%)
P1 (cross-genre technique)441261 (59.2%)
P2 (future genres)570249 (43.7%)

P0 is our own definition of "most important to understand first" (largely short-drama dialogue). Three quarters of it does not even know whether anyone is speaking.

Finding 2: Of clips tagged "lip-sync", 45% were never verified

This is the core of the piece. We took every case carrying the "lip-sync" tag (313 of them) and looked at its speech result:

Speech result among lip-sync casesCasesShare
dialogue_detected (speech actually found)11938.0%
no_segments (audio, no speech)5016.0%
not_run_this_round (never checked)14044.7%
suspect_hallucination20.6%
unmatched20.6%

Of 313 lip-sync cases, 140 (44.7%) had no speech data backing the tag when it was applied.

So where did those 140 tags come from? Only from looking at the picture — it looked like someone was talking. But "looks like speech" and "there is speech on the audio track" are two different claims: a shot of a mouth opening and closing can sit on a purely musical track, or on no track at all.

Finding 3: The reverse direction disagrees too — only 47% of detected speech carries the tag

Flipping the cross-tabulation:

Of the 251 cases with detected speechCasesShare
Also tagged "lip-sync"11947.4%
Not tagged "lip-sync"13252.6%

More than half of the cases where speech detection says "someone is speaking" never received the lip-sync tag. Some of that is legitimate — voiceover, narration and singing can register as speech without involving lips. But at a scale of 132 cases, a real portion must be misses.

What you can do with this

If you want to evaluate lip-sync capability: the usable sample here is 119 cases (speech detected *and* lip-sync tagged), not 313. Using 313 as the denominator dilutes any conclusion.

If you are picking a model for talking-head content: note the 73.8% gap in the P0 tier. It means the best practices for talking-head content have barely been systematically organised in the public corpus. That gap is an opportunity, not a blocker.

If you annotate data yourself: this piece is a ready-made cautionary tale. Never let "looks like" stand in for "measured." 45% of our lip-sync tags lack measurement, and that is more than enough to invalidate the field.

Method and limitations

Reproduce it

Filter by speech result, or cross-reference against the lip-sync technique tag: https://cases.aishifu.shop/