aishifu Research H3 Case Library About Contact 中文

49 first/last-frame cases carry 1.66x the tagging density and 2.1x the prompt length

By Alpha Lay ·

Our generation-mode piece explained the 73% "unverified" rate through source channels. This piece looks the other way, at the 49 cases explicitly labelled first/last-frame-to-video — and they look nothing like the rest.

Where the data comes from

Mode was determined from three sources: statements in the original post, structural features in the prompt (whether it declares a first or last frame), and official labels. The 49 are cases where at least one of the three is explicit. The sample is small, so this piece describes shape only and does not infer cause.

Finding 1: 1.66x the tagging density

MetricFirst/last-frameLibrary
Mean reviewed techniques6.003.61
Sample491,274

About 6 technique tags per case, 66% above the library average.

That is a large gap. While "tagged more thoroughly" and "genuinely more complex" are logically distinct claims (perhaps the annotator was simply more careful), the next two findings favour the latter.

Finding 2: Prompts run 2.1x the library median

MetricFirst/last-frameLibrary
Median prompt length3,279 chars1,563 chars

2.1x. It is the longest of any mode — longer than reference-to-video (2,942) and 2.5x text-to-video (1,313).

Three data points together (many techniques, long prompts, few cases) point to one portrait: people using this mode are fine-tuning. They are producing content that needs precise control over timing, which is exactly what a first/last-frame mode exists to specify.

Finding 3: Genres are highly dispersed, with no dominant category

GenreCases
Martial arts8
Short-drama dialogue8
Character performance7
Tool comparison demo6
Sci-fi5
Fantasy & magic5
Product ad5
Landscape & nature5

49 cases spread across 8+ genres, the largest at just 8 (16%).

Other modes are not like this: reference-to-video concentrates on technical demos, text-to-video on short drama. First/last-frame has no dominant category — it is used wherever start and end must align, making it a general-purpose tool rather than a genre's speciality.

Its technique profile says the same:

TechniqueCases
Reference image24
Environment reaction23
Lip-sync22
Emotional performance22
Cut timing21
Action causality chain19
Multi-shot switching15

The top seven span reference material, performance, editing and action — not one genre's techniques, but a full workflow's set.

Finding 4: It only took off in the later period

PeriodFirst/last-frame cases
Earlier half (before 08-07)10
Later half37

3.7x. Among all techniques and modes, this is one of the fastest-growing over time (vertical framing grew 42% over the same span, 2D animation 165%).

Our interpretation (an interpretation, not a finding): early on people were working out "can it generate at all," whereas first/last-frame requires a two-stage flow — a first frame image, then a last frame image — which is a second-phase tool. Its later arrival in the case stream fits the general pattern that workflow maturity lags model availability.

What you can do with this

If you need precise start/end control (transitions, action continuity, title cards): these 49 are the only directly relevant public cases, and their prompt structure (a 3,000-character order of magnitude) is a usable length expectation — do not plan around the library median of 1,500, or you will under-write.

If you evaluate tooling: note that "first/last frame" appears 8 times in the technique field but 49 times in the mode field (an inconsistency we flagged in the technique piece) — the two fields' scopes must be aligned first, or a cross-tab will count one concept twice.

If you build products: the "two reference images plus a long prompt" combination says users need help writing long prompts. Unlike "write me one sentence" tools, this direction has data behind the demand.

Method and limitations

Reproduce it

The library filters by generation mode for case-by-case inspection of prompts and technique tags: https://cases.aishifu.shop/