49 first/last-frame cases carry 1.66x the tagging density and 2.1x the prompt length
Our generation-mode piece explained the 73% "unverified" rate through source channels. This piece looks the other way, at the 49 cases explicitly labelled first/last-frame-to-video — and they look nothing like the rest.
Where the data comes from
Mode was determined from three sources: statements in the original post, structural features in the prompt (whether it declares a first or last frame), and official labels. The 49 are cases where at least one of the three is explicit. The sample is small, so this piece describes shape only and does not infer cause.
Finding 1: 1.66x the tagging density
| Metric | First/last-frame | Library |
|---|---|---|
| Mean reviewed techniques | 6.00 | 3.61 |
| Sample | 49 | 1,274 |
About 6 technique tags per case, 66% above the library average.
That is a large gap. While "tagged more thoroughly" and "genuinely more complex" are logically distinct claims (perhaps the annotator was simply more careful), the next two findings favour the latter.
Finding 2: Prompts run 2.1x the library median
| Metric | First/last-frame | Library |
|---|---|---|
| Median prompt length | 3,279 chars | 1,563 chars |
2.1x. It is the longest of any mode — longer than reference-to-video (2,942) and 2.5x text-to-video (1,313).
Three data points together (many techniques, long prompts, few cases) point to one portrait: people using this mode are fine-tuning. They are producing content that needs precise control over timing, which is exactly what a first/last-frame mode exists to specify.
Finding 3: Genres are highly dispersed, with no dominant category
| Genre | Cases |
|---|---|
| Martial arts | 8 |
| Short-drama dialogue | 8 |
| Character performance | 7 |
| Tool comparison demo | 6 |
| Sci-fi | 5 |
| Fantasy & magic | 5 |
| Product ad | 5 |
| Landscape & nature | 5 |
49 cases spread across 8+ genres, the largest at just 8 (16%).
Other modes are not like this: reference-to-video concentrates on technical demos, text-to-video on short drama. First/last-frame has no dominant category — it is used wherever start and end must align, making it a general-purpose tool rather than a genre's speciality.
Its technique profile says the same:
| Technique | Cases |
|---|---|
| Reference image | 24 |
| Environment reaction | 23 |
| Lip-sync | 22 |
| Emotional performance | 22 |
| Cut timing | 21 |
| Action causality chain | 19 |
| Multi-shot switching | 15 |
The top seven span reference material, performance, editing and action — not one genre's techniques, but a full workflow's set.
Finding 4: It only took off in the later period
| Period | First/last-frame cases |
|---|---|
| Earlier half (before 08-07) | 10 |
| Later half | 37 |
3.7x. Among all techniques and modes, this is one of the fastest-growing over time (vertical framing grew 42% over the same span, 2D animation 165%).
Our interpretation (an interpretation, not a finding): early on people were working out "can it generate at all," whereas first/last-frame requires a two-stage flow — a first frame image, then a last frame image — which is a second-phase tool. Its later arrival in the case stream fits the general pattern that workflow maturity lags model availability.
What you can do with this
If you need precise start/end control (transitions, action continuity, title cards): these 49 are the only directly relevant public cases, and their prompt structure (a 3,000-character order of magnitude) is a usable length expectation — do not plan around the library median of 1,500, or you will under-write.
If you evaluate tooling: note that "first/last frame" appears 8 times in the technique field but 49 times in the mode field (an inconsistency we flagged in the technique piece) — the two fields' scopes must be aligned first, or a cross-tab will count one concept twice.
If you build products: the "two reference images plus a long prompt" combination says users need help writing long prompts. Unlike "write me one sentence" tools, this direction has data behind the demand.
Method and limitations
- Only 49 cases (3.8%): every ratio here is affected by small-sample noise. Do not treat them as a stable baseline.
- Mode detection may be biased: cases classifiable as first/last-frame tend to be those where the original post is clearly written (the prompt declares a first or last frame). So these 49 skew toward "better documented" — which may itself explain why their prompts are long. A selection effect we cannot fully rule out.
- "More techniques" does not imply "more complex": it may simply mean the annotator was more careful on these cases.
- No causal claim: whether "grew in the later half" relates to improved model capability is not judged here.
- All figures use the 1,274 index records.
Reproduce it
The library filters by generation mode for case-by-case inspection of prompts and technique tags: https://cases.aishifu.shop/