The top four shot techniques are all about people — none of them about camera movement
When we built this schema we listed 28 shot techniques, from camera movement (dolly, pan, aerial) through performance (emotion, listener reaction) to technical devices (first/last frame, transformation effects).
After tagging 1,274 cases, the leaderboard did not look the way we expected. The top entries are not the technically difficult ones. They are the ones about people.
Where the data comes from
Technique tags are automatic candidates plus human review; unreviewed ones are marked "unreviewed" (225 cases, 17.7%) and excluded here. A case can carry several technique tags, so the values sum to more than the case count.
Finding 1: The top four all point at presenting a person
| Technique | Cases | Share of corpus |
|---|---|---|
| Environment reaction | 432 | 33.9% |
| Close-up | 403 | 31.6% |
| Emotional performance | 315 | 24.7% |
| Lip-sync | 313 | 24.6% |
| On-screen text | 287 | 22.5% |
| Aerial / high-low angle | 272 | 21.4% |
| Reference image | 252 | 19.8% |
| Action causality chain | 216 | 17.0% |
| Cut timing | 216 | 17.0% |
| Dolly in/out | 204 | 16.0% |
| … | ||
| Three or more people | 32 | 2.5% |
| Static camera | 30 | 2.4% |
| Listener reaction | 21 | 1.6% |
| First/last frame | 8 | 0.6% |
The top four are: environment reaction (a person reacting to their surroundings), close-up (shooting the person large), emotional performance (giving them an expression), lip-sync (making them speak). All four solve the same underlying problem — making the character look real.
Meanwhile the archetypal camera-movement techniques (aerial 272, dolly 204, orbit 158, tracking 144) are far from rare, but none of them breaks the top four.
Finding 2: Grouped into two families, "people" runs at nearly double "camera"
We sorted all 28 techniques into two semantic families and summed the counts:
| Family | Techniques included | Total |
|---|---|---|
| Presenting people | environment reaction, close-up, emotional performance, lip-sync, action causality chain, prop interaction, listener reaction, two-person dialogue, three-or-more, identity reference | 2,097 |
| Camera movement | aerial/high-low angle, dolly, orbit, tracking, static camera, multi-shot switching, continuous single shot | 1,068 |
(Note: these are tag occurrences, not case counts; a case can sit in both families.)
About 2 to 1. In the public corpus, creators spend twice as much attention on making characters hold up as on making the camera look good.
That matches practical experience: what breaks most often in AI video is faces and limbs. Once they break, no amount of camera craft saves the shot. Conversely, if the character is convincing, a static camera is fine.
Finding 3: The coldest two entries are multi-person blocking and first/last frame
The bottom of the leaderboard deserves a separate look:
| Technique | Cases | Note |
|---|---|---|
| First/last frame | 8 | Fewest of all 28 |
| Listener reaction | 21 | The cut to whoever is listening in a dialogue scene |
| Static camera | 30 | A camera that does not move |
| Three or more people | 32 | Three-plus characters in frame |
"First/last frame" has only 8 cases, yet the corresponding generation mode (first/last-frame-to-video) has 49. That is, most cases produced with that mode never got the technique tag — the two fields use different scopes. We count that as our own error.
More telling is "three or more people" at 32 and "listener reaction" at 21. These low values say the public corpus barely touches multi-person blocking — and multi-person framing happens to be where AI video fails most often (character merging, interpenetrating limbs). Where it breaks most, public reference is scarcest.
Finding 4: Tagging density varies 2.5x across genres
Each case carries an average of 3.61 reviewed techniques, but genres differ widely:
| Genre | Mean techniques | Sample |
|---|---|---|
| Transition editing | 6.65 | 43 |
| Wuxia / fantasy martial arts | 6.00 | 4 |
| Chase / escape | 5.56 | 16 |
| Action & physics | 5.09 | 70 |
| Martial arts combat | 4.58 | 71 |
| Short-drama dialogue | 4.11 | 224 |
| … | ||
| Music & dance | 2.96 | 121 |
| Product ad | 3.01 | 192 |
| Everyday life | 2.70 | 54 |
Action and transition genres are tagged much more densely. Reasonable enough — they are composed of multiple technical steps (weapon state, transformation effects, action causality chains), so there is more to tag. Product ads sit at 3.01 and music & dance at 2.96 — precisely the two categories with the largest distribution volumes.
What you can do with this
If you write short-drama prompts: focus on the top four techniques; they dominate the public corpus. Concretely: give characters a reaction to their environment, shoot tighter, put emotion on the face, match the lips. Get those four right and style plus camera work are secondary.
If you need multi-person scenes: know going in that public reference is nearly non-existent (three-or-more has 32 cases). You will have to build your own test set rather than copy existing cases.
If you annotate or build datasets: note the mismatch between the first/last-frame tag (8) and the first/last-frame generation mode (49). The same concept must be scoped identically across fields before you cross-tabulate, or every intersection is noise.
Method and limitations
- Tag boundaries are not strict: "environment reaction" and "emotional performance" can both apply, and different annotators will differ. We have not run inter-annotator agreement testing (that debt is written up separately).
- Correlation is not importance: a high tagging density may reflect "easier to tag" rather than "creators care more." We cannot rule either explanation out.
- Filling the 225 unreviewed cases could reorder the list, especially in the mid-range (17%–25%).
- All figures use the 1,274 index records.
Reproduce it
The technique filter supports case-by-case checking and multi-select intersections: https://cases.aishifu.shop/