90% of classifiable AI videos use one of just two visual styles
We defined 8 values for visual style: cinematic, photoreal, 2D animation, 3D animation, slow motion, stop-motion collage, black and white, plus "unreviewed."
After the statistics ran, the share of the top two values came out high enough that we re-checked the code. It wasn't a bug.
Where the data comes from
Style tags are automatic candidates plus human review. 421 cases (33.0%) carry the "unreviewed" marker, making this the slowest dimension to review of the four. To keep the conclusion free of missing-value contamination, every ratio below uses the 853 cases carrying at least one reviewed style tag as the denominator.
| Style | Cases | Share of corpus |
|---|---|---|
| Cinematic | 605 | 47.5% |
| Photoreal | 550 | 43.2% |
| Unreviewed | 421 | 33.0% |
| 2D animation | 193 | 15.1% |
| Slow motion | 176 | 13.8% |
| 3D animation | 107 | 8.4% |
| Stop-motion collage | 66 | 5.2% |
| Black and white | 24 | 1.9% |
Finding 1: 767 of 853 cases use one of only two answers
The top two values sum to 1,155 — more than the case count, because a case can carry both tags. So we de-duplicated:
| Condition | Cases | Share of classifiable (853) |
|---|---|---|
| Carries "cinematic" and/or "photoreal" | 767 | 89.9% |
| Carries neither | 86 | 10.1% |
90% of classifiable cases land on two visual answers.
Broken out further:
| Combination | Cases |
|---|---|
| Both cinematic and photoreal | 388 |
| Cinematic only | 217 |
| Photoreal only | 162 |
| Total | 767 |
The largest group carries both (388) — in most cases these two are two descriptions of one thing: shooting real people so they look real, with cinematic lighting and depth of field.
Finding 2: What the other 86 look like
The 10.1% that use neither answer have a clear composition:
| Style | Cases (within the 86) |
|---|---|
| 2D animation | 51 |
| Stop-motion collage | 23 |
| Slow motion | 19 |
| 3D animation | 9 |
| Black and white | 7 |
Those 86 are essentially the whole "animation / handmade texture" family. Which is to say the aesthetic map of the public corpus reduces to a stark binary: either it pursues looking like filmed cinema, or it is explicitly non-photoreal animation. The middle ground — illustration styles, picture-book, pixel art, collage experiments — barely exists across 1,274 cases.
Finding 3: Black and white appears 24 times, and not as a flex
We looked at the lowest value separately:
| Genre distribution of black and white | Cases |
|---|---|
| Music & dance | 6 |
| Character performance | 6 |
| Visual design | 5 |
| Fashion & beauty | 4 |
| Animation & games | 4 |
| Short-drama dialogue | 3 |
Not one of the 24 is a martial-arts clip. The bulk sits in music & dance and character performance — the categories that depend most on mood and least on information. In this corpus, black and white is not a stylistic experiment. It is a mood tool.
What you can do with this
If you are setting a visual direction for a new project: knowing which two answers 90% of people chose is itself decision input. Picking "cinematic" puts you in an extremely crowded lane where you must differentiate through content and hooks. Picking outside the two — stop-motion collage, at 66 cases — gives you very little reference but inherently higher recognisability.
If you wonder why your output "looks like everyone else's": the cause may not be the details of your prompt but the level of style selection itself. A 90% convergence means the space of style vocabulary has already been compressed very small.
If you evaluate video models: almost all public benchmarking currently happens inside the "cinematic plus photoreal" combination. Judging an animation model by that aesthetic will systematically underrate it.
Method and limitations
- The 421 unreviewed cases are the biggest risk: 33% of samples are unreviewed. If their style composition differs from the reviewed set, these ratios will move. We cannot currently verify this.
- Tags are "candidates plus review", not strict double annotation. Different annotators may place the "cinematic" threshold differently.
- "Cinematic" is not defined precisely in the schema — a real practical ambiguity we list as a to-improve item.
- All figures use the 1,274 index records; after de-duplication by original publication (924 cases) the convergence conclusion points the same way.
Reproduce it
Filter by visual style, optionally crossed with genre: https://cases.aishifu.shop/