aishifu Research H3 Case Library About Contact 中文

Research notes on AI video cases

We indexed 1,274 publicly shared H3-generated videos one at a time — tagging genre, shot technique, visual style, aspect ratio, duration and generation mode — then wrote up what the statistics show, across 19 pieces.

01

Only 15% of 1,274 AI videos are vertical — and short-drama clips rank 16th of 18

We indexed 1,274 publicly shared H3-generated videos one by one. Vertical framing accounts for 14.9% of them. Stranger still: the largest genre by sample size, short-drama dialogue, has the third-lowest vertical rate at just 10.7%.

Read →
02

73% of AI video cases don't say how they were made — and the reason isn't the creators

930 of 1,274 cases have no verified generation mode. We assumed creators simply omit it. Splitting by source channel showed something else: official sources are 0% missing, aggregation mirrors are 100% missing.

Read →
03

74% of AI videos land between 14 and 16 seconds. That's not a choice — it's a default

Everyone says AI video runs "8 to 16 seconds." Re-bucketing 1,274 cases at 2-second intervals shows a far more extreme shape: 945 cases fall in 14–16s, and 832 of them sit inside a 0.2-second window between 15.0 and 15.2 seconds.

Read →
04

45% of clips tagged 'lip-sync' were never checked for speech

We ran speech detection across 1,274 cases; 704 were never processed. The harder problem is elsewhere: of the 313 cases tagged 'lip-sync', 140 were never checked — while of the 251 where speech was actually detected, only 47% carry the tag.

Read →
05

The top four shot techniques are all about people — none of them about camera movement

432 environment reaction, 403 close-up, 315 emotional performance, 313 lip-sync. Grouping 28 techniques into "presenting people" versus "moving the camera" shows the first group runs at roughly double the second. Creators care about keeping characters believable.

Read →
06

90% of classifiable AI videos use one of just two visual styles

Excluding unreviewed records, 767 of 853 cases carry either "cinematic" or "photoreal" — only 86 carry neither. The aesthetic convergence in AI video is not an impression; it can be counted.

Read →
07

We got the "224 to 4" ratio wrong: after de-duplication it is 153 to 2

We previously wrote "224 short-drama cases versus 4 wuxia cases." A de-duplication check found that 350 of the 1,274 index records — 27.5% — are cross-host mirror duplicates. After de-duplicating by original publication, 924 remain, and wuxia accounts for just 2 original videos.

Read →
08

67.6% of these clips sit on a single free hosting project

After de-duplication, 625 of 924 unique videos are served from one Railway project domain. Separately: the official tutorial demos have a median duration of 8.1s and a 38.6% vertical rate, against 15.2s and 14.7% for community work — two entirely different populations.

Read →
09

Our own backlog: 98.9% of our highest-priority tier has never been reviewed

Of 1,274 records, 1,212 are marked "searchable, pending semantic review", and 707 of those only ever reached the automatic-candidate stage. The P0 tier we defined as most important to understand first (263 cases, mostly short drama) is 98.9% unreviewed. This piece is about our own problems.

Read →
10

These 1,274 cases span just 47 days — and publication essentially stopped on August 12

Of the cases we indexed, 980 carry a resolvable publication time, and all of them fall between 2026-07-29 and 2026-09-14. The first 14 days produced 69.7%; the single-day peak was 137. August 11 had 24. August 12 had 1.

Read →
11

Before and after August 7 are two different corpora: 2D animation up 2.6x, prompts 36% longer

Splitting 980 cases at the median publication time reveals a systematic shift: 2D animation 9.2%→24.4%, multi-shot switching 11.7%→20.8%, median prompt length 1,498→2,035 characters, while weapon state and reference images decline.

Read →
12

Prompts get longer when you supply a reference image: 2,942 characters versus 1,313

We measured prompt length across 1,274 cases. Counter-intuitively, the modes that need reference material have the longest prompts: reference-to-video at a median 2,942 characters and first/last-frame at 3,279, against 1,313 for pure text-to-video.

Read →
13

91% of these cases never reach 1080p — because half of them were compressed to 720p by X

Only 117 of 1,274 cases (9.2%) have a short side above 720 pixels. A single spec, 1280x720, accounts for 50.3%. And native X video (video.twimg.com) has a median short side of just 360 pixels.

Read →
14

88% of our classifications aren't based on the picture: how these 1,274 records were tagged

Every case records where its classification came from. 575 rest on a title, 540 rest on metadata alone, and only 102 actually read the prompt. About 8% of the library was classified by looking at content.

Read →
15

Of 322 authors, the most prolific contributed 1% of the cases

Only half of the 1,274 cases carry author information; de-duplicated, that is 322 authors. The top four accounts have 13 cases each, and the top twelve together are under 10%. Supply in this space is extremely fragmented.

Read →
16

One video has seven numbers here: 43% of cases carry more than one identity

550 records in the library carry IDs assigned by two or more different indexing projects. At least eight namespaces are in play: sns, beatapi, apimodels, morphic, ecomimagelab, imaginevid, minimax3, anil. The same video has different names depending on where you look.

Read →
17

120 cases have prompts under 200 characters, and 70% of those are verbatim from the original post

120 indexed cases have prompts shorter than 200 characters, median 97. 57.5% of them carry provenance 'verbatim from the publisher', and 69 sit on one mirror project — so this is not our scraping dropping content, it is what the original posts contained.

Read →
18

49 first/last-frame cases carry 1.66x the tagging density and 2.1x the prompt length

Only 49 cases are explicitly first/last-frame-to-video, but they average 6.00 reviewed technique tags (library average 3.61) and a median prompt of 3,279 characters (library 1,563). And the later half produced 3.7x more of them than the earlier half.

Read →
19

Every genre has a signature technique: short drama runs on two-person dialogue, martial arts on weapon state

Measuring relative concentration of 28 techniques per genre surfaces a distinct combination for each. It also works in reverse: the library-wide leader "environment reaction" appears in only 20% of product ads and food clips, well below the 33.9% average.

Read →