aishifu Research H3 Case Library About Contact 中文

88% of our classifications aren't based on the picture: how these 1,274 records were tagged

By Alpha Lay ·

An earlier piece covered the 95% of records sitting in "pending review." That one was about review not happening. This one is more fundamental: even where classification exists, the basis was not the picture.

Where the data comes from

Each case carries a basis field recording where that record's classification came from. There are six values:

basisMeaningCasesShare
titleBased on the case title57545.1%
metadata_onlyMetadata only, no content basis54042.4%
prompt_onlyBased on prompt text1028.0%
noneNo basis221.7%
synthetic_titleThe title itself was generated201.6%
case_id_slugBased on an identifier string151.2%

Finding 1: 87.5% of classifications have no content basis

Adding the top two: 575 + 540 = 1,115 records (87.5%).

Only 102 records (8.0%) had their prompt actually read.

In other words, roughly 92% of genre and technique classifications in this library are not based on observing the video content.

Finding 2: Some of the titles were generated by us

Among the 575 title-based records, 20 carried titles we generated ourselves (synthetic_title).

That is a self-referential loop: unable to get original information, we generated a title, then classified using that title. Strictly speaking these records have zero basis — yet they carry equal weight in every table and in every earlier piece in this series.

Let us state it plainly: this is a methodological defect, not a matter of dataset size. Adding more records will not fix it.

Finding 3: Titles average 36 characters

1,146 cases have a title (90.0%), averaging 36 characters.

How much information fits in 36 characters? Here are the most frequent tokens:

TokenOccurrences
minimax61
anime49
video46
character39
commercial32
film30
study30
motion29
fashion29
cinematic28
luxury21
storyboard21

These are genre-label words (commercial, fashion, anime, film), not descriptions of what happens on screen. How accurate is classifying a video as "fashion & beauty" because its title contains "fashion"? We do not know — we have never run a sampled human check.

Which is exactly the problem: a 36-character title, unaudited, does not constitute a reliable classification basis.

Finding 4: Only 11 titles claim verification

We also checked for confidence markers:

KeywordCases
Contains "verified"11
Contains "official"3

Of 1,146 titles, only 11 claim any verification. In other words, this corpus carries almost no "this case has been confirmed" markers — including on our own records.

What you can do with this

If you cite any earlier piece in this series: treat this as a standing disclosure. Conclusions resting on objective measurement (framing, duration, resolution, publication time) are sound. Conclusions resting on classification (genre distribution, technique rankings, style convergence) all sit on a base that is 87.5% non-content. The direction may hold; the precision should not be trusted.

If you build your own annotation pipeline: this is the lesson to take. "Basis" must be a first-class field, and it must be displayed in reporting. We recorded it (as basis) but never stratified by it in any analysis. Had we filtered on basis early, we would have seen 87.5% immediately.

If you evaluate any auto-annotated dataset: asking "what is the basis for the labels" is far more useful than asking "how many records." A 100,000-record library with title-based labels may be worth less than a 1,000-record library where someone looked at every frame.

Method and limitations

Reproduce it

Every case's classification basis is visible in its detail view: https://cases.aishifu.shop/