88% of our classifications aren't based on the picture: how these 1,274 records were tagged
An earlier piece covered the 95% of records sitting in "pending review." That one was about review not happening. This one is more fundamental: even where classification exists, the basis was not the picture.
Where the data comes from
Each case carries a basis field recording where that record's classification came from. There are six values:
| basis | Meaning | Cases | Share |
|---|---|---|---|
| title | Based on the case title | 575 | 45.1% |
| metadata_only | Metadata only, no content basis | 540 | 42.4% |
| prompt_only | Based on prompt text | 102 | 8.0% |
| none | No basis | 22 | 1.7% |
| synthetic_title | The title itself was generated | 20 | 1.6% |
| case_id_slug | Based on an identifier string | 15 | 1.2% |
Finding 1: 87.5% of classifications have no content basis
Adding the top two: 575 + 540 = 1,115 records (87.5%).
titlemeans: we read one sentence and classified the video's genre and techniques from it.metadata_onlyis blunter: no title either — only technical metadata like resolution, duration and host.
Only 102 records (8.0%) had their prompt actually read.
In other words, roughly 92% of genre and technique classifications in this library are not based on observing the video content.
Finding 2: Some of the titles were generated by us
Among the 575 title-based records, 20 carried titles we generated ourselves (synthetic_title).
That is a self-referential loop: unable to get original information, we generated a title, then classified using that title. Strictly speaking these records have zero basis — yet they carry equal weight in every table and in every earlier piece in this series.
Let us state it plainly: this is a methodological defect, not a matter of dataset size. Adding more records will not fix it.
Finding 3: Titles average 36 characters
1,146 cases have a title (90.0%), averaging 36 characters.
How much information fits in 36 characters? Here are the most frequent tokens:
| Token | Occurrences |
|---|---|
| minimax | 61 |
| anime | 49 |
| video | 46 |
| character | 39 |
| commercial | 32 |
| film | 30 |
| study | 30 |
| motion | 29 |
| fashion | 29 |
| cinematic | 28 |
| luxury | 21 |
| storyboard | 21 |
These are genre-label words (commercial, fashion, anime, film), not descriptions of what happens on screen. How accurate is classifying a video as "fashion & beauty" because its title contains "fashion"? We do not know — we have never run a sampled human check.
Which is exactly the problem: a 36-character title, unaudited, does not constitute a reliable classification basis.
Finding 4: Only 11 titles claim verification
We also checked for confidence markers:
| Keyword | Cases |
|---|---|
| Contains "verified" | 11 |
| Contains "official" | 3 |
Of 1,146 titles, only 11 claim any verification. In other words, this corpus carries almost no "this case has been confirmed" markers — including on our own records.
What you can do with this
If you cite any earlier piece in this series: treat this as a standing disclosure. Conclusions resting on objective measurement (framing, duration, resolution, publication time) are sound. Conclusions resting on classification (genre distribution, technique rankings, style convergence) all sit on a base that is 87.5% non-content. The direction may hold; the precision should not be trusted.
If you build your own annotation pipeline: this is the lesson to take. "Basis" must be a first-class field, and it must be displayed in reporting. We recorded it (as basis) but never stratified by it in any analysis. Had we filtered on basis early, we would have seen 87.5% immediately.
If you evaluate any auto-annotated dataset: asking "what is the basis for the labels" is far more useful than asking "how many records." A 100,000-record library with title-based labels may be worth less than a 1,000-record library where someone looked at every frame.
Method and limitations
basisis a field we wrote ourselves, and its accuracy has never been externally verified. This self-critique rests on our own self-report.titledoes not mean "unreliable": some titles are informative (naming shot, action, style). This piece argues that titles do not constitute a basis at the sample level, not that each one is wrong.metadata_onlyis not necessarily zero information: resolution, duration and framing are objective and can support some claims (e.g. "this is a landscape 15-second clip"). The problem is that they cannot support genre or technique classification.- Error rate is unknown: we have not sampled and re-checked, so we do not know how many of those 575 (or 1,115) classifications are wrong. That is the largest blank here, and it is our next outstanding debt.
- All figures use the 1,274 index records.
Reproduce it
Every case's classification basis is visible in its detail view: https://cases.aishifu.shop/