TL;DR
On information-dense video, TapVid leads the field.
This AI video benchmark is built for video that must inform — instructional, editorial, explanatory — not entertain. On the information-dense scenes it targets, TapVid takes the outright Factual Coverage lead, ties for first on TeachQuiz Information Delivery, and tops Overall Quality across seven systems. The independent, industry-standard automated metric, VBench, reproduces the human panel’s preference, confirming the eval’s rigor.
01
AI video benchmark verdict: first on the content that matters — balanced everywhere else
Information-dense scenes — professional, news, how-to, education — are the benchmark’s target domain, and they are where TapVid wins. Across all 25 scenes, TapVid is the clear 2nd of 7 on Informative Quality and the only code renderer strong on both informative axes .
Overall Quality — informative + aesthetic 75.2 T #1 of 7 on information-dense work — the only system high on every axis at once.
Aesthetic Preference vs Fable 5, the informative leader 58.7% Preferred over the field’s top informative system in side-by-side viewing — and significantly does-not-lose.
Factual Coverage — information-dense scenes 83.6% The outright field lead. Significantly ahead of 5 of 6 competitors at getting required facts on screen.
TeachQuiz Information Delivery — information-dense scenes 84.5% Tied for first. Viewers answer comprehension questions as well after TapVid as after any system in the field.
OVERALL QUALITY — INFORMATIVE + AESTHETIC · INFORMATION-DENSE SCENES
*(Content Fidelity + Visual Quality + Aesthetic Preference) / 3 · field-relative T-score (70 = field mean, 10 = 1 SD)*

Both frames are reported everywhere. Headline claims use the information-dense frame (UC1–UC4) the benchmark targets — shown here, where TapVid tops Overall Quality; on all 25 scenes TapVid is 2nd of 7 on Informative Quality (T 78 vs Fable 5’s 85). Full scorecards with confidence intervals are in the appendix .
THE BENCHMARK
02
Built for videos that must inform, not entertain
Each of five use cases draws on real-world source material and is evaluated on five intent-driven scenes: the brief specifies the facts a viewer should walk away with, and human annotators measure whether they do. The target is high-information-density, clear-intent video — instructional, editorial, explanatory — not aesthetic or entertainment content.
THE FIELD — SEVEN SYSTEMS, TWO PARADIGMS
| System | Harness / model | Family |
|---|---|---|
| TapVid | code-rendered video | code-render |
| Fable 5 | Claude Code · Fable 5 | code-render |
| Code-render H | Claude Code · Sonnet 5 | code-render |
| Code-render R | Codex · GPT-5.5 | code-render |
| Pixel-diffusion L | official studio · Pro mode, 1080p | pixel-diffusion |
| Seedance | Seedance 2.0 · Jimeng · VIP mode, 720p | pixel-diffusion |
| Veo | Veo 3.1 · Google Flow · Quality mode, 720p | pixel-diffusion |
All systems accessed 2026 Jul 6–12 : locally-run tools frozen at their latest release as of Jul 6, hosted systems at the latest flagship model on their official platform — every system at its maximum quality setting, 16:9. Exact version pins with sources are in appendix A .
FIVE USE CASES, REAL-WORLD SOURCE MATERIAL
| Use case | Source material |
|---|---|
| Professional Enterprise / self-learning | GDPval tasks — Professional & Technical Services, Government, Manufacturing, Finance |
| News Editorial | 2025–26 covers: Economist, FT, Bloomberg, Nature, Science |
| How-To Lifestyle | Top #HowTo YouTube — craft/DIY, art, fitness, mindset, puzzles |
| Explainer Education | Most-viewed “AP Exam explained” — history, psych, calculus, biology, geography |
| Marketing Product / sales | NASDAQ most-valuable-company launch videos across sectors |
Sampling rules fixed before generation. Each use case draws its five scenes by a predefined rule — cluster the source pool, randomly sample one per cluster (top industries, publications, content categories, AP subjects, NASDAQ sectors). Every brief is a fully-detailed video-generation prompt carrying the complete information payload, authored and frozen before any system generated anything ; its required-fact checklist (168 facts across the benchmark) and comprehension questions (149) were authored at the same time and human-reviewed. A complete worked example is in appendix B .
HOW THE FIELD WAS RUN — ONE SHOT, SAME BRIEF, TOP SETTINGS
- Identical inputs: every system received the same brief verbatim — zero per-system prompt engineering.
- Best-of-1: one generation per scene, never curated; hard failures retried until success (at most 3 retries observed).
- Durations as generated: target lengths set by the brief; each system’s actual output used as-is (system averages 1.1–4.2 min).
- The 173/175 accounting: Fable 5’s safety rules declined 2 of its 25 briefs. Rather than substitute its Opus fallback, we scored Fable 5 only on its own 23 generations — no competitor was quietly weakened, and missing cells are dropped from denominators, never imputed.
THE ANNOTATOR PANEL — INDEPENDENT, BLIND, FULLY CROSSED
- 12 independent annotators , outsourced freelancers unaffiliated with any system’s team — each holding a master’s degree or above from a QS World Ranking top-50 university, with certified English competence.
- Blind throughout: videos presented unlabeled in randomized order; annotators were never told which system produced a video.
- Fully crossed: all 12 annotators rated all 7 systems, two raters per video — a strict or lenient rater moves every system equally and cannot bias a comparison.
- Deliberate viewing, not click-work: ~20 minutes of annotation per video of under 5 minutes, paid at 4× local minimum wage. Annotated Jul 13–15, 2026.
Metrics grounded in published methods, measured by people
Every video is scored by a fully-crossed annotator panel — every annotator rated every system, so a strict or lenient rater moves all systems equally and cannot bias a comparison. The panel produced 6,760 human annotations across five stations grounded in published methods — three of them peer-reviewed at ICML, WACV, and IEEE TPAMI / CVPR — and every video is *also* scored by the VBench automated model: >3.7 million frame-level model evaluations .
- TeachQuiz Information Delivery (Code2Video, ICML 2026 · arXiv:2510.01174) — how much of the intended information a viewer actually acquires.
- Factual Coverage — of the facts the brief requires, how many appear on screen, verbatim and legible.
- T2VTextBench Text Accuracy (arXiv:2505.04946) — correctness of on-screen titles, numbers and labels.
- T2VWorldBench ( WACV 2026 · arXiv:2507.18107) — four quality dimensions: Video Quality, Realism, Relevance, Consistency.
- Aesthetic Preference — arena-style, side-by-side comparisons against TapVid.
- VBench (VBench++, IEEE TPAMI 2025 · original VBench, CVPR 2024 Highlight · the model behind the public HuggingFace leaderboard) — industry-standard automated video-quality suite, >3.7 million frame-level model evaluations .
VBench is an independent, machine-based cross-check of our human annotations : its Frame Quality evaluation reproduces our human preference ranking, so the two instruments corroborate each other ( Finding 05 ).
The rankings are robust to how you compute them: three independent estimators — raw rates, leave-one-out information gain, and mixed-effects logistic models — produce the identical ordering, and reliability-weighted and PCA variants leave it unchanged. Because Marketing behaves as a different task family from the four information-dense use cases, every result is reported in two frames : all 25 scenes, and the information-dense frame (Professional, News, How-To, Explainer). Full tables for both are in the appendix .
FINDING 01 · THE FULL-FIELD FRAME
03
Informative video is a trade-off frontier — and TapVid holds the balance point
WIN The field is a trade-off frontier: on Informative Quality the order is Fable 5 > TapVid > everyone else; on Aesthetic Preference the order inverts exactly. TapVid is the only code renderer high on both informative axes — Content Fidelity 77, Visual Quality 79, 2nd of 7 on each — the balance point of the frontier.
Aesthetic Preference is measured arena-style — two videos of the same brief, side by side — and captures something the informative rubrics do not: its correlation with every informative metric is negative-to-zero (−0.22 to −0.01). Code renderers buy informative quality; pixel-diffusion buys aesthetic appeal. The frontier below places all seven systems on that trade-off, in both frames.
THE INFORMATIVE–AESTHETIC FRONTIER
*x: Informative Quality · y: Aesthetic Preference · field-relative T-scores (70 = field mean) · information-dense scenes*

TapVid is the only system in the green quadrant — highly informative and aesthetically preferred — in both frames. The frontier runs downhill: Informative Quality orders Fable 5 > TapVid > the rest, and Aesthetic Preference inverts it end-for-end. On information-dense scenes TapVid moves to the frontier’s elbow (Informative Quality 82, within 3.5 T of the leader) while the pixel-diffusion aesthetic edge shrinks to parity; on all 25 scenes Seedance and Pixel-diffusion L sit above the preference range shown. Aesthetic Preference is reported beside Informative Quality, never averaged into it.
AESTHETIC WIN-RATE OF TAPVID — SIDE-BY-SIDE VIEWING, ALL 25 SCENES
*share of side-by-side comparisons TapVid wins (ties count half) · ×100*

Above 50 = viewers prefer TapVid. Significance from the mixed-logit non-loss / decisive-win tests; full CIs and the non-loss frame are in the appendix .
AESTHETIC WIN-RATE OF TAPVID — SIDE-BY-SIDE VIEWING, INFORMATION-DENSE SCENES
*share of side-by-side comparisons TapVid wins (ties count half) · ×100*

On the benchmark’s target domain the aesthetic deficits shrink to parity — only Seedance remains significantly preferred on decisive comparisons.
WIN Shown side-by-side against Fable 5 — the field’s informative leader — viewers prefer TapVid. Win-rate 58.7%, and TapVid significantly does-not-lose (P = 0.72). The system that beats TapVid on quality composites loses to it in front of viewers.
WIN On information-dense work, TapVid concedes nothing on aesthetics while leading content. Overall non-loss rises to 53.0% — parity — and no pixel-diffusion system holds a significant non-loss advantage there. Against Google’s Veo: ahead on Informative Quality (78 vs 65), significantly ahead on coverage, and a preference toss-up.
LIMITATION Across all 25 scenes, aesthetic preference is a real weakness. TapVid’s overall win-rate is 40.7% — significantly below the 50% no-preference line — and it is significantly less preferred than Pixel-diffusion L and Seedance, the bottom of the informative ranking. The inversion is the paradigm trade-off, concentrated in Marketing; the close-out plan addresses it directly.
FINDING 02 · FACTUAL COVERAGE
04
Getting facts on screen is the binding constraint — and TapVid does it best
WIN On information-dense scenes TapVid puts 83.6% of required facts on screen — the outright field lead, significantly ahead of 5 of 6 competitors. Across all 25 scenes it is statistically tied with the leader Fable 5 and significantly ahead of every pixel-diffusion system.
Coverage is the benchmark’s most discriminating metric: the field splits into two clean tiers, with the four code renderers (74–79% of facts on screen) sitting well clear of the three pixel-diffusion systems (54–61%).
FACTUAL COVERAGE — SHARE OF REQUIRED FACTS ON SCREEN
*mean · 95% bootstrap CI over scenes · information-dense scenes · ×100*

Info-dense frame: TapVid 83.6 [77.3–89.0] leads outright ; the gap to every system except Fable 5 is statistically significant (mixed-logit 95% CI excludes 0). All-25 frame: Fable 5 79.0, TapVid 77.3 — a statistical tie — with Seedance, Veo and Pixel-diffusion L significantly behind in both frames.
COVERAGE DRIVES DELIVERY — ONE POINT PER VIDEO, 173 VIDEOS
*x: Factual Coverage · y: TeachQuiz Information Delivery (correct-rate) · ×100*

WIN Once a fact is on screen, viewers acquire it ~95% of the time. Coverage and delivery correlate at r = 0.70 across 173 videos, and the fitted line runs from a 36% no-fact floor to a 96% on-screen ceiling. Comprehension is nearly a solved step — the game is coverage, and TapVid plays it best.
COVERAGE BY USE CASE — MODEL-PREDICTED P(FACT ON SCREEN)
*mixed-logit predicted probability per use case · ×100*

WIN TapVid leads or ties for the coverage lead in four of five use cases. It leads News (81.9) and How-To (86.7) outright, and on Professional it is significantly ahead of four competitors — including the overall leader Fable 5.
LIMITATION Marketing is the exception. TapVid’s coverage drops to 53.8% there — bottom of the code renderers, behind Fable 5, Code-render R and Code-render H. The same gap shapes every metric’s Marketing column; the close-out plan treats it as one problem.
FINDING 03 · TEACHQUIZ INFORMATION DELIVERY
05
Viewers learn as much from TapVid as from any system in the field
WIN On information-dense scenes TapVid ties for first on TeachQuiz Information Delivery — viewers answer 84.5% of comprehension questions correctly — and significantly out-delivers three systems (Code-render H, Seedance, Pixel-diffusion L). Across all 25 scenes the top five are statistically inseparable; TapVid is significantly ahead of Pixel-diffusion L and Seedance.
Delivery follows the TeachQuiz protocol: after watching, annotators answer multiple-choice questions about the facts the brief required, and each video is scored against a per-question field baseline. Three independent estimators — raw correct-rate, leave-one-out information gain, and a mixed-effects logit — agree on the ordering.
TEACHQUIZ INFORMATION DELIVERY — COMPREHENSION CORRECT-RATE
*raw correct-rate · significance vs TapVid from mixed-effects logit · information-dense scenes · ×100*

TapVid is top-tier in both frames. Full per-use-case win/tie/loss tables are in the appendix .
DELIVERY BY USE CASE — COMPREHENSION CORRECT-RATE
*raw correct-rate per use case · bars with 95% bootstrap CI over the use case’s scenes · ×100*

TapVid is top-tier on News (88.3) and How-To (89.7) — per the mixed-logit inference it significantly out-delivers Code-render R, Pixel-diffusion L and Seedance on News and Veo on How-To, losing to none in either. Professional and Explainer are mid-field; Marketing is the gap the close-out plan targets. Each use case rests on 5 scenes, so the CIs are wide; significance calls come from the per-use-case mixed logit in the appendix .
T2VTEXTBENCH TEXT ACCURACY — ON-SCREEN TITLES, NUMBERS AND LABELS
*four-tier grade mapped to [0,1] · 95% bootstrap CI · information-dense scenes · ×100*

WIN TapVid renders on-screen text best on information-dense scenes. Across all 25 scenes the top of the field is statistically inseparable.
FINDING 04 · VISUAL QUALITY
06
Statistically tied with the leader on three of four visual dimensions
WIN On the T2VWorldBench human evaluation, TapVid is statistically tied with the leader Fable 5 on Realism, Relevance and Consistency — in both frames — and a clear 2nd of 7 on the four-dimension composite (86.4), significantly ahead of Seedance and Pixel-diffusion L. Whatever TapVid renders, it renders cleanly.
The four dimensions are rated 1–5 by the crossed annotator panel against per-dimension rubrics. The one significant visual gap to Fable 5 is Video Quality (86.0 vs 90.0); the other three dimensions are statistical ties.
T2VWORLDBENCH — FOUR DIMENSIONS, ALL 25 SCENES
*mean grade · 95% bootstrap CI over scenes · ×100 · * = significant vs TapVid*

TapVid is 2nd on Quality, Realism and Consistency, 3rd on Relevance (behind Fable 5 and Code-render R). Relevance — “does it say the right things” — behaves as a content dimension, and is analyzed with the content metrics above.
WIN Where raters fully agree, TapVid rates highest. On the exact-agreement subset TapVid’s Video Quality rises to 0.97, overtaking Fable 5 — TapVid’s clear wins are unambiguous.
FINDING 05 · INDEPENDENT CORROBORATION
07
Human annotations validated by VBench industry-standard automated metric
CHECK Run blind on the same 25 scenes, VBench — the automated model behind the public HuggingFace leaderboard — scores every video independently of the human panel. It orders the systems the same way our annotators’ side-by-side preference does. Spearman ρ = +0.89 between VBench Frame Quality and human preference; +0.87 leaving any one system out. The human preference axis is not an artifact of our raters — an independent, paper-grounded machine metric lands in the same order.
A widely-used automated model, computed with no knowledge of the human labels, agrees with the human panel on how the systems rank — evidence that the eval captures a real, reproducible signal rather than rater idiosyncrasy.
AUTOMATED AND HUMAN INSTRUMENTS AGREE — EVERY VIDEO SHOWN
*x: VBench Frame Quality per video (automated) · y: each system’s human win-rate · 141 system×scene videos, coloured by system · ×100*

Each small dot is one video’s automated Frame Quality, coloured by system and placed at that system’s human win-rate; the large dots are the system means. The coloured clouds stack in win-rate order and their centres climb left-to-right — automated frame quality and human preference rank the systems together ( Spearman ρ = +0.89 ). Systems are left unnamed: the point is the agreement of the two instruments, not any one system’s position.
SRC VBench (VBench++, IEEE TPAMI 2025 · original VBench, CVPR 2024 Highlight) is the standard automated video-quality benchmark behind the public HuggingFace leaderboard, scoring six dimensions per clip. It was run with the official VBench-Long code path (commit e946a8d) on dedicated GPUs on Jul 13–14 — completed before any comparison against the human labels existed. The Frame Quality head (Aesthetic + Imaging Quality) is the aesthetic cross-check used here, computed quality-when-it-works so a failed generation does not distort a system’s score. Every composite was recomputed from the raw per-scene scores and matches the source report to the decimal; the full run protocol is in appendix A.
THE NEXT WINS TO CLOSE
08
Two targets: Marketing content, and aesthetic appeal without an informative trade
Both gaps are measured, localized, and separable from what already works. Neither calls for touching the coverage-and-delivery engine that wins the information-dense domain.
TAPVID BY USE CASE — THE MARKETING GAP IS CONTENT, NOT RENDERING
*Content Fidelity vs Visual Quality composite · field-relative T-score · per use case*

On Marketing, TapVid’s Visual Quality holds at 75 while Content Fidelity drops to 51 — the videos render cleanly but miss required facts. Everywhere else the two axes move together at 68–93.
THE AESTHETIC PROGRAM — MOVE UP THE AESTHETIC AXIS, HOLD THE INFORMATIVE LEAD
*x: Informative Quality · y: Aesthetic Preference · T-scores · all 25 scenes*

Because preference is orthogonal to every quality metric, the move is straight up — raising motion, polish and appeal shifts TapVid toward the frontier’s empty top-right corner without spending any of its informative lead.
NEXT Win to close #1 — Marketing content. The collapse is fact delivery in marketing scenes, not rendering — a targeted extension of the coverage engine that already leads four use cases, not a rebuild. Closing it converts TapVid’s information-dense coverage lead into a full-field one: excluding Marketing, TapVid already tops five of six systems on coverage outright.
NEXT Win to close #2 — aesthetic preference vs pixel-diffusion, without giving back the informative lead. Preference is empirically orthogonal to every quality metric, so closing it is its own program — motion, polish, appeal — not a trade forced against coverage or delivery. The information-dense frame shows the ceiling: where content is dense, TapVid already reaches preference parity while holding the content lead.
APPENDIX
09
Frozen protocol, seeded statistics, every result in full
System versions are pinned to the day with sources; sampling rules and fact checklists were fixed before any video existed; the annotator panel was blind and fully crossed (A–B). Three independent estimators agree on every ordering, all randomness is seeded, and an independent re-run reproduced every artifact (C–E). Complete leaderboards with confidence intervals for every metric, in both frames (F–K). Δ columns are mixed-logit log-odds contrasts vs TapVid unless noted; WIN / LOSS = the 95% interval excludes zero, from TapVid’s perspective; TIE = it does not.
KEY System naming. The body of this report refers to three systems by anonymized labels. Their real identities, stated here only: Code-render H = Hyperframes · Code-render R = Remotion · Pixel-diffusion L = LTX. The tables below use the real product names.
A · Experiment protocol — versions, timeline, annotation ledger
Timeline (2026). Briefs + fact checklists authored and frozen before Jul 6 → systems frozen Jul 6 → all generation Jul 6–12 → VBench-Long scoring Jul 13–14 → human annotation Jul 13–15 (first response Jul 13 09:41, last Jul 15 04:33) → statistics and human↔VBench alignment Jul 14–16. VBench scoring completed before any human-label comparison existed.
Version pins — locally-run tools at their latest release as of the Jul 6 freeze; hosted platforms at the latest flagship model during the window. All 16:9, maximum quality settings, storyboard/planning features as shipped by each platform.
| System | Version · platform · settings | Source |
|---|---|---|
| TapVid | internal build current at the access window | — |
| Fable 5 | Claude Code 2.1.202 · model claude-fable-5 (announced Jun 9, redeployed Jul 1) · defaults | announcement · changelog |
| Hyperframes | v0.7.37 · Claude Code · Sonnet 5 (released Jun 30) · defaults | release · Sonnet 5 |
| Remotion | v4.0.485 · Codex CLI 0.142.5 · GPT-5.5 (announced Apr 23; GPT-5.6 shipped Jul 9, after the freeze) · defaults | release · Codex · GPT-5.5 |
| Veo | Veo 3.1 (announced 2025-10-15) · Google Flow, storyboard planning by Flow · Quality mode, 720p | announcement · modes doc |
| LTX | LTX-2.3 (released Mar 2026) · LTX Studio, storyboard planning by LTX Studio using GPT-Image-2 · Pro mode, 1080p | release · GPT-Image-2 |
| Seedance | Seedance 2.0 (announced Feb 12) · Jimeng — ByteDance’s official consumer platform — Agent Mode, story-film skill · VIP mode (top quality), 720p | announcement · Jimeng |
Prompt provenance. All 25 briefs and their fact checklists were authored with Claude Opus on Claude Code, human-reviewed, and frozen before generation began. If AI authorship biases the briefs toward any system’s idiom, the beneficiary would be Fable 5 — a competitor — not TapVid.
The 6,760-annotation ledger. One row per annotator × scene × system × item: fact-check 2,326 (168 facts) + TeachQuiz comprehension 2,062 (149 questions) + T2VWorldBench four dimensions 1,384 (346 video×rater units × 4) + text accuracy 692 (346 units × grade + junk flag) + side-by-side preference 296 = 6,760 . Every count is recomputable from the raw export.
VBench run. Official VBench-Long code ( vbench2_beta_long , commit e946a8d ) on 5× NVIDIA L4, one shard per use case, ~6 h wall; environment pinned and reproduced from scratch on a fresh instance; per-cell result JSONs and run logs retained. Input normalization and the >3.7M evaluation arithmetic: appendix E .
B · Worked example — one brief through the whole pipeline
The brief (How-To scene 1): *“The Paper That Pops — A No-Glue Origami Fidget Toy”* , target 5:00, 16:9. Like every brief, it is a fully-specified production document, identical for all seven systems: art direction with an exact hex palette, lighting and camera language, a nine-scene shot list with full voice-over text and per-scene motion notes, and a single master prompt — e.g. *“…the three color-coded papers (indigo 6×6 outer box, amber 5×5 inside button, violet 9×4 spring), the shared blintz fold that builds both boxes… the no-glue lock where flaps tuck under adjacent edges…”* The full information payload lives in the brief, so coverage failures are attributable to the system, not the prompt.
Its required-fact checklist (authored with the brief, frozen before generation; station items rendered in English). A fact is credited only if it appears on screen verbatim and legible — annotators may replay frame-by-frame:
- Outer box paper: 6×6 cm
- Inner button paper: 5×5 cm
- Spring paper: 9×4 cm
- The no-glue promise: no glue, no tape, anywhere
- The key fold, by name: blintz fold (corners to center)
A TeachQuiz comprehension item (answered after viewing; the “video did not provide this” escape is scored incorrect): *“What is the key fold that both boxes share?”* — valley fold / blintz fold (corners to center) / squash fold / the video did not provide this information.
How it scored. On this scene both of TapVid’s raters credited 5/5 facts and answered 12/12 comprehension questions correctly. The same items caught real misses elsewhere in the field on this scene — including the fold-name question — which is the discrimination the benchmark is built to measure.
C · Statistical methods — estimators, seeds, and independent re-verification
Estimators. Binary stations (coverage, delivery): Bayesian mixed logistic regression value ~ system + (1|scene) + (1|item) , TapVid as reference. Graded stations (text, dimensions): scene-cluster-robust OLS with annotator fixed effects. Delivery additionally via leave-one-out information gain. Three independent estimators — raw rates, information gain, mixed logit — produce the identical system ordering in both frames.
Uncertainty. 95% CIs from scene-level bootstrap (resample the 25 scenes, B = 3000, fixed seed); significance = the 95% interval excludes zero; results reported as tiers where intervals overlap, never false-precision ranks. VBench: Skillings–Mack omnibus per dimension (method differences significant at p = 8e-18 to 1e-9 over 23 complete scene blocks), post-hoc pairwise Wilcoxon signed-rank with Holm correction. Human↔VBench alignment: Spearman rank correlation, stress-tested by 2000× bootstrap, leave-one-system-out, and permutation tests (all seeded).
Reproducibility — independently re-verified. All analysis randomness is seeded, raw data (the full annotation export and per-scene VBench scores) ships with the analysis, and each metric has its own one-command reproduce script under pinned dependencies (Python 3.14.5 · numpy 2.5.1 · pandas 3.0.3 · scipy 1.18.0 · statsmodels 0.14.6 · matplotlib 3.11.0 · openpyxl 3.1.5). An independent re-run in a fresh environment regenerated all 55 committed result artifacts: every deterministic output byte-identical, all 15 figures pixel-identical, and the variational mixed-model columns matching to ≤6e-06 log-odds — with no significance verdict changing anywhere .
D · Reliability, method, and limitations
Rater robustness by design. The panel is fully crossed — every annotator rated every system — so rater strictness cannot bias a between-system comparison. Adding an annotator random effect moves contrasts by at most 0.16 (delivery) / 0.31 (coverage) log-odds with no significance flips.
Inter-rater agreement per metric. Delivery: 73.6% raw agreement, Gwet AC1 0.59. Coverage: 67.2%, AC1 0.42 — the annotator effect is the dominant nuisance source (sd 0.98), so the absolute coverage level is rater-sensitive even though the comparison is not. Text: within-one-grade 91%, weighted κ 0.27 — the noisiest rating. WorldBench dimensions: within-one 79–85%, weighted κ 0.24–0.35. Aesthetic Preference: weighted κ 0.07 — aggregate win-rates over 296 comparisons are meaningful; individual judgments are not, which also attenuates the orthogonality correlations toward zero.
Limitations. Delivery is relative to this seven-system field and not comparable across benchmark runs with a different field. Comprehension is administered open-book; the 0.86 median ceiling compresses system differences, and a closed-book run is the highest-leverage sharpening. 25 scenes (5 per use case) with 2 raters per video bound per-use-case resolution — the top tiers do not fully separate. Coverage credits a fact only when verbatim and legible. Mixed-logit intervals are variational (anti-conservative); the large significant contrasts are safe. The arena’s star topology supports no competitor-vs-competitor ranking and no absolute aesthetic score. T-scores are field-relative conveniences, not absolute grades. A facts-per-second bandwidth station is defined but was not computed in this round.
E · VBench ↔ human alignment — method
Instruments. VBench Frame Quality is the mean of the model’s Aesthetic Quality and Imaging Quality dimensions, computed quality-when-it-works (a failed generation is excluded, not scored 0). Human preference is the side-by-side win-rate over TapVid. The correlation is over the six competitor systems — TapVid is the star-topology anchor of every comparison and so is not a plotted point.
| Alignment (VBench Frame Quality ↔ human preference) | Spearman ρ |
|---|---|
| Overall | +0.89 |
| Leave-one-system-out (worst case) | +0.87 |
| Professional | +0.94 |
| News | +0.77 |
| Marketing | +0.70 |
| How-To | +0.48 |
| Explainer | −0.33 |
The overall sign survives 2000× bootstrap resampling and leave-one-system-out; n = 6 systems bounds formal significance. Explainer is the one segment where viewers reward motion over frame-polish, so the frame-quality metric tracks a different cue there. Every VBench composite was recomputed from the raw per-scene scores and matches the source report to the decimal.
Run provenance and the >3.7M count. VBench-Long ( vbench2_beta_long at commit e946a8d ; VBench++, IEEE TPAMI 2025) was run on NVIDIA L4 GPUs on Jul 13–14, 2026 — before any human-label comparison existed. Inputs were normalized per scene so no system gains an encoding artifact: 720p cap (never upscaled), one uniform 24 fps frame grid, near-lossless re-encode. Six dimensions per video; four of them score *every frame* (24 model forwards/s each), motion smoothness runs 11.5/s and dynamic degree 7.5/s — 116 model forwards per second of footage. Across the 32,141 seconds of scored video: 3.73 million frame-level model evaluations . A further consistency re-check on length-trimmed builds (not counted above) verified that video length does not confound the comparison: native−trim deltas all < 0.0025.
F · Composite scorecard — Informative Quality and the three axes
Field-relative T-scores: 70 = the mean of these seven systems, every 10 points = 1 SD. Informative Quality = equal-weight mean of the Content Fidelity and Visual Quality axes. Aesthetic Preference is arena-derived and reported beside the informative axes, never inside them — it runs inverse to them. Overall Quality, (Content + Visual + Aesthetic Preference)/3, is the named information-dense scenario headline.
All 25 scenes
| System | Informative Quality | 95% CI | Content | Visual | Aesthetic Preference |
|---|---|---|---|---|---|
| Fable 5 | 85.4 | [78.3, 92.1] | 82.6 | 88.2 | 52.3 |
| TapVid | 78.2 | [69.3, 86.6] | 77.1 | 79.2 | 61.7 |
| Remotion | 72.3 | [61.5, 82.2] | 76.6 | 68.1 | 69.2 |
| Hyperframes | 71.8 | [59.3, 83.0] | 73.3 | 70.2 | 68.1 |
| Veo | 64.8 | [51.7, 77.0] | 66.0 | 63.6 | 74.6 |
| Seedance | 61.3 | [46.8, 75.1] | 60.3 | 62.2 | 83.1 |
| LTX | 56.3 | [41.8, 70.8] | 54.1 | 58.5 | 81.0 |
Information-dense scenes (Professional, News, How-To, Explainer)
| System | Overall Quality (C+V+P)/3 | Content | Visual | Aesthetic Preference | Informative Quality |
|---|---|---|---|---|---|
| TapVid | 75.2 | 83.6 | 80.3 | 61.7 | 82.0 |
| Fable 5 | 75.0 | 82.8 | 88.1 | 54.2 | 85.5 |
| Seedance | 73.5 | 71.7 | 69.7 | 79.1 | 70.7 |
| LTX | 70.0 | 60.2 | 74.8 | 75.1 | 67.5 |
| Remotion | 69.7 | 77.5 | 65.9 | 65.7 | 71.7 |
| Hyperframes | 69.0 | 72.9 | 68.3 | 65.7 | 70.6 |
| Veo | 68.4 | 71.0 | 61.7 | 72.4 | 66.3 |
Bootstrap CIs over the 25 scenes overlap widely: the tiers the data supports are Fable 5 · TapVid · Remotion/Hyperframes · the pixel trio. The Overall Quality TapVid–Fable 5 gap (75.2 vs 75.0) is within noise. The ranking is robust to method (PCA vs axis-mean), reliability weighting, and where Relevance is grouped. TapVid’s Aesthetic Preference is 0.5-anchored by the arena’s star topology; its per-use-case variation lives in the competitors.
G · Factual Coverage — both frames
All 25 scenes — share of required facts on screen, verbatim and legible; 95% bootstrap CI; mixed-logit contrast.
| System | Coverage | 95% CI | Δ vs TapVid | 95% CI | Result |
|---|---|---|---|---|---|
| Fable 5 | 0.790 | [0.717, 0.860] | +0.094 | [−0.183, +0.370] | TIE |
| TapVid | 0.773 | [0.699, 0.843] | — | ||
| Remotion | 0.759 | [0.668, 0.845] | −0.166 | [−0.417, +0.085] | TIE |
| Hyperframes | 0.745 | [0.642, 0.842] | −0.199 | [−0.448, +0.050] | TIE |
| Seedance | 0.611 | [0.506, 0.710] | −0.826 | [−1.052, −0.601] | WIN |
| Veo | 0.596 | [0.494, 0.697] | −0.867 | [−1.091, −0.642] | WIN |
| LTX | 0.541 | [0.433, 0.658] | −1.179 | [−1.400, −0.958] | WIN |
Information-dense scenes
| System | Coverage | 95% CI | Δ vs TapVid | 95% CI | Result |
|---|---|---|---|---|---|
| TapVid | 0.836 | [0.773, 0.890] | — | ||
| Fable 5 | 0.792 | [0.710, 0.863] | −0.291 | [−0.601, +0.020] | TIE |
| Remotion | 0.765 | [0.669, 0.857] | −0.520 | [−0.801, −0.239] | WIN |
| Hyperframes | 0.727 | [0.607, 0.835] | −0.681 | [−0.952, −0.409] | WIN |
| Seedance | 0.677 | [0.578, 0.770] | −0.922 | [−1.182, −0.662] | WIN |
| Veo | 0.601 | [0.497, 0.714] | −1.213 | [−1.464, −0.963] | WIN |
| LTX | 0.545 | [0.418, 0.671] | −1.553 | [−1.798, −1.307] | WIN |
Coverage ↔ delivery: Pearson r = 0.70 across 173 videos; pooled OLS gives P(comprehended | on screen) = 0.96 and a 0.36 off-screen floor; a mixed logit agrees (0.945 / 0.30). Coverage’s absolute level is rater-sensitive (see D); the between-system comparison is not.
H · TeachQuiz Information Delivery — both frames
All 25 scenes — raw comprehension correct-rate, mixed-logit predicted P(correct) at a typical scene/question, and contrast.
| System | Raw rate | Predicted P | Δ vs TapVid | 95% CI | Result |
|---|---|---|---|---|---|
| Fable 5 | 0.821 | 0.868 | +0.243 | [−0.091, +0.576] | TIE |
| Remotion | 0.812 | 0.857 | +0.147 | [−0.167, +0.461] | TIE |
| TapVid | 0.795 | 0.838 | — | ||
| Hyperframes | 0.789 | 0.834 | −0.028 | [−0.330, +0.274] | TIE |
| Veo | 0.785 | 0.830 | −0.052 | [−0.353, +0.249] | TIE |
| LTX | 0.715 | 0.756 | −0.509 | [−0.786, −0.233] | WIN |
| Seedance | 0.691 | 0.730 | −0.647 | [−0.918, −0.377] | WIN |
Information-dense scenes
| System | Raw rate | Δ vs TapVid | 95% CI | Result |
|---|---|---|---|---|
| Fable 5 | 0.846 | +0.056 | [−0.342, +0.453] | TIE |
| TapVid | 0.845 | — | ||
| Veo | 0.840 | +0.005 | [−0.368, +0.378] | TIE |
| Remotion | 0.832 | −0.067 | [−0.433, +0.299] | TIE |
| Hyperframes | 0.790 | −0.393 | [−0.731, −0.054] | WIN |
| Seedance | 0.786 | −0.423 | [−0.759, −0.087] | WIN |
| LTX | 0.752 | −0.650 | [−0.971, −0.329] | WIN |
Delivery is relative to this seven-system field (leave-one-out gain / field-baselined logit); administration is open-book, which compresses differences toward the 0.86 median ceiling. Three estimators (raw, gain D, mixed logit) agree on the ordering in both frames.
I · T2VTextBench Text Accuracy and T2VWorldBench dimensions
T2VTextBench Text Accuracy — four-tier grade on [0,1]; cluster-robust contrast vs TapVid. * = 95% CI excludes zero.
| System | All 25 | 95% CI | Info-dense | 95% CI | Δ all-25 |
|---|---|---|---|---|---|
| Fable 5 | 0.783 | [0.714, 0.850] | 0.750 | [0.681, 0.825] | +0.062 |
| TapVid | 0.740 | [0.660, 0.820] | 0.769 | [0.675, 0.856] | — |
| Hyperframes | 0.700 | [0.615, 0.785] | 0.706 | [0.613, 0.794] | −0.009 |
| Remotion | 0.695 | [0.610, 0.785] | 0.688 | [0.594, 0.781] | −0.050 |
| Seedance | 0.640 | [0.535, 0.740] | 0.700 | [0.594, 0.806] | −0.087 |
| Veo | 0.630 | [0.530, 0.730] | 0.650 | [0.531, 0.762] | −0.090 |
| LTX | 0.580 | [0.500, 0.660] | 0.619 | [0.537, 0.694] | −0.145* |
T2VWorldBench, all 25 scenes — mean grade per dimension and the paper composite. * = significant vs TapVid (cluster-robust OLS).
| System | Quality | Realism | Relevance | Consistency | Composite |
|---|---|---|---|---|---|
| Fable 5 | 0.900* | 0.909 | 0.891 | 0.909 | 0.902* |
| TapVid | 0.860 | 0.880 | 0.844 | 0.872 | 0.864 |
| Remotion | 0.800 | 0.860 | 0.860 | 0.824 | 0.836 |
| Hyperframes | 0.812 | 0.852 | 0.816 | 0.844 | 0.831 |
| Veo | 0.800 | 0.804 | 0.808 | 0.824 | 0.809 |
| Seedance | 0.808 | 0.808 | 0.772* | 0.796* | 0.796* |
| LTX | 0.792 | 0.784* | 0.688* | 0.792* | 0.764* |
The composite is internally valid (Cronbach α = 0.91) but folds one content-facing axis: Relevance correlates with Delivery/Coverage (+0.77) more than with its sibling visual dimensions (0.61–0.66). Info-dense frame: TapVid stays 2nd on every dimension; only the Quality and Composite gaps to Fable 5 remain significant.
J · Aesthetic Preference (arena-style) — both frames, both readings
Star topology: every comparison is TapVid vs one competitor on the same brief (296 comparisons). Win-rate counts ties as half; non-loss counts ties for TapVid; both are reported. Verdicts from binomial mixed logits vs the 0.5 no-preference line.
All 25 scenes
| vs | Win-rate | 95% CI | Non-loss | 95% CI | Verdict |
|---|---|---|---|---|---|
| Fable 5 | 0.587 | [0.482, 0.683] | 0.696 | [0.577, 0.808] | TapVid does-not-lose* |
| Hyperframes | 0.440 | [0.347, 0.531] | 0.560 | [0.458, 0.667] | toss-up |
| Remotion | 0.430 | [0.338, 0.519] | 0.540 | [0.433, 0.643] | toss-up |
| Veo | 0.380 | [0.297, 0.464] | 0.480 | [0.389, 0.571] | toss-up |
| LTX | 0.320 | [0.217, 0.417] | 0.380 | [0.267, 0.500] | TapVid loses* |
| Seedance | 0.300 | [0.206, 0.393] | 0.400 | [0.286, 0.500] | TapVid loses* |
| Overall | 0.407 | [0.35, 0.47] | 0.507 | [0.44, 0.57] | less preferred than the field* |
Information-dense scenes — overall win-rate 0.432 [0.37, 0.50], non-loss 0.530: parity. The significant non-loss deficits vanish; on strict decisive comparisons only Seedance remains significantly preferred.
| vs | Win-rate | 95% CI | Non-loss | 95% CI |
|---|---|---|---|---|
| Fable 5 | 0.569 | [0.450, 0.692] | 0.667 | [0.542, 0.800] |
| Remotion | 0.463 | [0.354, 0.568] | 0.575 | [0.458, 0.692] |
| Hyperframes | 0.463 | [0.361, 0.563] | 0.600 | [0.500, 0.714] |
| Veo | 0.400 | [0.304, 0.500] | 0.500 | [0.385, 0.607] |
| LTX | 0.375 | [0.250, 0.500] | 0.425 | [0.292, 0.545] |
| Seedance | 0.338 | [0.225, 0.450] | 0.425 | [0.286, 0.567] |
Preference is empirically orthogonal to the informative metrics: correlating TapVid’s per-video advantage on each metric with the arena outcome gives −0.22 to −0.01 (n = 148 cells) — all negative or near zero. Non-loss is the favourable reading (20% of comparisons are ties); both are shown for that reason.
K · Per-use-case standings — where TapVid wins and loses
Significant wins/losses vs TapVid per use case (95% interval excludes zero). Cells list competitors; “—” = none.
| Use case | Coverage: TapVid beats | loses to | Delivery: TapVid beats | loses to |
|---|---|---|---|---|
| Professional | Fable 5, Hyperframes, Veo, LTX | — | Fable 5, LTX | Seedance |
| News | Remotion, Veo, LTX, Seedance | — | Remotion, LTX, Seedance | — |
| How-To | Remotion, Hyperframes, Veo, LTX | — | Veo | — |
| Explainer | LTX, Seedance | Fable 5 | Seedance | Fable 5 |
| Marketing | Seedance | Fable 5, Remotion, Hyperframes | Seedance | Fable 5, Remotion, Hyperframes |
No system wins everywhere: Seedance leads Professional delivery yet collapses on Marketing (P 0.30); Veo ties the News delivery lead yet is worst-in-field on How-To. TapVid’s arena win-rate by use case: Professional 0.450, News 0.433, Explainer 0.429, How-To 0.417, Marketing 0.308. TapVid’s composite Informative Quality by use case: News 87, How-To 89, Explainer 81, Professional 71, Marketing 63.
10
Try TapVid
Turn supplied assets and approved copy into a reviewable explainer-video workflow. Explore TapVid.
11
References and citation
Supporting evaluation methods include T2VTextBench, T2VWorldBench, and the Text-to-Video Human Evaluation protocol.
Recommended citation (APA): Law, D. (2026, September 6). AI video benchmark: Accuracy vs aesthetics (Version 1.0). TapVid Research. https://tapvid.ai/blog/information-dense-ai-video-benchmark
@techreport{law2026aiVideoBenchmark,
author = {Derek Law},
title = {AI Video Benchmark: Accuracy vs Aesthetics},
institution = {TapVid Research},
year = {2026},
month = {September},
url = {https://tapvid.ai/blog/information-dense-ai-video-benchmark},
note = {Version 1.0; July 2026 benchmark run}
}12
About the author

Derek Law is CTO and co-founder of TapVid. His work is guided by a simple thesis: the constraint on AI is the interface, not intelligence. Previously, he worked on Alexa AI and GenUI for Amazon AGI.



